論文の概要: Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body
- arxiv url: http://arxiv.org/abs/2609.34840v1
- Date: Mon, 28 Sep 2026 10:37:35 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 02:22:46.77867
- Title: Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body
- Title(参考訳): コントロールプリミティブとしての侵害:1体に展開するエージェントの知覚チャネルと侵害記憶
- Abstract要約: 本研究では, 固定重量政策のパラメータを, 身体の描画前に設定し, 生命の更新を行わないエフェポッチ1の設定について検討する。
私たちは、感覚のコストが、最も穏やかな仕事よりも、まだ感じられていない仕事への割り当てを移動したことを証明しています。
- 参考スコア(独自算出の注目度): 1.6158652096705033
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: An agent deployed in a single body cannot learn how fast that body wears, because every trial that would reveal its wear resistance wears the body it would protect. We study this \emph{epoch-one} setting, in which the parameters of a fixed-weight policy are set before the body is drawn and never updated in life. The agent carries a load-gated nociceptive channel and a memory that retains what was felt. We prove that felt cost moves the allocation to the best-\emph{paid} work not yet felt rather than the gentlest, that an agent without retention never sees the felt-cost constraint bind, and that the channel pays only where the threat is individually unpredictable, cheap to avoid and expensive to ignore. We measure per body, setting the agent with channel and memory against the same individual without them, where neither carries a schedule learned across lives. On $2{,}000$ simulated floor-layer knees, with wear anchored to published loss rates, feeling, retaining and substituting extends the working life from age $55.2$ to $59.6$ and raises career output from $33.7$ to $36.1$. $69.3\%$ of bodies gain and \textbf{none lose}. A body that feels but retains nothing past the day gains one of the $+4.4$ years, and retention carries the rest. A population-trained agent gains $+0.65$ years from the same channel at $-0.54$ output. The difference is what a species prior already supplies, and a single body has none. The two are related by an identity, the ablation mean reporting $(1-χ)$ of the per-body value with $χ$ the share a blind schedule already captures, so we report both. Where the regime map predicts value, a care robot sextuples its certified service life and a field-anchored fleet writes off $0.15$ of its machines instead of $0.55$. Where it predicts none, a rover gains little over blind caution, so the map holds in both directions.
- Abstract(参考訳): 1つの体に配備されたエージェントは、その体がどれだけ速く着るかを学べない。
本研究では,この「emph{epoch-one} 設定」について検討し,身体が描画される前に固定重み付けポリシーのパラメータが設定され,生命において更新されないようにする。
エージェントは、ロードゲートされた侵害受容チャネルと、感じられたものを保持するメモリを担持する。
我々は、感覚コストが、最良なemph{pay}作業への割り当てを、温和さよりもまだ感じていないこと、保留しないエージェントが、フィールコストの制約を決して見ていないこと、チャネルが個々の脅威が予測不能で、避けられやすく、無視する費用がかかる場合にのみ支払うことを証明した。
我々は、人体ごとの測定を行い、エージェントをチャンネルとメモリで、同じ個人に対してチャンネルなしで設定する。
2$,}000ドルのシミュレートされたフロアレイヤーの膝には、失業率、感覚、保持、置換により、労働寿命を55.2ドルから59.6ドルに延長し、キャリアのアウトプットを33.7ドルから36.1ドルに引き上げている。
699.3\%のボディゲインと \textbf{none lose}。
感じるが一日中何も持たない体は、+4.4ドルの年金の1つを獲得し、残りは保留される。
人口訓練されたエージェントは、同じチャンネルから$-0.54$の出力で$+0.65$の年収を得る。
違いは、すでにある種が供給しているものであって、単一の体には何も存在しないことだ。
両者はアイデンティティによって関連付けられており、アブレーションは1対1で1対1で、シェアは1対1で、ブラインドスケジュールは1対1ですでに取得している。
システムマップが価値を予測する場合、ケアロボットは認定されたサービス寿命を調整し、フィールド対応の車両は0.55ドルではなく0.15ドルを請求する。
予測できない場所では、ローバーは目が見えないほど注意を引いてしまうため、地図は両方向に保持される。
関連論文リスト
- Agora: Git as Shared Memory for Collective AutoResearch [54.31992252587096]
別々のセッションで作業する調査エージェントは、他の人が試したことと、構築可能な結果を知る必要がある。
各コミットは結果、洞察、仮説、検証、報告を記録し、それを以前の作業とリンクする。
約12日間に渡り,13人の言語モデル労働者がアゴラを重み移動問題の解決に利用したことを報告した。
論文 参考訳(メタデータ) (2026-09-16T04:00:54Z) - Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents [16.24087700490158]
大規模言語モデルエージェントは、自律ループとしてますますデプロイされる。
これは実装の詳細ではなく、構成の失敗であることを示す。
LoopHarnessは、ループレベルでの永続的非遅延安全状態を復元する。
論文 参考訳(メタデータ) (2026-08-27T13:52:31Z) - Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games [50.880420636090896]
9-player Werewolf環境において,隠れた役割に対する外部信頼状態を維持するための監査可能なフレームワークを構築した。
我々は,その効果を関連づけとして報告し,そのメカニズムを未解決として扱う。
論文 参考訳(メタデータ) (2026-07-12T16:03:30Z) - Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix [46.122203287541005]
目標条件付き予測器は、精度0.90$の精度で到達するが、これは知覚ではなく、インテンストラクションの書き起こしである。
ゴールを破って(0.90!to!0.27$, three seed)チャンスを逃し、予測アンカーに9.5%の時間を与える。
診断は、修正を規定する: ゴールを(プランナーのコストに属する)ダイナミクスから遠ざけ、エンフレッドパスを監督する。
論文 参考訳(メタデータ) (2026-07-08T02:38:43Z) - Forget to Improve: On-Device LLM-Agent Continual Learning via Budget-Curated Memory [7.254575731198877]
オンデバイス言語モデルエージェントは、重みを更新するよりも、取得したメモリでの経験を蓄積することで改善する。
エージェントのエクスペリエンス・メモリのライフサイクルを管理する1バイト当たりのネット値スコアであるsysを提案する。
論文 参考訳(メタデータ) (2026-06-23T19:42:07Z) - Hidden-State Privacy Has an Empty Middle [51.56484100374058]
すべてのフルランクガウス解放を$O(1)$ Fisher utility で表すと、マハラノビス信号が隠れた幅で直線的に成長する方向を認める。
スクラッチからトレーニングされたスプリットメモリトランスフォーマーは、[20, 33]$90MでG_mathrmMahに達し、固定言語損失ペナルティにおいて、30Mから1Bまでの同じ予算のGPTベースラインに対して6ドル~24ドルという優位性を維持する。
論文 参考訳(メタデータ) (2026-05-21T20:12:09Z) - Stochastic Shortest Path with Adversarially Changing Costs [57.90236104782219]
最短経路 (SSP) は計画と制御においてよく知られた問題である。
また, 時間とともにコストの逆変化を考慮に入れた逆SSPモデルを提案する。
我々は、この自然な逆SSPの設定を最初に検討し、それに対するサブ線形後悔を得る。
論文 参考訳(メタデータ) (2020-06-20T12:10:35Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。