論文の概要: Toward a mechanistic understanding of inference in visual cortex and diffusion models
- arxiv url: http://arxiv.org/abs/2607.15693v1
- Date: Fri, 17 Jul 2026 07:11:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-20 17:56:52.778249
- Title: Toward a mechanistic understanding of inference in visual cortex and diffusion models
- Title(参考訳): 視覚野と拡散モデルにおける推論の力学的理解に向けて
- Abstract要約: 一次視覚野(V1)における知覚的推論モデルについて述べる。
我々は,これらのリカレントダイナミクスを,認知的スコアマッチング目標と暗黙的微分を用いて効果的に訓練する。
このモデルは、極度の視覚的あいまいさの中で、拡張輪郭のような画像の特徴を回復する、非常に優れた装飾性能を示す。
- 参考スコア(独自算出の注目度): 46.86106767256413
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: We describe a model of perceptual inference in primary visual cortex (V1) equivalent to a minimal diffusion model whose function can be readily understood from its parameters. The model is based on sparse coding with a non-factorial prior over latent variables in the form of an unconstrained, pairwise interaction matrix, extending standard sparse coding inference to a general recurrent dynamical system. We efficiently train these recurrent dynamics using a denoising score-matching objective and implicit differentiation. After training on natural images, the learned interaction matrix mirrors the structure of horizontal connections in superficial layers of V1 that link neurons of similar orientation tuning. This model exhibits exceptionally good denoising performance, restoring image features such as extended contours amid extreme visual ambiguity, nearly matching the behavior of standard, black-box diffusion architectures in generalization regime. Owing to the model's simplicity, the network's Jacobian can be decomposed directly in terms of the interaction matrix between latent variables, revealing mechanistically how the recurrent dynamics assign high probability over a continuous family of natural structural deformations. Intriguingly, within this circuit, a large fraction of latent variables learn to disconnect from visual input altogether, essentially forming a hierarchical representation that appears to enforce global consistency among image features. Together, the model and results bridge two distinct domains: for neuroscience, it generates concrete, testable hypotheses regarding functional connectivity in recurrent neural circuits during perceptual inference tasks; for machine learning, it elucidates the internal mechanisms learned by diffusion models that allow them to generate infinitely many novel images from a finite training set.
- Abstract(参考訳): 本稿では,一次視覚野(V1)における知覚的推論モデルと,そのパラメータから容易に理解できる最小拡散モデルについて述べる。
このモデルは、非ファクトリ事前遅延変数によるスパース符号化を、制約のないペアワイズ相互作用行列の形でベースとし、標準スパース符号推論を一般的なリカレント力学系に拡張する。
我々は,これらのリカレントダイナミクスを,認知的スコアマッチング目標と暗黙的微分を用いて効果的に訓練する。
自然画像のトレーニングの後、学習された相互作用行列は、類似の向き調整のニューロンをリンクするV1の表層における水平接続の構造を反映する。
このモデルは、極度の視覚的あいまいさの中で、拡張輪郭などの画像特性を復元し、一般化体制における標準のブラックボックス拡散アーキテクチャの挙動とほぼ一致する、非常に優れた装飾性能を示す。
モデルの単純さにより、ネットワークのヤコビアンは潜伏変数間の相互作用行列によって直接分解され、リカレント力学が自然構造変形の連続族に対して高い確率を割り当てる機構が明らかにされる。
興味深いことに、この回路内では、少数の潜伏変数が視覚入力から完全に切り離すことを学び、本質的には画像特徴間の大域的な一貫性を強制するように見える階層的な表現を形成する。
モデルと結果は2つの異なる領域を橋渡しする: 神経科学では、知覚的推論タスク中に繰り返し発生する神経回路における機能的接続に関する具体的な、テスト可能な仮説を生成する; 機械学習では、拡散モデルによって学習された内部メカニズムを解明し、有限のトレーニングセットから無限に多くの新しい画像を生成することができる。
関連論文リスト
- Understanding Representation Dynamics of Diffusion Models via Low-Dimensional Modeling [29.612011138019255]
拡散モデルにおける一様表現ダイナミクスの出現について検討する。
この一様性は、ノイズスケールをまたいだデノイング強度とクラス信頼の相互作用から生じる。
分類タスクでは、単調力学の存在は拡散モデルの一般化を確実に反映する。
論文 参考訳(メタデータ) (2025-02-09T01:58:28Z) - A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data [51.03144354630136]
最近の進歩は、拡散モデルが高品質な画像を生成することを示している。
我々はこの現象を階層的なデータ生成モデルで研究する。
t$の後に作用する後方拡散過程は相転移によって制御される。
論文 参考訳(メタデータ) (2024-02-26T19:52:33Z) - DIFFormer: Scalable (Graph) Transformers Induced by Energy Constrained
Diffusion [66.21290235237808]
本稿では,データセットからのインスタンスのバッチを進化状態にエンコードするエネルギー制約拡散モデルを提案する。
任意のインスタンス対間の対拡散強度に対する閉形式最適推定を示唆する厳密な理論を提供する。
各種タスクにおいて優れた性能を有する汎用エンコーダバックボーンとして,本モデルの適用性を示す実験を行った。
論文 参考訳(メタデータ) (2023-01-23T15:18:54Z) - A simple probabilistic neural network for machine understanding [0.0]
本稿では,機械理解のためのモデルとして,確率的ニューラルネットワークと内部表現の固定化について論じる。
内部表現は、それが最大関係の原理と、どのように異なる特徴が組み合わされるかについての最大無知を満たすことを要求して導出する。
このアーキテクチャを持つ学習機械は、パラメータやデータの変化に対する表現の連続性など、多くの興味深い特性を享受している、と我々は主張する。
論文 参考訳(メタデータ) (2022-10-24T13:00:15Z) - Dynamic Inference with Neural Interpreters [72.90231306252007]
本稿では,モジュールシステムとしての自己アテンションネットワークにおける推論を分解するアーキテクチャであるNeural Interpretersを提案する。
モデルへの入力は、エンドツーエンドの学習方法で一連の関数を通してルーティングされる。
ニューラル・インタープリタは、より少ないパラメータを用いて視覚変換器と同等に動作し、サンプル効率で新しいタスクに転送可能であることを示す。
論文 参考訳(メタデータ) (2021-10-12T23:22:45Z) - The Role of Isomorphism Classes in Multi-Relational Datasets [6.419762264544509]
アイソモーフィックリークは,マルチリレーショナル推論の性能を過大評価することを示す。
モデル評価のためのアイソモーフィック・アウェア・シンセサイティング・ベンチマークを提案する。
また、同型類は単純な優先順位付けスキームによって利用することができることを示した。
論文 参考訳(メタデータ) (2020-09-30T12:15:24Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。