論文の概要: ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
- arxiv url: http://arxiv.org/abs/2607.20092v1
- Date: Wed, 22 Jul 2026 12:47:29 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-23 18:51:38.078065
- Title: ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
- Title(参考訳): ENRAP-VL : 視覚・言語モデルにおける二重環境適応のための分類学的プローブ
- Authors: Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani,
- Abstract要約: 文脈エントレメント(Contextual entrainment)とは、その文脈が関連性、真性、意味があるかどうかに関わらず、入力中の補助的コンテキストを出力を引き出すモデルの動向である。
我々は8つのカテゴリーにまたがる1500項目を手作業でキュレートしたデータセットであるENTRAP-VLを紹介した。
我々は、その現象を厳格に調査できるように、その機器、それを動機づける分類学、それを可能にする評価プロトコルを提供する。
- 参考スコア(独自算出の注目度): 1.189955933770711
- License: http://creativecommons.org/licenses/by-sa/4.0/
- Abstract: Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.
- Abstract(参考訳): 文脈エントレメント(Contextual entrainment)は、その文脈が関連性、真性、あるいは意味のあるものであるかどうかに関わらず、入力中の補助的コンテキストを出力を引き出すモデルの動向である。
近年、一助言語モデルにおいて機械的記述が特定され、与えられている。
視覚言語モデル(VLM)でどのように現れるかは、対照的に、ほとんど検討されていない。
我々は,VLMにおける文脈制約の研究には,既存のテキストのみのベンチマークをマルチモーダル・セッティングに移植するよりも,手元にある項目を中心に条件が構築されている2つのモダリティ・インスツルメンツ(テキスト・ストリームの描写画像,ビジュアル・ストリームのテキスト・クエリ)を必要とする。
VLMへの移行は漸進的というよりむしろ現実的であると我々は主張する。
エントレーニングは、テキストと視覚的文脈によって独立して引き起こされる二重現象であり、前作の単調な世界知識のみの定式化に匹敵しない正確さの区別(これは、世界において可能とされている描写シーンの誤りである)を開放する。
この位置を具体的かつ実用的なものにするために, ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language) を導入し, 8つのカテゴリにまたがる1500項目を手作業で収集し, 2つの軸にまたがる分類,すなわち, 項目の関連と真実との関係を整理し, テキスト・エントレインメント・ストリーム (8つの文脈条件) と視覚・エントレインメント・ストリーム (3つの文脈条件) に分割した。
我々は,特定のモデルにおけるエントレーニングの測定を主張せず,その指標,モチベーションを動機づける分類,およびそれを可能にする評価プロトコルを提供し,コミュニティが厳密にこの現象を調査できるようにしている。
データセットとそのドキュメントを公開します。
関連論文リスト
- Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B [51.74607395697567]
In-Context Learning (ICL)は、大規模言語モデル(LLM)の興味深い能力である。
我々は5つの自然主義ICLタスクに対してGemma-2 2Bにおける情報フローを因果介入を用いて同定する。
このモデルでは,2段階戦略を用いてタスク情報を推論し,コンテキスト化-then-aggregateと呼ぶ。
論文 参考訳(メタデータ) (2025-03-31T18:33:55Z) - Vision-Language Models Struggle to Align Entities across Modalities [13.100184125419695]
クロスモーダルなエンティティリンクは、マルチモーダルコード生成のような現実世界のアプリケーションに必要な基本的なスキルである。
我々のベンチマークであるMATEは5.5kの評価インスタンスで構成されており、視覚シーンはテキスト表現と一致している。
現状のビジョン・ランゲージ・モデル(VLM)と人間をこの課題で評価し,VLMが人間と比べ有意に苦労していることを見いだした。
論文 参考訳(メタデータ) (2025-03-05T19:36:43Z) - Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks [0.8999666725996978]
本稿では,大規模な視覚言語モデル(VLM)によって生成されたテキスト記述を,高価な手作業による注釈コストを伴わずに補助的なモダリティとして統合する新しいRSSCフレームワークを提案する。
5つのRSSCデータセットの定量的および定性的な評価実験により、我々のフレームワークがベースラインモデルより一貫して優れていることが示された。
論文 参考訳(メタデータ) (2024-12-03T16:24:16Z) - VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning [66.23296689828152]
我々は、視覚・言語モデルの機能を活用し、文脈内感情分類を強化する。
まず、VLLMに対して、視覚的文脈に関連して、被験者の明らかな感情を自然言語で記述するように促す。
第二に、記述は視覚入力とともに、トランスフォーマーベースのアーキテクチャのトレーニングに使用される。
論文 参考訳(メタデータ) (2024-04-10T15:09:15Z) - Can Large Language Models Understand Context? [17.196362853457412]
本稿では,生成モデルの評価に適合する既存のデータセットを適応させることにより,文脈理解ベンチマークを提案する。
実験結果から, 事前学習された高密度モデルでは, 最先端の微調整モデルと比較して, よりニュアンスな文脈特徴の理解に苦慮していることが明らかとなった。
LLM圧縮は研究と実世界のアプリケーションの両方において重要度が高くなっているため、文脈学習環境下での量子化モデルの文脈理解を評価する。
論文 参考訳(メタデータ) (2024-02-01T18:55:29Z) - A Multi-Modal Context Reasoning Approach for Conditional Inference on
Joint Textual and Visual Clues [23.743431157431893]
共同文と視覚的手がかりの条件推論は多モーダル推論タスクである。
我々はModCRというマルチモーダルコンテキスト推論手法を提案する。
2つの対応するデータセットに対して広範囲な実験を行い、実験結果により性能が大幅に向上した。
論文 参考訳(メタデータ) (2023-05-08T08:05:40Z) - Understanding ME? Multimodal Evaluation for Fine-grained Visual
Commonsense [98.70218717851665]
モデルが、限られた評価データ資源のために、視覚的シーンと基礎となるコモンセンス知識を本当に理解しているかどうかは不明だ。
本稿では,視覚シーン,テキスト,関連知識に対するモデルの理解をテストするために,質問応答ペアを自動的に生成するマルチモーダル評価(ME)パイプラインを提案する。
次に、MEデータによるトレーニングが標準VCR評価におけるモデルの性能を高めることを示すために、さらに一歩踏み出します。
論文 参考訳(メタデータ) (2022-11-10T21:44:33Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。