論文の概要: Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering
- arxiv url: http://arxiv.org/abs/2610.01637v1
- Date: Thu, 01 Oct 2026 13:04:59 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:24.142201
- Title: Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering
- Title(参考訳): ベトナム語視覚質問応答のための多層Fusing Transformerによる視覚・テキスト表現
- Abstract要約: VQA (Visual Question Answering) は、コンピュータが自然に画像に関する質問を理解し答えることを要求する研究分野である。
VQAの英語に関する広範な研究と開発にもかかわらず、他の言語、特にベトナム語に対する同様の取り組みはほとんど行われていない。
このギャップを埋めることで、ベトナムのVQAの分野は、人工知能の研究の多様性を豊かにするだけでなく、様々な分野における実践的な応用を可能にしている。
- 参考スコア(独自算出の注目度): 3.2434811678562685
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: In recent decades, artificial intelligence has made significant progress in understanding and interacting with images. One of the important applications of this technology is Visual Question Answering (VQA), a research field that requires computers to understand and answer questions about images in a natural manner. Despite extensive research and development in VQA for English, there have been very few similar efforts made for other languages, especially Vietnamese. This gap presents a significant challenge and opportunity for the advancement of VQA technology in the Vietnamese language context. By bridging this gap, the field of Vietnamese VQA not only enriches the diversity of research in artificial intelligence but also enables practical applications in various domains, such as education, healthcare, and entertainment, catering to Vietnamese-speaking populations worldwide. Thus, the exploration and development of Vietnamese VQA systems hold immense potential for advancing both research and practical applications in the intersection of computer vision and natural language processing. In this paper, we propose a Multi-layer Fusing Transformer model utilizing a cross attention module to combine multiple modality features of images and texts from different layers in an aggregated representation. Our architecture allows us extract information from low level to high level. Through detailed experiments and ablation studies, our model achieves promising results against the competitive baselines in ViVQA dataset for Vietnamese language.
- Abstract(参考訳): 近年、人工知能は画像の理解と相互作用に大きな進歩を遂げている。
この技術の重要な応用の1つに視覚質問回答(VQA)がある。
VQAの英語に関する広範な研究と開発にもかかわらず、他の言語、特にベトナム語に対する同様の取り組みはほとんど行われていない。
このギャップはベトナム語の文脈におけるVQA技術の進歩に大きな課題と機会をもたらす。
このギャップを埋めることによって、ベトナムのVQAの分野は、人工知能の研究の多様性を豊かにするだけでなく、教育、医療、エンターテイメントといった様々な分野における実践的な応用を可能にし、世界中のベトナム語を話す人口に対応している。
このように、ベトナムのVQAシステムの探索と開発は、コンピュータビジョンと自然言語処理の交差点における研究と実践の両方を前進させる大きな可能性を秘めている。
本稿では,多層Fusing Transformerモデルを提案する。
我々のアーキテクチャは、低レベルから高レベルまで情報を抽出することを可能にする。
ベトナム語に対するVVQAデータセットの競争的ベースラインに対して,詳細な実験とアブレーション研究により有望な結果が得られた。
関連論文リスト
- AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering [2.4577252294937444]
VQA(Visual Question Answering)は、モデルが視覚情報とテキスト情報を共同で理解する必要がある基本的なマルチモーダルタスクである。
近年の研究では、VQAタスクにおいて、大規模言語モデルによって自動評価と人的判断の整合性がさらに向上することが示唆されている。
論文 参考訳(メタデータ) (2026-03-10T13:57:52Z) - Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration [0.40964539027092917]
本研究は,ベトナムの視覚質問応答データセットを用いて実験を行うことにより,ギャップを埋めることを目的とする。
画像表現能力を向上し,VVQAシステム全体の性能を向上させるモデルを開発した。
実験結果から,本モデルが競合するベースラインを超え,有望な性能を達成できることが示唆された。
論文 参考訳(メタデータ) (2024-07-30T22:32:50Z) - CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark [68.21939124278065]
言語と文化の豊富なセットをカバーするために設計された、文化的に多言語なビジュアル質問回答ベンチマーク。
CVQAには文化的に駆動されたイメージと、4大陸30カ国の質問が含まれ、31の言語と13のスクリプトをカバーし、合計10万の質問を提供する。
CVQA上で複数のマルチモーダル大言語モデル (MLLM) をベンチマークし、現在の最先端モデルではデータセットが困難であることを示す。
論文 参考訳(メタデータ) (2024-06-10T01:59:00Z) - Language Guided Visual Question Answering: Elevate Your Multimodal
Language Model Using Knowledge-Enriched Prompts [54.072432123447854]
視覚的質問応答(VQA)は、画像に関する質問に答えるタスクである。
疑問に答えるには、常識知識、世界知識、イメージに存在しないアイデアや概念についての推論が必要である。
本稿では,論理文や画像キャプション,シーングラフなどの形式で言語指導(LG)を用いて,より正確に質問に答えるフレームワークを提案する。
論文 参考訳(メタデータ) (2023-10-31T03:54:11Z) - ViCLEVR: A Visual Reasoning Dataset and Hybrid Multimodal Fusion Model
for Visual Question Answering in Vietnamese [1.6340299456362617]
ベトナムにおける様々な視覚的推論能力を評価するための先駆的な収集であるViCLEVRデータセットを紹介した。
我々は、現代の視覚的推論システムの包括的な分析を行い、その強みと限界についての貴重な洞察を提供する。
PhoVITは、質問に基づいて画像中のオブジェクトを識別する総合的なマルチモーダル融合である。
論文 参考訳(メタデータ) (2023-10-27T10:44:50Z) - OpenViVQA: Task, Dataset, and Multimodal Fusion Models for Visual
Question Answering in Vietnamese [2.7528170226206443]
ベトナム初の視覚的質問応答のための大規模データセットであるOpenViVQAデータセットを紹介する。
データセットは37,000以上の質問応答ペア(QA)に関連付けられた11,000以上の画像で構成されている。
提案手法は,SAAA,MCAN,LORA,M4CなどのSOTAモデルと競合する結果が得られる。
論文 参考訳(メタデータ) (2023-05-07T03:59:31Z) - Achieving Human Parity on Visual Question Answering [67.22500027651509]
The Visual Question Answering (VQA) task using both visual image and language analysis to answer a textual question to a image。
本稿では,人間がVQAで行ったのと同じような,あるいは少しでも良い結果が得られるAliceMind-MMUに関する最近の研究について述べる。
これは,(1)包括的視覚的・テキスト的特徴表現による事前学習,(2)参加する学習との効果的な相互モーダル相互作用,(3)複雑なVQAタスクのための専門的専門家モジュールを用いた新たな知識マイニングフレームワークを含む,VQAパイプラインを体系的に改善することで達成される。
論文 参考訳(メタデータ) (2021-11-17T04:25:11Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。