論文の概要: SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs
- arxiv url: http://arxiv.org/abs/2608.09006v1
- Date: Mon, 10 Aug 2026 01:47:30 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-11 19:16:37.034298
- Title: SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs
- Title(参考訳): SignLlama: LLMの視覚的特徴の優先順位付けによるグロスフリー手話翻訳の強化
- Authors: Shiwei Gan, Xiao Liu, Yafeng Yin, Zhiwei Jiang, Bowen Guo, Lie Xie, Sanglu Lu, Hongkai Wen,
- Abstract要約: 本稿では,Gloss-Free Sign Language TranslationタスクにLarge Language Modelsを適用するために解決しなければならない2つの重要な課題について述べる。
まず,Filted Pseudo-Gloss CTC Pretraining という簡単な手法を提案する。
第2に,テキスト入力がマスクされた視覚のみの予測パスを定義し,視覚入力のみに依存するターゲットシーケンスを生成するためにモデルが必要である。
- 参考スコア(独自算出の注目度): 24.378058252362646
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Translation (GFSLT) remains a challenge. In this paper, we investigate how to effectively adapt LLMs to the GFSLT task. We show that there are two key issues that need to be solved: (1) the inherent distributional gap between visual feature inputs and text feature inputs makes it difficult for LLMs to interpret visual inputs; and (2) existing approaches typically concatenate visual and textual features in an autoregressive framework, which leads to the model overemphasizing textual inputs and deprioritizing visual cues, as LLMs are pretrained predominantly on text-centric data. To address the first challenge, we propose a simple yet effective method named Filtered Pseudo-Gloss CTC Pretraining, which leverages filtered pseudo-gloss sequences generated from text sequences to supervise the training of the visual backbone. To tackle the second issue, we introduce a Visual-Prioritized Distillation training strategy. Specifically, we define a visual-only prediction path in which text inputs are masked, and the model is required to generate the target sequence relying solely on visual inputs. To guide this path, the outputs from the standard visual-textual prediction are then distilled into the visual-only prediction path, encouraging the model to prioritize visual features. Comprehensive experiments and qualitative analyses demonstrate the effectiveness of the proposed model. The proposed SignLlama achieves very competitive performance on multiple datasets for GFSLT tasks, without using any extra modalities or external sign language datasets for pretraining.
- Abstract(参考訳): 大規模言語モデル(LLM)は、幅広いタスクで大きな成功を収めています。
しかし、Gross-Free Sign Language Translation (GFSLT) のための微調整 LLM は依然として課題である。
本稿では, GFSLT タスクに LLM を効果的に適応させる方法について検討する。
1) 視覚的特徴入力とテキスト特徴入力の分布ギャップが視覚的入力の解釈を困難にすること,2) 既存のアプローチが視覚的特徴とテキスト的特徴を自己回帰的フレームワークに結合させることによって,テキスト入力を過度に強調し,視覚的手がかりを優先順位付けするモデルが,LLMが主にテキスト中心のデータに基づいて事前訓練されていること,である。
最初の課題に対処するために、テキストシーケンスから生成されたフィルタされた擬似グロスシーケンスを利用して視覚的バックボーンのトレーニングを監督する、Filted Pseudo-Gloss CTC Pretrainingというシンプルな方法を提案する。
第2の課題に対処するために、我々は、Visual-Prioritized Distillationトレーニング戦略を導入する。
具体的には、テキスト入力がマスクされた視覚のみの予測パスを定義し、そのモデルが視覚入力のみに依存するターゲットシーケンスを生成する必要がある。
この経路を導くために、標準的な視覚テキスト予測からの出力を視覚のみの予測経路に蒸留し、モデルに視覚的特徴の優先順位付けを促す。
総合的な実験と定性的分析により,提案モデルの有効性が示された。
提案したSignLlamaは、事前トレーニングに余分なモダリティや外部手話データセットを使わずに、GFSLTタスクの複数のデータセット上で非常に競合的なパフォーマンスを実現している。
関連論文リスト
- Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance [67.26434607115392]
大規模視覚言語モデル(LVLM)は様々な視覚言語タスクにおいて印象的な成果を上げている。
LVLMは言語バイアスによる幻覚に悩まされ、画像や非効果的な視覚的理解に焦点が当てられなくなった。
MDA (Multimodal duAl-attention meChanIsm) aNd soft-image Guidance (IFG) を用いたLVLMの言語バイアスに対処するためのLACingを提案する。
論文 参考訳(メタデータ) (2024-11-21T16:33:30Z) - VILA: On Pre-training for Visual Language Models [74.08039416548209]
ステップ・バイ・ステップ制御可能な比較によるVLM事前学習の設計オプションについて検討した。
私たちは、最先端のモデルよりも一貫して優れたVisual LanguageモデルファミリであるVILAを構築します。
論文 参考訳(メタデータ) (2023-12-12T18:58:18Z) - Visually-augmented pretrained language models for NLP tasks without
images [77.74849855049523]
既存のソリューションはしばしば視覚的知識増強のために明示的なイメージに依存している。
我々は、新しいtextbfVisually-textbfAugmented fine-tuningアプローチを提案する。
我々のアプローチは、BERT、RoBERTa、BART、T5を異なるスケールで継続的に改善することができる。
論文 参考訳(メタデータ) (2022-12-15T16:13:25Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。