論文の概要: SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
- arxiv url: http://arxiv.org/abs/2610.00686v1
- Date: Wed, 30 Sep 2026 20:29:37 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:23.758803
- Title: SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
- Title(参考訳): SemanTok:効率的な自動回帰ビデオ生成のための予測可能なセマンティックトークン
- Abstract要約: 我々は、凍結したDINO機能をエンコーダに供給するフレキシブルビデオトークンであるSemanTokを紹介した。
SemanTokは、すべてのARモデルサイズにおいて、高いセマンティックアライメントとビデオ忠実性を実現する。
- 参考スコア(独自算出の注目度): 8.581749377608215
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model $3.4\times$ its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.
- Abstract(参考訳): 最近のビデオベースの世界モデルは、自己回帰(AR)予測のスケーラビリティと拡散モデルの視覚的品質を組み合わせている。
シーントークン化器の選択は、忠実さとセマンティクスの両方の観点から、それぞれの最適なパフォーマンスにおいて最重要である。
最初の粗いトークンはクリップのグローバルなセマンティクスを持ち、後のトークンは詳細を指定します。
既存のフレキシブルなトークンライザは、初期デコーダ隠蔽状態に表現調整(REPA)損失のみを適用し、デコーダのターゲットはそのノイズ入力から部分的に満たされる。
SemanTokは、凍結したDINO機能をエンコーダに供給するフレキシブルなビデオトークンライザで、保持されているトークンプレフィックスからのみ再構成する軽量なヘッドを追加します。
SemanTokは、すべてのARモデルサイズにおいて高いセマンティックアライメントとビデオ忠実性を実現している。201MのSemanTok ARモデルは、VideoFlexTok ARモデルにマッチまたは打ち勝つ。
配布外のクラスにセマンティックアライメントを保持し、純粋なノイズを含むあらゆるノイズレベルでデコーダにセマンティックアライメントを与える。
復元と生成の両方で良好に機能し、短いトークンプレフィックスは、後続のトークンに遅延したピクセルディテールにより、予測と生成精度の向上のために安価である。
関連論文リスト
- EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation [80.13014959623452]
EVATokは、$textbfE$fficient $textbfV$ideo $textbfA$daptive $textbfTok$enizersを生成するフレームワークである。
EVATok は UCF-101 上でより優れた再構成と最先端のクラス・ビデオ生成を実現する。
論文 参考訳(メタデータ) (2026-03-12T17:59:59Z) - Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation [34.112157859384645]
自己回帰(AR)モデリングは、最先端の言語と視覚的生成モデルを支える。
伝統的に、トークン'' は最小の予測単位として扱われ、しばしば言語における離散的なシンボルまたは視覚における量子化されたパッチとして扱われる。
トークンの概念をエンティティXに拡張するフレームワークであるxARを提案する。
論文 参考訳(メタデータ) (2025-02-27T18:59:08Z) - WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling [63.8735398698683]
言語モデルの重要な構成要素は、高次元の自然信号を低次元の離散トークンに圧縮するトークン化器である。
本稿では,従来の音響領域におけるSOTA音響モデルよりもいくつかの利点があるWavTokenizerを紹介する。
WavTokenizerは、優れたUTMOSスコアを持つ最先端の再構築品質を実現し、本質的によりリッチなセマンティック情報を含んでいる。
論文 参考訳(メタデータ) (2024-08-29T13:43:36Z) - Fast End-to-End Speech Recognition via a Non-Autoregressive Model and
Cross-Modal Knowledge Transferring from BERT [72.93855288283059]
LASO (Listen Attentively, and Spell Once) と呼ばれる非自動回帰音声認識モデルを提案する。
モデルは、エンコーダ、デコーダ、および位置依存集合体(PDS)からなる。
論文 参考訳(メタデータ) (2021-02-15T15:18:59Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。