論文の概要: NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation
- arxiv url: http://arxiv.org/abs/2609.36756v3
- Date: Thu, 01 Oct 2026 04:49:21 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:23.51939
- Title: NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation
- Title(参考訳): NesTok: 自己回帰画像生成のためのNested Self-Aligned 1D Tokenizer
- Abstract要約: NesTokは、動的ビジュアルトークン化ツールに適したネストされた自己アライメントフレームワークである。
標準トレーニングよりも大幅に改善され、ImageNetで0.98のrFIDスコアを達成している。
- 参考スコア(独自算出の注目度): 53.369637657900235
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256$\times$256 among existing variable-length autoregressive image generation methods. Code will be available at https://github.com/Jiawei804/NesTok.
- Abstract(参考訳): 1次元 (1D) 可変長の視覚的トークン化器は、トークンの数を変えることで適応的な圧縮を可能にし、下流の自己回帰(AR)モデルでは、単一のトークン化器を用いて生成品質を計算コストに対して柔軟にトレードオフすることができる。
しかし、ネストしたドロップアウトに基づく既存のアプローチは、しばしばトークン化器の表現能力を完全に活用することができず、画像再構成と生成の両方において最適以下の性能をもたらす。
本研究では,動的視覚的トークン化に適したネスト型自己アライメントフレームワークNesTokを紹介する。
NesTokはクロス長トレーニングを導入し、フル長のシーケンスを使用してトークンの長さ間の再構築を共同で最適化し、短いトークンシーケンスがフル長のシーケンスの再構築品質に近づくことを可能にする。
ImageNetでは、NesTokは標準トレーニングよりも大幅に改善されており、rFIDスコアは0.98である。
下流画像生成では、既存の可変長自己回帰画像生成手法のうち、ImageNet 256$\times$256で1.46の最先端のgFIDスコアを得る。
コードはhttps://github.com/Jiawei804/NesTok.comで入手できる。
関連論文リスト
- Improving Flexible Image Tokenizers for Autoregressive Image Generation [53.238708824055664]
textbfReToKは、アンダーライン冗長なアンダーラインToken Paddingと階層的セマンティック正規化を備えたフレキシブルなトークンライザである。
本手法は, フレキシブルかつ固定長のトークン化器と比較して, 優れた生成性能を実現する。
論文 参考訳(メタデータ) (2026-01-04T14:11:45Z) - Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models [92.18057318458528]
Token-ShuffleはTransformerにおける画像トークンの数を減らす新しい方法である。
我々の戦略は、事前訓練されたテキストエンコーダを必要とせず、MLLMが超高解像度画像合成をサポートできるようにする。
GenAIベンチマークでは、2.7Bモデルがハードプロンプトで0.77点、ARモデルLlamaGenが0.18点、拡散モデルLDMが0.15点である。
論文 参考訳(メタデータ) (2025-04-24T17:59:56Z) - FlexTok: Resampling Images into 1D Token Sequences of Flexible Length [16.76602756308683]
可変長の1Dトークンシーケンスに2D画像を投影するトークンライザであるFlexTokを紹介する。
簡単なGPT型変換器を用いて, 自己回帰生成設定によるアプローチの評価を行った。
論文 参考訳(メタデータ) (2025-02-19T18:59:44Z) - ImageFolder: Autoregressive Image Generation with Folded Tokens [51.815319504939396]
トークン長の増大は、画像再構成の品質を改善するための一般的なアプローチである。
トークン長に関する復元と生成品質の間にはトレードオフがある。
本稿では,自己回帰モデルにおいて折り畳み可能な空間整列型画像トークンを提供するセマンティック・トークンライザであるイメージを提案する。
論文 参考訳(メタデータ) (2024-10-02T17:06:39Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。