論文の概要: IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence
- arxiv url: http://arxiv.org/abs/2609.36860v1
- Date: Tue, 29 Sep 2026 07:03:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-30 21:28:47.254529
- Title: IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence
- Title(参考訳): IronLLM: リアルタイム・エンボディード・インテリジェンスのためのコンパクトエッジネイティブ言語モデルの作成
- Abstract要約: デバイス上での効率的な推論のための言語モデルであるIronLLM-0.6Bを提案する。
モデルは品質指向のデータパイプラインを使用して、約6.2兆のトークンで事前訓練される。
また、RMSNormをDynamic Tanhに置き換えたIronLLM-0.6B-Lightも紹介する。
- 参考スコア(独自算出の注目度): 20.839888550150153
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-free drafting, achieving a 1.48x decoding speedup. The model is pretrained on approximately 6.2 trillion tokens using a quality-oriented data pipeline and is further post-trained with Multi-Domain On-Policy Distillation to integrate capabilities from domain-specialized teachers. To better meet the low-latency requirements of on-device scenarios, IronLLM-0.6B adopts an Instruct-Only design. Evaluations show that IronLLM-0.6B achieves competitive performance relative to larger models such as Qwen3.5-0.8B and MiniCPM5-1B, while producing more concise responses on many tasks. We further present IronLLM-0.6B-Light, which replaces RMSNorm with Dynamic Tanh and simplifies several computationally expensive components to improve inference and quantization efficiency. Together, the IronLLM models provide an effective performance-efficiency trade-off for resource-constrained deployment.
- Abstract(参考訳): 654Mパラメータ言語モデルであるIronLLM-0.6Bを提案する。
IronLLM-0.6Bは、ハイブリッドアテンションアーキテクチャとX-MTPを組み合わせる。これは、KVキャッシュ毎の詳細なリプレイを排除し、ロールバックフリーのドラフトに軽量な検証ヘッドを使用し、1.48倍のデコードスピードアップを達成する軽量な共有KVマルチトークン予測設計である。
このモデルは、品質指向のデータパイプラインを使用して約6.2兆のトークンで事前訓練され、ドメイン特化教師の能力を統合するために、マルチドメイン・オン・ポリティ蒸留(Multi-Domain On-Policy Distillation)でさらに訓練されている。
オンデバイスシナリオの低レイテンシ要件を満たすため、IronLLM-0.6Bはインストラクトオンリー設計を採用する。
評価の結果、IronLLM-0.6BはQwen3.5-0.8BやMiniCPM5-1Bのような大型モデルと比較して競争性能が向上し、多くのタスクに対してより簡潔な応答が得られた。
さらに, RMSNorm を Dynamic Tanh に置き換えた IronLLM-0.6B-Light について述べる。
同時に、IronLLMモデルは、リソース制約されたデプロイメントに対して、効果的なパフォーマンス効率トレードオフを提供する。
関連論文リスト
- Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models [2.1665689529884697]
本稿では,コンパクトな埋め込みモデルと汎用SLMの2つのコンポーネントからなる効率的なRAGアーキテクチャであるB1adeについて述べる。
5つの事前訓練エンコーダのパラメータフリー融合により構築された335Mパラメータ検索モデルであるB1ade-embedは、追加トレーニングをゼロとした500M未満モデルの上位MTEBスコアを達成する。
B1ade-1Bは42.4%の反応で回収された通路を引用し、トレーニング分布の寄与率を5.5ポイント上回っている。
論文 参考訳(メタデータ) (2026-07-29T22:41:25Z) - MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling [80.48332380100915]
MiniCPM-SALAは、疎注意の高忠実長文モデリングと線形注意のグローバル効率を統合するハイブリッドモデルである。
1つのNVIDIA A6000D GPUでは、256Kトークンのシーケンス長におけるフルアテンションモデルの推論速度が3.5倍に達する。
論文 参考訳(メタデータ) (2026-02-12T09:37:05Z) - MC#: Mixture Compressor for Mixture-of-Experts Large Models [86.64315380917827]
Mixture-of-Experts (MoE)は、大きな言語モデル(LLM)と視覚言語モデル(VLM)をスパースアクティベーションによって拡張することで効果的にスケールする。
静的量子化と動的エキスパートプルーニングを組み合わせたフレームワークであるMC#(Mixture-Compressor-sharp)を提案する。
論文 参考訳(メタデータ) (2025-10-13T03:12:46Z) - Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models [49.911784762244814]
TraceRLは拡散言語モデル(DLM)のための軌道対応強化学習フレームワークである
我々は最先端の拡散言語モデル、すなわち TraDo を導出する。
TraDo-8B-InstructはQwen2.5-7B-Instructで6.1%、Llama3.1-8B-Instructで51.3%の精度向上を実現している。
論文 参考訳(メタデータ) (2025-09-08T17:58:06Z) - MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization [74.04867639197445]
MiroMind-M1 は Qwen-2.5 ベースのベンチマーク上に構築された完全なオープンソース RLM のセットである。
我々のモデルは2つの段階で訓練されている: SFT on a carefully curated corpus of 719K math-reasoning problem with confirmed CoT trajectories, then RLVR on 62K challenge and verible problem。
論文 参考訳(メタデータ) (2025-07-19T16:21:23Z) - FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing [59.12511498024836]
本稿では,重み付けスコアに基づいてモデルブロックを選択的にプルーする大規模言語モデル(LLM)をプルーする手法を提案する。
重み共有機構を用いて各刈り込みブロックを置換する原理的計量を提案する。
経験的評価は、既存の方法よりも大幅にパフォーマンスが向上したことを示している。
論文 参考訳(メタデータ) (2025-01-24T18:46:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。