論文の概要: Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference
- arxiv url: http://arxiv.org/abs/2608.04428v1
- Date: Wed, 05 Aug 2026 04:17:03 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:43.719588
- Title: Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference
- Title(参考訳): Deltoris氏:bit-level SparsityとSpeculative Inferenceを通じて、Embodied AIでリアルタイムVLA推論を実現する
- Authors: Zheng Liu, Zeyu Guo, Zihan Liu, Anbang Wu, Han Zhao, Fangxin Liu, Zhezhi He, Yinhe Han, Jingwen Leng, Minyi Guo, Yiming Gan, Yu Feng,
- Abstract要約: 視覚アクション(VLA)モデルは、具体化されたAIの重要なコンポーネントとして登場した。
本稿では,効率的な拡散型VLA推論のためのアルゴリズムハードウェア協調設計フレームワークであるDeltorisを提案する。
- 参考スコア(独自算出の注目度): 40.7778980764007
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it imposes strict latency and energy constraints on edge devices. In this work, we present Deltoris, an algorithm-hardware co-design framework for efficient diffusion-based VLA inference. First, we exploit the temporal similarity of consecutive inputs and propose a \textit{temporal-aware bit-sparsity} algorithm that computes only the differences between consecutive inputs, eliminating redundant bit-level operations. To further address the extra off-chip traffic introduced by our algorithm, we propose a \textit{speculative inference} technique, which amortizes data loading across multiple control steps. Lastly, to support these techniques, we co-design a dedicated accelerator with customized 1D systolic bit-serial PE arrays that eliminate PE workload imbalance. Our evaluation shows that Deltoris achieves up to 34.2$\times$ speedup over mobile GPUs and 6.1$\times$ over prior accelerators, while maintaining comparable accuracy.
- Abstract(参考訳): 視覚言語アクション(VLA)モデルは、組み込みAIの重要なコンポーネントとして登場した。
既存のアプローチの中で、拡散に基づくVLAモデルはより優れた運動品質と一般化を実現する。
しかし、拡散ベースのVLAモデルは計算集約的であり、例えば50-200Hzの高制御周波数で実行する必要がある。
したがって、エッジデバイスに厳格なレイテンシとエネルギー制約を課す。
本稿では,効率的な拡散に基づくVLA推論のためのアルゴリズムハードウェア協調設計フレームワークであるDeltorisを紹介する。
まず、連続入力の時間的類似性を利用して、連続入力間の差のみを計算し、冗長なビットレベル演算を除去する「textit{temporal-aware bit-sparsity}」アルゴリズムを提案する。
提案アルゴリズムは,複数の制御ステップにまたがるデータローディングを復号化するための<textit{speculative inference} 手法を提案する。
最後に、これらの技術をサポートするために、PEの負荷不均衡を解消する1次元シストリックなビットシリアルPEアレイをカスタマイズした専用加速器を設計する。
我々の評価によると、DeltorisはモバイルGPUよりも34.2$\times$スピードアップし、以前のアクセラレータよりも6.1$\times$を達成でき、精度は同等である。
関連論文リスト
- NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows [75.70583906344815]
拡散モデルは、複雑なマルチモーダルな動作分布をモデル化できるため、アクションデコーダとして広く採用されている。
我々は、Vision-Language-Action(VLA)モデルのための拡散型デコーダの高速かつ表現性の高い代替品であるNinAを提案する。
論文 参考訳(メタデータ) (2025-08-23T00:02:15Z) - SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration [70.72227437717467]
VLA(Vision-Language-Action)モデルは、その強力な制御能力に注目が集まっている。
計算コストが高く、実行頻度も低いため、ロボット操作や自律ナビゲーションといったリアルタイムタスクには適さない。
本稿では,共同スケジューリングモデルとプルーニングトークンにより,VLAモデルを高速化する統一フレームワークSP-VLAを提案する。
論文 参考訳(メタデータ) (2025-06-15T05:04:17Z) - MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices [24.1144641404561]
本稿では,メモリ制約付きエッジアクセラレータ上での正確なアテンション推定高速化手法を提案する。
エッジコンピューティングのシナリオではFLAT (State-of-the-art attention fusion Method) と比較して,2.75倍のスピードアップと54%のエネルギー消費削減が見られた。
論文 参考訳(メタデータ) (2024-11-20T19:44:26Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。