論文の概要: Omni Interaction Agent Technical Report
- arxiv url: http://arxiv.org/abs/2609.08977v2
- Date: Wed, 09 Sep 2026 09:47:31 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-10 19:44:08.694432
- Title: Omni Interaction Agent Technical Report
- Title(参考訳): オムニ・インタラクション・エージェント技術報告
- Authors: Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddy Sun, Steve Yves, Zhou Zhao,
- Abstract要約: Ganderは、Omni omniリアルタイムインタラクションとエージェント機能を単一のフレームワークに統合する、エンドツーエンドモデルである。
Ganderは、ビデオ、音声、テキストを含む複数のモードで入力を受け取り、完全な自然な対話を可能にする。
我々は、会話能力、全能理解、対話能力、エージェントインテリジェンスという4つの側面にわたって、ガンダーの包括的評価を行う。
- 参考スコア(独自算出の注目度): 48.78411737142625
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives streaming inputs across multiple modalities, including video, speech, and text, enabling natural full-duplex interaction in both everyday conversations and complex workflow-oriented agent scenarios. Users can interrupt the model at any time, while the model can also proactively provide intermediate feedback or ask follow up questions. To natively support these capabilities, Gander adopts two key architectural designs: 1) It employs a Cerebellum-Brain collaborative framework, in which the Cerebellum is responsible for realtime interaction and omni conversational capabilities, while the Brain handles complex reasoning and higher-level agentic tasks. The two components interact continuously through tool calling and the agent orchestration runtime. 2) The Cerebellum is built upon a streaming Thinker-Talker architecture, user inputs and model outputs are further flattened into an ordered token stream at the chunk level, providing a unified representation for low latency, continuous interaction. We conduct comprehensive evaluations of Gander across four dimensions: conversational ability, omni understanding, interactive capability, and agentic intelligence. Internal human evaluations demonstrate that Gander maintains the natural and expressive spoken dialogue capabilities of SOTA open source models while achieving competitive performance in omni interaction. Gander also demonstrates robustness in challenging real-world scenarios, including background noise interference, multi-party interactions, and backchannel communication. We release Gander together with its models, code, and data to facilitate further research and development in the community.
- Abstract(参考訳): 本稿では,単一フレームワーク内でのオムニ認識,リアルタイムインタラクション,エージェント機能を統一するエンドツーエンドモデルであるGanderを紹介する。
ターンベースの従来のパラダイムとは対照的に、Ganderはビデオ、音声、テキストを含む複数のモードにわたるストリーミング入力を継続的に受信し、日々の会話と複雑なワークフロー指向のエージェントシナリオの両方で自然なフル二重のインタラクションを可能にする。
ユーザーはいつでもモデルを中断することができるが、モデルは積極的に中間的なフィードバックを提供したり、フォローアップ質問をすることができる。
これらの機能をネイティブにサポートするために、Gander氏は2つの重要なアーキテクチャ設計を採用した。
1)脳と脳の協調的な枠組みを用いており、脳は複雑な推論と高レベルのエージェント的タスクを処理し、脳はリアルタイムの相互作用と全会話能力に責任を負う。
この2つのコンポーネントは、ツール呼び出しとエージェントオーケストレーションランタイムを通じて継続的に相互作用する。
2) CerebellumはストリーミングのThinker-Talkerアーキテクチャ上に構築されており,ユーザ入力とモデル出力はさらに,チャンクレベルで順序付けられたトークンストリームにフラット化され,低レイテンシ,継続的なインタラクションの統一表現を提供する。
我々は、会話能力、全能理解、対話能力、エージェントインテリジェンスという4つの側面にわたって、ガンダーの包括的評価を行う。
内部の人間による評価は、ガンダーがOTAオープンソースモデルの自然かつ表現力のある対話能力を維持しつつ、オムニ相互作用における競合性能を実現していることを示している。
Gander氏はまた、バックグラウンドノイズ干渉、マルチパーティインタラクション、バックチャネル通信など、現実のシナリオに挑戦する上で、堅牢性を示す。
私たちはGanderをそのモデル、コード、データとともにリリースし、コミュニティにおけるさらなる研究と開発を促進します。
関連論文リスト
- MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models [62.05118198431989]
非同期のフル音声モデルは、AI停止のフルタイムの対話性と自然な性質によって区別される。
本フレームワークは,外部情報における知識要求型対話クエリと接地応答の同定を可能にする。
本設計では,再学習を伴わないプラグ・アンド・プレイ検索手法をサポートし,アウト・オブ・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー・ツー
論文 参考訳(メタデータ) (2026-04-14T16:17:52Z) - Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion Models [80.28579390566298]
テキスト条件付き自己回帰拡散モデルであるInteract2Arを導入する。
ハンドキネマティクスは専用のパラレルブランチを通じて組み込まれ、高忠実度フルボディ生成を可能にする。
我々のモデルは、時間的動きの合成、外乱へのリアルタイム適応、ディヤディックからマルチパーソンシナリオへの拡張など、一連のダウンストリームアプリケーションを可能にする。
論文 参考訳(メタデータ) (2025-12-22T18:59:50Z) - End-to-end Listen, Look, Speak and Act [22.047534228540783]
ELLSAは、より自然で一般的な対話型人工知能への一歩であり、人工知能の幅広い追求に寄与している。
中心となるのはSA-MoE(Attention Mixture-of-Experts)で、それぞれのモダリティを専門の専門家にルーティングすることで、統一された注意バックボーンを通じてそれらを融合させる。
論文 参考訳(メタデータ) (2025-10-19T08:45:46Z) - Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset [113.25650486482762]
4000時間以上の対面インタラクション映像の大規模な収集であるSeamless Interactionデータセットを紹介した。
このデータセットは、ダイドの具体的ダイナミクスを理解するAIテクノロジの開発を可能にする。
そこで我々は,このデータセットを用いて,人間の発話に適応した動作ジェスチャーと表情を生成するモデル群を開発した。
論文 参考訳(メタデータ) (2025-06-27T18:09:49Z) - Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication [4.49451692966442]
本稿では,効果的なコミュニケーションのための話者と聞き手の拡散間生成モデルを提案する。
初めて、リスナーのフルボディジェスチャーを生成フレームワークに統合する。
論文 参考訳(メタデータ) (2025-05-08T07:00:58Z) - It Takes Two: Real-time Co-Speech Two-person's Interaction Generation via Reactive Auto-regressive Diffusion Model [34.94330722832987]
会話中の2文字の動的動きを合成するための音声駆動自動回帰システムを提案する。
我々の知る限りでは、オンライン方式で2文字の対話型フルボディモーションを生成できる最初のシステムである。
論文 参考訳(メタデータ) (2024-12-03T12:31:44Z) - A Probabilistic Model Of Interaction Dynamics for Dyadic Face-to-Face
Settings [1.9544213396776275]
我々は,対面設定における対の参加者間の相互作用のダイナミクスを捉える確率論的モデルを開発した。
この相互作用エンコーディングは、あるエージェントの将来のダイナミクスを予測する際に、生成に影響を与えるために使用される。
我々のモデルは, 相互作用する力学に基づいて, モード間のデライン化に成功していることを示す。
論文 参考訳(メタデータ) (2022-07-10T23:31:27Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。