論文の概要: From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping
- arxiv url: http://arxiv.org/abs/2607.00530v1
- Date: Wed, 01 Jul 2026 07:19:00 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-02 19:56:07.773876
- Title: From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping
- Title(参考訳): 技術メトリクスからユーザ知覚へ:オブジェクト検出とグラフピングのためのマルチモーダルヒューマンロボットインタラクションシステムのユーザスタディ
- Authors: Jian Song, Tian Zi, Shen Guanting,
- Abstract要約: 本稿では、エンド・ツー・エンドのタスク成功における15パーセントの利得が、ユーザ知覚の一貫性と測定可能な相違をもたらすのに十分かどうかを検討する。
ベースラインシステムは、音声認識用Whisper、オープン語彙オブジェクト検出用Florence-2、アクション抽出用LLaMA 3.1、動作実行用インターバル型2ファジィロジックコントローラを組み合わせる。
- 参考スコア(独自算出の注目度): 11.17111944571978
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Improvements in the technical performance of human--robot interaction (HRI) systems do not automatically translate into differences that human users can detect during live interaction. This paper investigates whether a 15 percentage point gain in end-to-end task success (from 75% in a multimodal baseline system to 90% in an improved configuration identified through a prior ablation study) is sufficient to produce consistent and measurable differences in user perception. The baseline system combines Whisper for speech recognition, Florence-2 for open-vocabulary object detection, LLaMA 3.1 for action extraction, and an interval Type-2 fuzzy logic controller for motion execution. The improved configuration replaces the perception and language modules with Grounding DINO + SAM and Qwen 3.5 9B, respectively, while retaining the same controller. A within-subject user study with 24 participants compared both systems on the same tabletop object-grasping task. After interacting with each configuration, participants rated perceived speed, reliability, and overall competence and fluency on a 7-point Likert scale. Results show that 17 out of 24 participants (70.83%) preferred the improved system (exact binomial test, p = 0.043, h = 0.43), and all three perceptual constructs were rated significantly higher for the improved configuration after Holm correction, with large to very large effect sizes (p < 0.001). These findings confirm that the identified technical improvements are perceptible to users in direct interaction and underscore the importance of complementing benchmark evaluation with user-centred evidence when assessing robotic manipulation pipelines.
- Abstract(参考訳): 人間-ロボットインタラクション(HRI)システムの技術的パフォーマンスの改善は、人間がライブインタラクション中に検出できる違いに自動的に変換されない。
本稿では、エンド・ツー・エンドのタスク成功率の15パーセント(マルチモーダルベースラインシステムでは75%から90%)が、ユーザ認知の一貫性と測定可能な相違をもたらすのに十分かどうかを検討する。
ベースラインシステムは、音声認識用Whisper、オープン語彙オブジェクト検出用Florence-2、アクション抽出用LLaMA 3.1、動作実行用インターバル型2ファジィロジックコントローラを組み合わせる。
改良された構成は、認識モジュールと言語モジュールをそれぞれ、同じコントローラを保持しながら、Grounding DINO + SAMとQwen 3.5 9Bで置き換える。
被験者24名を対象に、同じテーブルトップ・オブジェクト・グラスピング・タスクにおいて、両方のシステムを比較した。
各構成と対話した後、参加者は7ポイントのLikertスケールで、認識されたスピード、信頼性、全体的な能力と流布度を評価した。
その結果,24名中17名(70.83%)が改善系を好んだ(p = 0.043,h = 0.43)。
これらの結果から,ロボット操作パイプラインの評価において,これらの技術改善がユーザにとって直接的インタラクションにおいて認識可能であることが確認され,ベンチマーク評価とユーザ中心のエビデンスを補完することの重要性が強調された。
関連論文リスト
- An Interpretable Closed-Loop Intelligent Tutoring System for Multimodal Affective Feedback in Asynchronous Presentation Training [0.0]
ITSはマルチモーダル入力をエビデンスベースのフィードバックにマッピングし、観測可能なパフォーマンスキューに遡ることができる。
このシステムは、専門家のレーティングに匹敵するパフォーマンスレベルでルーリック整合スコアを達成した。
論文 参考訳(メタデータ) (2026-05-17T14:12:40Z) - Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task [0.0]
本論文は,従来のマルチモーダルな人間-ロボットインタラクションシステムを拡張し,エンドツーエンドのパフォーマンスに最も強い影響を与える3つのモジュールについて,制御されたアブレーション研究を導入する。
目標は、パイプライン全体を再設計するのではなく、共通の実験プロトコルの下で各コンポーネントのコントリビューションを分離し、エンドツーエンドで最高の組み合わせを評価することだ。
論文 参考訳(メタデータ) (2026-05-01T15:04:53Z) - Predicting the Intention to Interact with a Service Robot:the Role of Gaze Cues [51.58558750517068]
サービスロボットは、接近する人が対話する意図をできるだけ早く知覚する必要がある。
我々は,この認識課題を,対話を意図した潜在的なユーザ意図のシーケンス・ツー・シーケンス分類器を用いて解決する。
我々の主な貢献は、この文脈における人の視線を表す特徴の利点の研究である。
論文 参考訳(メタデータ) (2024-04-02T14:22:54Z) - MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition [62.89464258519723]
異なるレベルのオーディオ/視覚エンコーダに融合することで、各モードの表現を促進する多層クロスアテンション融合に基づくAVSR手法を提案する。
提案手法は第1位システムを超え,新たなSOTA cpCERの29.13%をこのデータセット上に構築する。
論文 参考訳(メタデータ) (2024-01-07T08:59:32Z) - ERNIE-SPARSE: Learning Hierarchical Efficient Transformer Through
Regularized Self-Attention [48.697458429460184]
情報ボトルネック感度と異なる注目トポロジ間の不整合の2つの要因がスパース変換器の性能に影響を及ぼす可能性がある。
本稿では,ERNIE-Sparseというモデルを提案する。
i) 局所情報とグローバル情報を逐次統一する階層スパース変換器(HST) と、(ii) 注意トポロジの異なる変換器の距離を最小化する自己注意正規化(SAR) の2つの特徴がある。
論文 参考訳(メタデータ) (2022-03-23T08:47:01Z) - Effects of Word-frequency based Pre- and Post- Processings for Audio
Captioning [49.41766997393417]
音響シーン・イベントの検出・分類のタスク6(自動音声キャプション)に使用したシステム(DCASE)2020 Challengeは,音声キャプションのためのデータ拡張,マルチタスク学習,ポストプロセッシングという3つの要素を組み合わせる。
このシステムは評価スコアが最も高いが、個々の要素のどれがパーフォーマンスに最も貢献したかはまだ明らかになっていない。
論文 参考訳(メタデータ) (2020-09-24T01:07:33Z) - Soliciting Human-in-the-Loop User Feedback for Interactive Machine
Learning Reduces User Trust and Impressions of Model Accuracy [8.11839312231511]
混合開始システムにより、ユーザは対話的にフィードバックを提供し、システムパフォーマンスを向上させることができる。
本研究は,フィードバックの提供行為が知的システムのユーザ理解とその正確性に与える影響について検討する。
論文 参考訳(メタデータ) (2020-08-28T16:46:41Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。