Sticky-Glance: Robust Intent Recognition for Human Robot Collaboration via Single-Glance
- URL: http://arxiv.org/abs/2603.06121v1
- Date: Fri, 06 Mar 2026 10:27:12 GMT
- Title: Sticky-Glance: Robust Intent Recognition for Human Robot Collaboration via Single-Glance
- Abstract summary: We propose an object-centric gaze grounding framework that stabilizes intent through a sticky-glance algorithm.<n>The inferred intent remains anchored to the object even under short glances with minimal 3 gaze samples.<n> Experiments across dynamic tracking, multi-perspective alignment, a baseline comparison, user studies, and ablation studies demonstrate improved robustness, efficiency, and reduced workload.
- Score: 24.0711921276237
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Gaze is a valuable means of communication for impaired people with extremely limited motor capabilities. However, robust gaze-based intent recognition in multi-object environments is challenging due to gaze noise, micro-saccades, viewpoint changes, and dynamic objects. To address this, we propose an object-centric gaze grounding framework that stabilizes intent through a sticky-glance algorithm, jointly modeling geometric distance and direction trends. The inferred intent remains anchored to the object even under short glances with minimal 3 gaze samples, achieving a tracking rate of 0.94 for dynamic targets and selection accuracy of 0.98 for static targets. We further introduce a continuous shared control and multi-modal interaction paradigm, enabling high-readiness control and human-in-loop feedback, thereby reducing task duration for nearly 10 \%. Experiments across dynamic tracking, multi-perspective alignment, a baseline comparison, user studies, and ablation studies demonstrate improved robustness, efficiency, and reduced workload compared to representative baselines.
Related papers
- Cooperative Informative Sensing for Monitoring Dynamic Indoor Environments via Multi-Agent Reinforcement Learning [56.64821510576244]
We formulate cooperative active observation as a decentralized control problem in which multiple robots adjust their motion to directly optimize monitoring accuracy under partial observability.<n>We propose a learning-based framework for cooperative policies from decentralized observations using multi-agent reinforcement learning (MARL), supported by an architecture that handles variable numbers of humans and temporal dependencies.
arXiv Detail & Related papers (2026-04-25T07:20:15Z) - Learn Weightlessness: Imitate Non-Self-Stabilizing Motions on Humanoid Robot [56.377161933570335]
Humans naturally exploit a "weightless" state during non-self-stabilizing (NSS) motions.<n>Inspired by this biological mechanism, we propose a weightlessness-state auto-labeling strategy for dataset annotation.<n>Our work bridges the gap between precise trajectory tracking and adaptive environmental interaction, offering a biologically-inspired solution for contact-rich humanoid control.
arXiv Detail & Related papers (2026-04-23T07:10:05Z) - Olfactory pursuit: catching a moving odor source in complex flows [0.5541644538483947]
Odor signals are intermittent, strongly mixed by turbulent-like transport, and typically lag behind the true target position.<n>We formulate olfactory pursuit as a partially observable Markov decision process in which an agent maintains a joint belief over the target's position and velocity.<n>Our results identify predictive inference of target motion as the key ingredient for effective olfactory pursuit.
arXiv Detail & Related papers (2026-04-13T15:52:24Z) - ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation [55.467742403416175]
We introduce a physics-driven neural algorithm that translates large-scale motion capture to humanoid embodiments.<n>We learn a unified multimodal controller that supports both dense references and sparse task specifications.<n>Results show that ULTRA generalizes to autonomous, goal-conditioned whole-body loco-manipulation from egocentric perception.
arXiv Detail & Related papers (2026-03-03T18:59:29Z) - HAIC: Humanoid Agile Object Interaction Control via Dynamics-Aware World Model [56.4392302336014]
We present HAIC, a framework for robust interaction across diverse object dynamics without external state estimation.<n>Our key contribution is a dynamics predictor that estimates high-order object states (velocity, acceleration) solely from proprioceptive history.<n> Experiments on a humanoid robot show HAIC achieves high success rates in agile tasks.
arXiv Detail & Related papers (2026-02-12T09:34:35Z) - A2VISR: An Active and Adaptive Ground-Aerial Localization System Using Visual Inertial and Single-Range Fusion [17.59898575353575]
It's a practical approach using the ground-aerial collaborative system to enhance the localization robustness of flying robots in cluttered environments.<n>We improve the ground-aerial localization framework in a more comprehensive manner, which integrates active vision, single-ranging, inertialometry, and optical flow.<n>The proposed approach achieves robust online localization, with an average root mean square error of approximately 0.09 m, while maintaining resilience to capture loss and sensor failures.
arXiv Detail & Related papers (2025-12-18T10:07:06Z) - Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking [16.366398265001422]
3D multi-object tracking is a critical and challenging task in the field of autonomous driving.<n>We introduce the Dynamic Scene Cue-Consistency Tracker (DSC-Track) to implement this principle.
arXiv Detail & Related papers (2025-08-15T08:48:13Z) - DINO-CoDT: Multi-class Collaborative Detection and Tracking with Vision Foundation Models [11.34839442803445]
We propose a multi-class collaborative detection and tracking framework tailored for diverse road users.<n>We first present a detector with a global spatial attention fusion (GSAF) module, enhancing multi-scale feature learning for objects of varying sizes.<n>Next, we introduce a tracklet RE-IDentification (REID) module that leverages visual semantics with a vision foundation model to effectively reduce ID SWitch (IDSW) errors.
arXiv Detail & Related papers (2025-06-09T02:49:10Z) - Open-World Drone Active Tracking with Goal-Centered Rewards [62.21394499788672]
Drone Visual Active Tracking aims to autonomously follow a target object by controlling the motion system based on visual observations.<n>We propose DAT, the first open-world drone active air-to-ground tracking benchmark.<n>We also propose GC-VAT, which aims to improve the performance of drone tracking targets in complex scenarios.
arXiv Detail & Related papers (2024-12-01T09:37:46Z) - Interpretable Dynamic Graph Neural Networks for Small Occluded Object Detection and Tracking [0.0]
This paper introduces DGNN-YOLO, a novel framework that integrates dynamic graph neural networks (DGNNs) with YOLO11 to address limitations.<n>Unlike standard GNNs, DGNNs are chosen for their superior ability to dynamically update graph structures in real-time.<n>This framework constructs and regularly updates its graph representations, capturing objects as nodes and their interactions as edges.
arXiv Detail & Related papers (2024-11-26T09:29:27Z) - MotionTrack: Learning Robust Short-term and Long-term Motions for
Multi-Object Tracking [56.92165669843006]
We propose MotionTrack, which learns robust short-term and long-term motions in a unified framework to associate trajectories from a short to long range.
For dense crowds, we design a novel Interaction Module to learn interaction-aware motions from short-term trajectories, which can estimate the complex movement of each target.
For extreme occlusions, we build a novel Refind Module to learn reliable long-term motions from the target's history trajectory, which can link the interrupted trajectory with its corresponding detection.
arXiv Detail & Related papers (2023-03-18T12:38:33Z) - TRiPOD: Human Trajectory and Pose Dynamics Forecasting in the Wild [77.59069361196404]
TRiPOD is a novel method for predicting body dynamics based on graph attentional networks.
To incorporate a real-world challenge, we learn an indicator representing whether an estimated body joint is visible/invisible at each frame.
Our evaluation shows that TRiPOD outperforms all prior work and state-of-the-art specifically designed for each of the trajectory and pose forecasting tasks.
arXiv Detail & Related papers (2021-04-08T20:01:00Z) - Learning Global Structure Consistency for Robust Object Tracking [57.736915865309165]
This work considers the emphtransient variations of the whole scene.
We propose an effective and efficient short-term model that learns to exploit the global structure consistency in a short time.
We empirically verify that the proposed tracker can tackle the two challenging scenarios and validate it on large scale benchmarks.
arXiv Detail & Related papers (2020-08-26T19:12:53Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.