Szczegóły publikacji
Opis bibliograficzny
UNCOM: zero-shot context-aware command understanding for tabletop scenarios / Antonio Galiza Cerdeira Gonzalez, Paweł GAJEWSKI, Bipin Indurkhya // W: 2026 IEEE International Conference on Advanced Robotics and its Social Impacts (ARSO) [Dokument elektroniczny] : 10-12 June 2026, Vienna, Austria : proceedings. — Wersja do Windows. — Dane tekstowe. — Piscataway : IEEE, 2026. — ( Conference proceedings (IEEE Workshop on Advanced Robotics and its Social Impacts) ; ISSN 2162-7568 ). — Dod. ISBN: 979-8-3315-6446-9 (print on demand). — e-ISBN: 979-8-3315-6445-2. — S. 163–168. — Wymagania systemowe: Adobe Reader. — Bibliogr. s. 168, Abstr. — P. Gajewski - dod. afiliacja: Jagiellonian University, Krakow, Poland
Autorzy (3)
- Gonzalez Antonio Galiza Cerdeira
- AGHGajewski Paweł
- Indurkhya Bipin
Dane bibliometryczne
| ID BaDAP | 168667 |
|---|---|
| Data dodania do BaDAP | 2026-07-23 |
| Tekst źródłowy | URL |
| DOI | 10.1109/ARSO68304.2026.11536136 |
| Rok publikacji | 2026 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Wydawca | Institute of Electrical and Electronics Engineers (IEEE) |
| Czasopismo/seria | Conference proceedings (IEEE Workshop on Advanced Robotics and its Social Impacts) |
Abstract
This paper presents UNCOM, a novel hybrid framework for interpreting natural human commands in tabletop scenarios. The system integrates multiple sources of information - speech, gestures, and scene context - to extract structured, actionable instructions for robots. Addressing the need for general-purpose human-robot interaction in domestic environments, UNCOM is designed for zero-shot operation, without reliance on predefined object models or training data specific to a given task. Using foundational and task-specific deep learning models, it allows out-of-the-box speech recognition, natural language understanding, gesture detection, and object segmentation. The modular architecture enhances transparency and explainability by explicitly parsing commands into object-action-target representations, enabling integration with symbolic robotic frameworks. We demonstrate the system in a TIAGo++ robot and provide an evaluation on a real-world data set of human-robot interaction scenarios; achieving an 82.39% success rate over our benchmark data set, highlighting the robustness of the system to diversity, noise, and communication ambiguity. The data set, evaluation scenarios, and the code are publicly available to support future research. © 2026 IEEE.