ORCID

Abstract

Effective communication is vital in healthcare, especially across language barriers, where non-verbal cues and gestures are critical. This paper presents a privacy-preserving vision-language framework for medical interpreter robots that detects specific speech acts (consent and instruction) and generates corresponding robotic gestures. Built on locally deployed open-source models, the system utilizes a Large Language Model (LLM) with few-shot prompting for intent detection. We also introduce a novel dataset of clinical conversations annotated for speech acts and paired with gesture clips. Our identification module achieved 0.90 accuracy, 0.93 weighted precision, and a 0.91 weighted F1-Score. Our approach significantly improves computational efficiency and, in user studies, outperforms the speech-gesture generation baseline in human-likeness while maintaining comparable appropriateness.

Keywords

Gesture, Healthcare, Human-Robot Interaction, Large Language Model, Medical Interpreter, Pose Estimation

Publication Date

2026-03-16

Event

21st ACM/IEEE International Conference on Human-Robot Interaction, HRI Companion 2026

Publication Title

Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction, HRI Companion 2026

Publisher

Association for Computing Machinery (ACM)

ISBN

9798400723216

First Page

74

Last Page

79

Deposit Date

2026-08-13

Funding

This research was conducted with the financial support of Research Ireland under Grant Agreement No. 13/RC/2106_P2 at ADAPT, the SFI Research Centre for AI-Driven Digital Content Technology.

Creative Commons License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.


Share

COinS