In this paper, we investigate ergonomic and technical factors limiting the effectiveness of gestural interfaces for Extended Reality based on hand pose recognition. For this purpose, we recorded and analyzed a novel dataset including hand poses known to be associated with semantics, performed by heterogeneous subjects wearing a popular Virtual Reality headset (Meta Quest 3) with finger-tracking capabilities. Unlike most related literature, which focuses on testing classifiers for recognizing static poses or dynamic gestures, we used the collected data to analyze these factors. We analyzed the variability of gesture execution and, through human visual labeling of renderings of the captured skeleton, verified the correspondence between the intended gesture labels and the recorded finger articulations. We investigated the effects of hand position, rotation, and skin color on the quality of the recorded poses. We evaluated the discriminability of the executions in feature spaces describing the hand poses, considering both the intended gesture labels and the visually assigned ones. We also collected information on semantic associations, familiarity with the gestures, and ergonomic issues through specific questionnaires. The outcomes of our study show that many poses commonly used to control interfaces and strongly associated with semantic priors are not optimal, as they are not always well executed or robustly tracked. The gesture labels corresponding to participants’ intentions do not always match those assigned during visual annotation, and poses are inadvertently modified during hand rotation. Our results provide valuable guidelines for designing effective gesture-controlled XR interfaces, including optimal dictionaries and system design choices.
Meaningful and ergonomic hand poses for extended reality interaction
Marco Emporio;Hassan Fattahi;Ariel Caputo;Andrea Giachetti
2026-01-01
Abstract
In this paper, we investigate ergonomic and technical factors limiting the effectiveness of gestural interfaces for Extended Reality based on hand pose recognition. For this purpose, we recorded and analyzed a novel dataset including hand poses known to be associated with semantics, performed by heterogeneous subjects wearing a popular Virtual Reality headset (Meta Quest 3) with finger-tracking capabilities. Unlike most related literature, which focuses on testing classifiers for recognizing static poses or dynamic gestures, we used the collected data to analyze these factors. We analyzed the variability of gesture execution and, through human visual labeling of renderings of the captured skeleton, verified the correspondence between the intended gesture labels and the recorded finger articulations. We investigated the effects of hand position, rotation, and skin color on the quality of the recorded poses. We evaluated the discriminability of the executions in feature spaces describing the hand poses, considering both the intended gesture labels and the visually assigned ones. We also collected information on semantic associations, familiarity with the gestures, and ergonomic issues through specific questionnaires. The outcomes of our study show that many poses commonly used to control interfaces and strongly associated with semantic priors are not optimal, as they are not always well executed or robustly tracked. The gesture labels corresponding to participants’ intentions do not always match those assigned during visual annotation, and poses are inadvertently modified during hand rotation. Our results provide valuable guidelines for designing effective gesture-controlled XR interfaces, including optimal dictionaries and system design choices.| File | Dimensione | Formato | |
|---|---|---|---|
|
s10055-026-01445-9_reference-2.pdf
accesso aperto
Tipologia:
Versione dell'editore
Licenza:
Creative commons
Dimensione
12.44 MB
Formato
Adobe PDF
|
12.44 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



