Estimating the upper-body pose is essential for a wide range of applications, from rehabilitation therapy to industrial injury prevention. Current solutions face several limitations, such as heavy setup requirements, and often suffer from occlusion and drift issues. We propose a novel, sparse sensor fusion framework that estimates upper-limb kinematics using two shoulder-mounted imu and a chest-mounted camera with an integrated imu, without visual markers. The proposed pipeline integrates an orientation estimation module with a qpik solver. This formulation enables the system to enforce biomechanical constraints and utilize low-frequency visual input to correct inertial drift and recover joint positions that are unobservable by sparse imu alone. We evaluate the framework on the Total Capture dataset, where egocentric targets are simulated from motion-capture keypoints masked by a virtual chest-camera field of view, allowing the fusion layer to be assessed under controlled visual rate and accuracy. In this setting, our method achieves lower tracking error than adapted EKF-based visual-inertial baselines while using fewer wearable sensors. Compared with an IMU-only data-driven method, the proposed fusion pipeline achieves lower error with sufficiently frequent and accurate visual updates, demonstrating the benefit of exploiting intermittent egocentric constraints.

Upper-Limb Pose Estimation from Sparse Body-Worn Visual-Inertial Sensors

Ferdinando Pompanin;Enrico Martini
;
Andrea Calanca;Franco Fummi;Nicola Bombieri
2026-01-01

Abstract

Estimating the upper-body pose is essential for a wide range of applications, from rehabilitation therapy to industrial injury prevention. Current solutions face several limitations, such as heavy setup requirements, and often suffer from occlusion and drift issues. We propose a novel, sparse sensor fusion framework that estimates upper-limb kinematics using two shoulder-mounted imu and a chest-mounted camera with an integrated imu, without visual markers. The proposed pipeline integrates an orientation estimation module with a qpik solver. This formulation enables the system to enforce biomechanical constraints and utilize low-frequency visual input to correct inertial drift and recover joint positions that are unobservable by sparse imu alone. We evaluate the framework on the Total Capture dataset, where egocentric targets are simulated from motion-capture keypoints masked by a virtual chest-camera field of view, allowing the fusion layer to be assessed under controlled visual rate and accuracy. In this setting, our method achieves lower tracking error than adapted EKF-based visual-inertial baselines while using fewer wearable sensors. Compared with an IMU-only data-driven method, the proposed fusion pipeline achieves lower error with sufficiently frequent and accurate visual updates, demonstrating the benefit of exploiting intermittent egocentric constraints.
2026
Accelerometers , Biomechanics , Kalman filters , Motion capture , Motion analysis , Multimodal sensors , Multisensor systems , Sensor fusion
File in questo prodotto:
File Dimensione Formato  
2026_SPL.pdf

accesso aperto

Licenza: Creative commons
Dimensione 831.91 kB
Formato Adobe PDF
831.91 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11562/1203527
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact