dorsal/arxiv
View SchemaTowards Egocentric 3D Hand Pose Estimation in Unseen Domains
| Authors | Wiktor Mucha, Michael Wray, Martin Kampel |
|---|---|
| Categories | |
| ArXiv ID | 2601.06537vv1 |
| URL | https://arxiv.org/abs/2601.06537 |
| License | http://creativecommons.org/licenses/by-nc-sa/4.0/ |
Abstract
We present V-HPOT, a novel approach for improving the cross-domain performance of 3D hand pose estimation from egocentric images across diverse, unseen domains. State-of-the-art methods demonstrate strong performance when trained and tested within the same domain. However, they struggle to generalise to new environments due to limited training data and depth perception -- overfitting to specific camera intrinsics. Our method addresses this by estimating keypoint z-coordinates in a virtual camera space, normalised by focal length and image size, enabling camera-agnostic depth prediction. We further leverage this invariance to camera intrinsics to propose a self-supervised test-time optimisation strategy that refines the model's depth perception during inference. This is achieved by applying a 3D consistency loss between predicted and in-space scale-transformed hand poses, allowing the model to adapt to target domain characteristics without requiring ground truth annotations. V-HPOT significantly improves 3D hand pose estimation performance in cross-domain scenarios, achieving a 71% reduction in mean pose error on the H2O dataset and a 41% reduction on the AssemblyHands dataset. Compared to state-of-the-art methods, V-HPOT outperforms all single-stage approaches across all datasets and competes closely with two-stage methods, despite needing approximately x3.5 to x14 less data.
{
"annotation_id": "57013a8b-6cda-4cc1-aa0b-7f61cadf195e",
"date_created": "2026-02-17T05:53:07.771000Z",
"date_modified": "2026-02-17T05:53:07.771000Z",
"file_hash": "4bd487382872acccbe6e9e89974945c00883a5af1ed2c4e3e4302c045ceffdc7",
"private": false,
"record": {
"abstract": "We present V-HPOT, a novel approach for improving the cross-domain performance of 3D hand pose estimation from egocentric images across diverse, unseen domains. State-of-the-art methods demonstrate strong performance when trained and tested within the same domain. However, they struggle to generalise to new environments due to limited training data and depth perception -- overfitting to specific camera intrinsics. Our method addresses this by estimating keypoint z-coordinates in a virtual camera space, normalised by focal length and image size, enabling camera-agnostic depth prediction. We further leverage this invariance to camera intrinsics to propose a self-supervised test-time optimisation strategy that refines the model\u0027s depth perception during inference. This is achieved by applying a 3D consistency loss between predicted and in-space scale-transformed hand poses, allowing the model to adapt to target domain characteristics without requiring ground truth annotations. V-HPOT significantly improves 3D hand pose estimation performance in cross-domain scenarios, achieving a 71% reduction in mean pose error on the H2O dataset and a 41% reduction on the AssemblyHands dataset. Compared to state-of-the-art methods, V-HPOT outperforms all single-stage approaches across all datasets and competes closely with two-stage methods, despite needing approximately x3.5 to x14 less data.",
"arxiv_id": "2601.06537",
"authors": [
"Wiktor Mucha",
"Michael Wray",
"Martin Kampel"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by-nc-sa/4.0/",
"title": "Towards Egocentric 3D Hand Pose Estimation in Unseen Domains",
"url": "https://arxiv.org/abs/2601.06537",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "d1af4e19-c1ad-4667-9f9f-9c04598be41a",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}