dorsal/arxiv
View SchemaMotion Focus Recognition in Fast-Moving Egocentric Video
| Authors | Daniel Hong, James Tribble, Hao Wang, Chaoyi Zhou, Ashish Bastola, Siyu Huang, Abolfazl Razi |
|---|---|
| Categories | |
| ArXiv ID | 2601.07154vv1 |
| URL | https://arxiv.org/abs/2601.07154 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of motion analysis in sports and other fast-movement scenarios. To bridge this gap, we propose a real-time motion focus recognition method that estimates the subject's locomotion intention from any egocentric video. Our approach leverages the foundation model for camera pose estimation and introduces system-level optimizations to enable efficient and scalable inference. Evaluated on a collected egocentric action dataset, our method achieves real-time performance with manageable memory consumption through a sliding batch inference strategy. This work makes motion-centric analysis practical for edge deployment and offers a complementary perspective to existing egocentric studies on sports and fast-movement activities.
{
"annotation_id": "d4e9d984-4f82-45c3-884f-eb22ba160390",
"date_created": "2026-02-17T05:53:11.279000Z",
"date_modified": "2026-02-17T05:53:11.279000Z",
"file_hash": "839142f6aba2404d40477a1cbbd5f40696716f6c530ea916747e36d53778b05b",
"private": false,
"record": {
"abstract": "From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of motion analysis in sports and other fast-movement scenarios. To bridge this gap, we propose a real-time motion focus recognition method that estimates the subject\u0027s locomotion intention from any egocentric video. Our approach leverages the foundation model for camera pose estimation and introduces system-level optimizations to enable efficient and scalable inference. Evaluated on a collected egocentric action dataset, our method achieves real-time performance with manageable memory consumption through a sliding batch inference strategy. This work makes motion-centric analysis practical for edge deployment and offers a complementary perspective to existing egocentric studies on sports and fast-movement activities.",
"arxiv_id": "2601.07154",
"authors": [
"Daniel Hong",
"James Tribble",
"Hao Wang",
"Chaoyi Zhou",
"Ashish Bastola",
"Siyu Huang",
"Abolfazl Razi"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Motion Focus Recognition in Fast-Moving Egocentric Video",
"url": "https://arxiv.org/abs/2601.07154",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "70b47e2d-7748-41a4-a5ab-f34b171c0bca",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}