dorsal/arxiv
View SchemaMotion Focus Recognition in Fast-Moving Egocentric Video
| Authors | Daniel Hong, James Tribble, Hao Wang, Chaoyi Zhou, Ashish Bastola, Siyu Huang, Abolfazl Razi |
|---|---|
| Categories | |
| ArXiv ID | 2601.07154vv2 |
| URL | https://arxiv.org/abs/2601.07154 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of motion analysis in sports and other fast-movement scenarios. To bridge this gap, we propose a real-time motion focus recognition method that estimates the subject's locomotion intention from any egocentric video. Our approach leverages the foundation model for camera pose estimation and introduces system-level optimizations to enable efficient and scalable inference. Evaluated on a collected egocentric action dataset, our method achieves real-time performance with manageable memory consumption through a sliding batch inference strategy. This work makes motion-centric analysis practical for edge deployment and offers a complementary perspective to existing egocentric studies on sports and fast-movement activities.
{
"annotation_id": "a0f84b5a-8e61-438f-998b-925f4670df1c",
"date_created": "2026-02-17T05:53:11.957000Z",
"date_modified": "2026-02-17T05:53:11.957000Z",
"file_hash": "b8d66655a6ec635d65ec81d58120c14ad6dfc87df6703b847b48ff24b0f8df7a",
"private": false,
"record": {
"abstract": "From Vision-Language-Action (VLA) systems to robotics, existing egocentric datasets primarily focus on action recognition tasks, while largely overlooking the inherent role of motion analysis in sports and other fast-movement scenarios. To bridge this gap, we propose a real-time motion focus recognition method that estimates the subject\u0027s locomotion intention from any egocentric video. Our approach leverages the foundation model for camera pose estimation and introduces system-level optimizations to enable efficient and scalable inference. Evaluated on a collected egocentric action dataset, our method achieves real-time performance with manageable memory consumption through a sliding batch inference strategy. This work makes motion-centric analysis practical for edge deployment and offers a complementary perspective to existing egocentric studies on sports and fast-movement activities.",
"arxiv_id": "2601.07154",
"authors": [
"Daniel Hong",
"James Tribble",
"Hao Wang",
"Chaoyi Zhou",
"Ashish Bastola",
"Siyu Huang",
"Abolfazl Razi"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Motion Focus Recognition in Fast-Moving Egocentric Video",
"url": "https://arxiv.org/abs/2601.07154",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "008088e5-e809-40d4-8d28-2a451ddcd17c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}