dorsal/arxiv
View SchemaAction100M: A Large-scale Video Action Dataset
| Authors | Delong Chen, Tejaswi Kasarla, Yejin Bang, Mustafa Shukor, Willy Chung, Jade Yu, Allen Bolourchi, Theo Moutakanni, Pascale Fung |
|---|---|
| Categories | |
| ArXiv ID | 2601.10592vv1 |
| URL | https://arxiv.org/abs/2601.10592 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Inferring physical actions from visual observations is a fundamental capability for advancing machine intelligence in the physical world. Achieving this requires large-scale, open-vocabulary video action datasets that span broad domains. We introduce Action100M, a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding O(100 million) temporally localized segments with open-vocabulary action supervision and rich captions. Action100M is generated by a fully automated pipeline that (i) performs hierarchical temporal segmentation using V-JEPA 2 embeddings, (ii) produces multi-level frame and segment captions organized as a Tree-of-Captions, and (iii) aggregates evidence with a reasoning model (GPT-OSS-120B) under a multi-round Self-Refine procedure to output structured annotations (brief/detailed action, actor, brief/detailed caption). Training VL-JEPA on Action100M demonstrates consistent data-scaling improvements and strong zero-shot performance across diverse action recognition benchmarks, establishing Action100M as a new foundation for scalable research in video understanding and world modeling.
{
"annotation_id": "be214e38-86dd-4e4c-bcae-b93e6b89127e",
"date_created": "2026-02-17T05:53:23.744000Z",
"date_modified": "2026-02-17T05:53:23.744000Z",
"file_hash": "8edecc2c28b304e7b130fad54a69de174a586a89b4bfb52a7a7623b8ca54b708",
"private": false,
"record": {
"abstract": "Inferring physical actions from visual observations is a fundamental capability for advancing machine intelligence in the physical world. Achieving this requires large-scale, open-vocabulary video action datasets that span broad domains. We introduce Action100M, a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding O(100 million) temporally localized segments with open-vocabulary action supervision and rich captions. Action100M is generated by a fully automated pipeline that (i) performs hierarchical temporal segmentation using V-JEPA 2 embeddings, (ii) produces multi-level frame and segment captions organized as a Tree-of-Captions, and (iii) aggregates evidence with a reasoning model (GPT-OSS-120B) under a multi-round Self-Refine procedure to output structured annotations (brief/detailed action, actor, brief/detailed caption). Training VL-JEPA on Action100M demonstrates consistent data-scaling improvements and strong zero-shot performance across diverse action recognition benchmarks, establishing Action100M as a new foundation for scalable research in video understanding and world modeling.",
"arxiv_id": "2601.10592",
"authors": [
"Delong Chen",
"Tejaswi Kasarla",
"Yejin Bang",
"Mustafa Shukor",
"Willy Chung",
"Jade Yu",
"Allen Bolourchi",
"Theo Moutakanni",
"Pascale Fung"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Action100M: A Large-scale Video Action Dataset",
"url": "https://arxiv.org/abs/2601.10592",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "00137023-36c6-4535-bab3-93d219ed9627",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}