dorsal/arxiv
View SchemaRigMo: Unifying Rig and Motion Learning for Generative Animation
| Authors | Hao Zhang, Jiahao Luo, Bohui Wan, Yizhou Zhao, Zongrui Li, Michael Vasilkovsky, Chaoyang Wang, Jian Wang, Narendra Ahuja, Bing Zhou |
|---|---|
| Categories | |
| ArXiv ID | 2601.06378vv1 |
| URL | https://arxiv.org/abs/2601.06378 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Despite significant progress in 4D generation, rig and motion, the core structural and dynamic components of animation are typically modeled as separate problems. Existing pipelines rely on ground-truth skeletons and skinning weights for motion generation and treat auto-rigging as an independent process, undermining scalability and interpretability. We present RigMo, a unified generative framework that jointly learns rig and motion directly from raw mesh sequences, without any human-provided rig annotations. RigMo encodes per-vertex deformations into two compact latent spaces: a rig latent that decodes into explicit Gaussian bones and skinning weights, and a motion latent that produces time-varying SE(3) transformations. Together, these outputs define an animatable mesh with explicit structure and coherent motion, enabling feed-forward rig and motion inference for deformable objects. Beyond unified rig-motion discovery, we introduce a Motion-DiT model operating in RigMo's latent space and demonstrate that these structure-aware latents can naturally support downstream motion generation tasks. Experiments on DeformingThings4D, Objaverse-XL, and TrueBones demonstrate that RigMo learns smooth, interpretable, and physically plausible rigs, while achieving superior reconstruction and category-level generalization compared to existing auto-rigging and deformation baselines. RigMo establishes a new paradigm for unified, structure-aware, and scalable dynamic 3D modeling.
{
"annotation_id": "8381b283-0a83-4a89-8f85-a96130b1b756",
"date_created": "2026-02-17T05:53:07.571000Z",
"date_modified": "2026-02-17T05:53:07.571000Z",
"file_hash": "7a41e0939a914c6a6661ec459673669d7c839792974fb2c7b9fb673f3c24d08a",
"private": false,
"record": {
"abstract": "Despite significant progress in 4D generation, rig and motion, the core structural and dynamic components of animation are typically modeled as separate problems. Existing pipelines rely on ground-truth skeletons and skinning weights for motion generation and treat auto-rigging as an independent process, undermining scalability and interpretability. We present RigMo, a unified generative framework that jointly learns rig and motion directly from raw mesh sequences, without any human-provided rig annotations. RigMo encodes per-vertex deformations into two compact latent spaces: a rig latent that decodes into explicit Gaussian bones and skinning weights, and a motion latent that produces time-varying SE(3) transformations. Together, these outputs define an animatable mesh with explicit structure and coherent motion, enabling feed-forward rig and motion inference for deformable objects. Beyond unified rig-motion discovery, we introduce a Motion-DiT model operating in RigMo\u0027s latent space and demonstrate that these structure-aware latents can naturally support downstream motion generation tasks. Experiments on DeformingThings4D, Objaverse-XL, and TrueBones demonstrate that RigMo learns smooth, interpretable, and physically plausible rigs, while achieving superior reconstruction and category-level generalization compared to existing auto-rigging and deformation baselines. RigMo establishes a new paradigm for unified, structure-aware, and scalable dynamic 3D modeling.",
"arxiv_id": "2601.06378",
"authors": [
"Hao Zhang",
"Jiahao Luo",
"Bohui Wan",
"Yizhou Zhao",
"Zongrui Li",
"Michael Vasilkovsky",
"Chaoyang Wang",
"Jian Wang",
"Narendra Ahuja",
"Bing Zhou"
],
"categories": [
"cs.GR"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "RigMo: Unifying Rig and Motion Learning for Generative Animation",
"url": "https://arxiv.org/abs/2601.06378",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "674497a4-14e3-4d22-93b5-0c16ce8b0e2e",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}