dorsal/arxiv
View SchemaAffostruction: 3D Affordance Grounding with Generative Reconstruction
| Authors | Chunghyun Park, Seunghyeon Lee, Minsu Cho |
|---|---|
| Categories | |
| ArXiv ID | 2601.09211vv1 |
| URL | https://arxiv.org/abs/2601.09211 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
This paper addresses the problem of affordance grounding from RGBD images of an object, which aims to localize surface regions corresponding to a text query that describes an action on the object. While existing methods predict affordance regions only on visible surfaces, we propose Affostruction, a generative framework that reconstructs complete geometry from partial observations and grounds affordances on the full shape including unobserved regions. We make three core contributions: generative multi-view reconstruction via sparse voxel fusion that extrapolates unseen geometry while maintaining constant token complexity, flow-based affordance grounding that captures inherent ambiguity in affordance distributions, and affordance-driven active view selection that leverages predicted affordances for intelligent viewpoint sampling. Affostruction achieves 19.1 aIoU on affordance grounding (40.4\% improvement) and 32.67 IoU for 3D reconstruction (67.7\% improvement), enabling accurate affordance prediction on complete shapes.
{
"annotation_id": "629ea227-473d-42c7-8406-aa20af2e70b9",
"date_created": "2026-02-17T05:53:19.851000Z",
"date_modified": "2026-02-17T05:53:19.851000Z",
"file_hash": "002740b48edd7d90089ddae3de55e04471eb35e9344ef0238b8f55e4709543a3",
"private": false,
"record": {
"abstract": "This paper addresses the problem of affordance grounding from RGBD images of an object, which aims to localize surface regions corresponding to a text query that describes an action on the object. While existing methods predict affordance regions only on visible surfaces, we propose Affostruction, a generative framework that reconstructs complete geometry from partial observations and grounds affordances on the full shape including unobserved regions. We make three core contributions: generative multi-view reconstruction via sparse voxel fusion that extrapolates unseen geometry while maintaining constant token complexity, flow-based affordance grounding that captures inherent ambiguity in affordance distributions, and affordance-driven active view selection that leverages predicted affordances for intelligent viewpoint sampling. Affostruction achieves 19.1 aIoU on affordance grounding (40.4\\% improvement) and 32.67 IoU for 3D reconstruction (67.7\\% improvement), enabling accurate affordance prediction on complete shapes.",
"arxiv_id": "2601.09211",
"authors": [
"Chunghyun Park",
"Seunghyeon Lee",
"Minsu Cho"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Affostruction: 3D Affordance Grounding with Generative Reconstruction",
"url": "https://arxiv.org/abs/2601.09211",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "50b107ec-ec02-42ef-9ce7-8fe94292f77c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}