dorsal/arxiv
View SchemaOpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
| Authors | Sheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang, Cheng Sun |
|---|---|
| Categories | |
| ArXiv ID | 2601.09575vv1 |
| URL | https://arxiv.org/abs/2601.09575 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
We propose OpenVoxel, a training-free algorithm for grouping and captioning sparse voxels for the open-vocabulary 3D scene understanding tasks. Given the sparse voxel rasterization (SVR) model obtained from multi-view images of a 3D scene, our OpenVoxel is able to produce meaningful groups that describe different objects in the scene. Also, by leveraging powerful Vision Language Models (VLMs) and Multi-modal Large Language Models (MLLMs), our OpenVoxel successfully build an informative scene map by captioning each group, enabling further 3D scene understanding tasks such as open-vocabulary segmentation (OVS) or referring expression segmentation (RES). Unlike previous methods, our method is training-free and does not introduce embeddings from a CLIP/BERT text encoder. Instead, we directly proceed with text-to-text search using MLLMs. Through extensive experiments, our method demonstrates superior performance compared to recent studies, particularly in complex referring expression segmentation (RES) tasks. The code will be open.
{
"annotation_id": "8f97029f-111f-4b81-88f5-4b2bfe99399a",
"date_created": "2026-02-17T05:53:19.964000Z",
"date_modified": "2026-02-17T05:53:19.964000Z",
"file_hash": "9a12aebecd92da5cd19d90c28b79b6ed44e2017637f68297072ec28b8bffad02",
"private": false,
"record": {
"abstract": "We propose OpenVoxel, a training-free algorithm for grouping and captioning sparse voxels for the open-vocabulary 3D scene understanding tasks. Given the sparse voxel rasterization (SVR) model obtained from multi-view images of a 3D scene, our OpenVoxel is able to produce meaningful groups that describe different objects in the scene. Also, by leveraging powerful Vision Language Models (VLMs) and Multi-modal Large Language Models (MLLMs), our OpenVoxel successfully build an informative scene map by captioning each group, enabling further 3D scene understanding tasks such as open-vocabulary segmentation (OVS) or referring expression segmentation (RES). Unlike previous methods, our method is training-free and does not introduce embeddings from a CLIP/BERT text encoder. Instead, we directly proceed with text-to-text search using MLLMs. Through extensive experiments, our method demonstrates superior performance compared to recent studies, particularly in complex referring expression segmentation (RES) tasks. The code will be open.",
"arxiv_id": "2601.09575",
"authors": [
"Sheng-Yu Huang",
"Jaesung Choe",
"Yu-Chiang Frank Wang",
"Cheng Sun"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding",
"url": "https://arxiv.org/abs/2601.09575",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "c11a37eb-503d-40a0-a10b-2fd2173dd7ab",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}