dorsal/arxiv
View SchemaReal2Sim based on Active Perception with automatically VLM-generated Behavior Trees
| Authors | Alessandro Adami, Sebastian Zudaire, Ruggero Carli, Pietro Falco |
|---|---|
| Categories | |
| ArXiv ID | 2601.08454vv1 |
| URL | https://arxiv.org/abs/2601.08454 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Constructing an accurate simulation model of real-world environments requires reliable estimation of physical parameters such as mass, geometry, friction, and contact surfaces. Traditional real-to-simulation (Real2Sim) pipelines rely on manual measurements or fixed, pre-programmed exploration routines, which limit their adaptability to varying tasks and user intents. This paper presents a Real2Sim framework that autonomously generates and executes Behavior Trees for task-specific physical interactions to acquire only the parameters required for a given simulation objective, without relying on pre-defined task templates or expert-designed exploration routines. Given a high-level user request, an incomplete simulation description, and an RGB observation of the scene, a vision-language model performs multi-modal reasoning to identify relevant objects, infer required physical parameters, and generate a structured Behavior Tree composed of elementary robotic actions. The resulting behavior is executed on a torque-controlled Franka Emika Panda, enabling compliant, contact-rich interactions for parameter estimation. The acquired measurements are used to automatically construct a physics-aware simulation. Experimental results on the real manipulator demonstrate estimation of object mass, surface height, and friction-related quantities across multiple scenarios, including occluded objects and incomplete prior models. The proposed approach enables interpretable, intent-driven, and autonomously Real2Sim pipelines, bridging high-level reasoning with physically-grounded robotic interaction.
{
"annotation_id": "6145f5bd-b525-4255-b15e-d42b43a98c2e",
"date_created": "2026-02-17T05:53:15.633000Z",
"date_modified": "2026-02-17T05:53:15.633000Z",
"file_hash": "8a6775a34e9ea0ef1a410c74468e0445da36917428c8193a1d2610a281c801fd",
"private": false,
"record": {
"abstract": "Constructing an accurate simulation model of real-world environments requires reliable estimation of physical parameters such as mass, geometry, friction, and contact surfaces. Traditional real-to-simulation (Real2Sim) pipelines rely on manual measurements or fixed, pre-programmed exploration routines, which limit their adaptability to varying tasks and user intents. This paper presents a Real2Sim framework that autonomously generates and executes Behavior Trees for task-specific physical interactions to acquire only the parameters required for a given simulation objective, without relying on pre-defined task templates or expert-designed exploration routines. Given a high-level user request, an incomplete simulation description, and an RGB observation of the scene, a vision-language model performs multi-modal reasoning to identify relevant objects, infer required physical parameters, and generate a structured Behavior Tree composed of elementary robotic actions. The resulting behavior is executed on a torque-controlled Franka Emika Panda, enabling compliant, contact-rich interactions for parameter estimation. The acquired measurements are used to automatically construct a physics-aware simulation. Experimental results on the real manipulator demonstrate estimation of object mass, surface height, and friction-related quantities across multiple scenarios, including occluded objects and incomplete prior models. The proposed approach enables interpretable, intent-driven, and autonomously Real2Sim pipelines, bridging high-level reasoning with physically-grounded robotic interaction.",
"arxiv_id": "2601.08454",
"authors": [
"Alessandro Adami",
"Sebastian Zudaire",
"Ruggero Carli",
"Pietro Falco"
],
"categories": [
"cs.RO"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Real2Sim based on Active Perception with automatically VLM-generated Behavior Trees",
"url": "https://arxiv.org/abs/2601.08454",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "69a9a0b6-bd15-4615-9eb3-8d7135fd5ce0",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}