dorsal/arxiv
View SchemaSIRR-LMM: Single-image Reflection Removal via Large Multimodal Model
| Authors | Yu Guo, Zhiqiang Lao, Xiyun Song, Yubin Zhou, Heather Yu |
|---|---|
| Categories | |
| ArXiv ID | 2601.07209vv1 |
| URL | https://arxiv.org/abs/2601.07209 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Glass surfaces create complex interactions of reflected and transmitted light, making single-image reflection removal (SIRR) challenging. Existing datasets suffer from limited physical realism in synthetic data or insufficient scale in real captures. We introduce a synthetic dataset generation framework that path-traces 3D glass models over real background imagery to create physically accurate reflection scenarios with varied glass properties, camera settings, and post-processing effects. To leverage the capabilities of Large Multimodal Model (LMM), we concatenate the image layers into a single composite input, apply joint captioning, and fine-tune the model using task-specific LoRA rather than full-parameter training. This enables our approach to achieve improved reflection removal and separation performance compared to state-of-the-art methods.
{
"annotation_id": "6fb11ee5-0ff3-4726-97b3-611a9d8ae6ac",
"date_created": "2026-02-17T05:53:12.001000Z",
"date_modified": "2026-02-17T05:53:12.001000Z",
"file_hash": "54eb165c3ff445eb1c951a0730518d2f170503836cc14d88959f1614f8aef1f7",
"private": false,
"record": {
"abstract": "Glass surfaces create complex interactions of reflected and transmitted light, making single-image reflection removal (SIRR) challenging. Existing datasets suffer from limited physical realism in synthetic data or insufficient scale in real captures. We introduce a synthetic dataset generation framework that path-traces 3D glass models over real background imagery to create physically accurate reflection scenarios with varied glass properties, camera settings, and post-processing effects. To leverage the capabilities of Large Multimodal Model (LMM), we concatenate the image layers into a single composite input, apply joint captioning, and fine-tune the model using task-specific LoRA rather than full-parameter training. This enables our approach to achieve improved reflection removal and separation performance compared to state-of-the-art methods.",
"arxiv_id": "2601.07209",
"authors": [
"Yu Guo",
"Zhiqiang Lao",
"Xiyun Song",
"Yubin Zhou",
"Heather Yu"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.GR"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "SIRR-LMM: Single-image Reflection Removal via Large Multimodal Model",
"url": "https://arxiv.org/abs/2601.07209",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "5b5b91e0-b14b-4255-b792-7d839a4e61d9",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}