dorsal/arxiv
View SchemaLearning Domain-Invariant Representations for Cross-Domain Image Registration via Scene-Appearance Disentanglement
| Authors | Jiahao Qin, Yiwen Wang |
|---|---|
| Categories | |
| ArXiv ID | 2601.08875vv2 |
| URL | https://arxiv.org/abs/2601.08875 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Image registration under domain shift remains a fundamental challenge in computer vision and medical imaging: when source and target images exhibit systematic intensity differences, the brightness constancy assumption underlying conventional registration methods is violated, rendering correspondence estimation ill-posed. We propose SAR-Net, a unified framework that addresses this challenge through principled scene-appearance disentanglement. Our key insight is that observed images can be decomposed into domain-invariant scene representations and domain-specific appearance codes, enabling registration via re-rendering rather than direct intensity matching. We establish theoretical conditions under which this decomposition enables consistent cross-domain alignment (Proposition 1) and prove that our scene consistency loss provides a sufficient condition for geometric correspondence in the shared latent space (Proposition 2). Empirically, we validate SAR-Net on bidirectional scanning microscopy, where coupled domain shift and geometric distortion create a challenging real-world testbed. Our method achieves 0.885 SSIM and 0.979 NCC, representing 3.1x improvement over the strongest baseline, while maintaining real-time performance (77 fps). Ablation studies confirm that both scene consistency and domain alignment losses are necessary: removing either degrades performance by 90% SSIM or causes 223x increase in latent alignment error, respectively. Code and data are available at https://github.com/D-ST-Sword/SAR-NET.
{
"annotation_id": "923770e3-7570-4466-9e97-f2a1ef05ab5b",
"date_created": "2026-02-17T05:53:20.166000Z",
"date_modified": "2026-02-17T05:53:20.166000Z",
"file_hash": "16b081beca41ae36bb2accd2233c70eb2c470105364e3d4ca447f1f1bdfa9730",
"private": false,
"record": {
"abstract": "Image registration under domain shift remains a fundamental challenge in computer vision and medical imaging: when source and target images exhibit systematic intensity differences, the brightness constancy assumption underlying conventional registration methods is violated, rendering correspondence estimation ill-posed. We propose SAR-Net, a unified framework that addresses this challenge through principled scene-appearance disentanglement. Our key insight is that observed images can be decomposed into domain-invariant scene representations and domain-specific appearance codes, enabling registration via re-rendering rather than direct intensity matching. We establish theoretical conditions under which this decomposition enables consistent cross-domain alignment (Proposition 1) and prove that our scene consistency loss provides a sufficient condition for geometric correspondence in the shared latent space (Proposition 2). Empirically, we validate SAR-Net on bidirectional scanning microscopy, where coupled domain shift and geometric distortion create a challenging real-world testbed. Our method achieves 0.885 SSIM and 0.979 NCC, representing 3.1x improvement over the strongest baseline, while maintaining real-time performance (77 fps). Ablation studies confirm that both scene consistency and domain alignment losses are necessary: removing either degrades performance by 90% SSIM or causes 223x increase in latent alignment error, respectively. Code and data are available at https://github.com/D-ST-Sword/SAR-NET.",
"arxiv_id": "2601.08875",
"authors": [
"Jiahao Qin",
"Yiwen Wang"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Learning Domain-Invariant Representations for Cross-Domain Image Registration via Scene-Appearance Disentanglement",
"url": "https://arxiv.org/abs/2601.08875",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "5ce49f4a-ac75-424e-9ea0-555ef5f1ea54",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}