dorsal/arxiv
View SchemaWildRayZer: Self-supervised Large View Synthesis in Dynamic Environments
| Authors | Xuweiyi Chen, Wentao Zhou, Zezhou Cheng |
|---|---|
| Categories | |
| ArXiv ID | 2601.10716vv1 |
| URL | https://arxiv.org/abs/2601.10716 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
We present WildRayZer, a self-supervised framework for novel view synthesis (NVS) in dynamic environments where both the camera and objects move. Dynamic content breaks the multi-view consistency that static NVS models rely on, leading to ghosting, hallucinated geometry, and unstable pose estimation. WildRayZer addresses this by performing an analysis-by-synthesis test: a camera-only static renderer explains rigid structure, and its residuals reveal transient regions. From these residuals, we construct pseudo motion masks, distill a motion estimator, and use it to mask input tokens and gate loss gradients so supervision focuses on cross-view background completion. To enable large-scale training and evaluation, we curate Dynamic RealEstate10K (D-RE10K), a real-world dataset of 15K casually captured dynamic sequences, and D-RE10K-iPhone, a paired transient and clean benchmark for sparse-view transient-aware NVS. Experiments show that WildRayZer consistently outperforms optimization-based and feed-forward baselines in both transient-region removal and full-frame NVS quality with a single feed-forward pass.
{
"annotation_id": "ec5cd16a-d4b4-478e-be63-31f947d28028",
"date_created": "2026-02-17T05:53:26.460000Z",
"date_modified": "2026-02-17T05:53:26.460000Z",
"file_hash": "b256f623f382649def619df5768832729224ca9f27ee4ba9ee63a50b7b6d4e78",
"private": false,
"record": {
"abstract": "We present WildRayZer, a self-supervised framework for novel view synthesis (NVS) in dynamic environments where both the camera and objects move. Dynamic content breaks the multi-view consistency that static NVS models rely on, leading to ghosting, hallucinated geometry, and unstable pose estimation. WildRayZer addresses this by performing an analysis-by-synthesis test: a camera-only static renderer explains rigid structure, and its residuals reveal transient regions. From these residuals, we construct pseudo motion masks, distill a motion estimator, and use it to mask input tokens and gate loss gradients so supervision focuses on cross-view background completion. To enable large-scale training and evaluation, we curate Dynamic RealEstate10K (D-RE10K), a real-world dataset of 15K casually captured dynamic sequences, and D-RE10K-iPhone, a paired transient and clean benchmark for sparse-view transient-aware NVS. Experiments show that WildRayZer consistently outperforms optimization-based and feed-forward baselines in both transient-region removal and full-frame NVS quality with a single feed-forward pass.",
"arxiv_id": "2601.10716",
"authors": [
"Xuweiyi Chen",
"Wentao Zhou",
"Zezhou Cheng"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments",
"url": "https://arxiv.org/abs/2601.10716",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "9424314f-1115-4631-b1eb-0ffd28d3d277",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}