dorsal/arxiv
View SchemaROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding
| Authors | Xueyun Tian, Wei Li, Bingbing Xu, Heng Dong, Yuanzhuo Wang, Huawei Shen |
|---|---|
| Categories | |
| ArXiv ID | 2601.10323vv1 |
| URL | https://arxiv.org/abs/2601.10323 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing approaches suffer from disjointed capabilities: they typically exhibit incomplete modality support or lack autonomous proactive monitoring. To address this, we present ROMA, a real-time omni-multimodal assistant for unified reactive and proactive interaction. ROMA processes continuous inputs as synchronized multimodal units, aligning dense audio with discrete video frames to handle granularity mismatches. For online decision-making, we introduce a lightweight speak head that decouples response initiation from generation to ensure precise triggering without task conflict. We train ROMA with a curated streaming dataset and a two-stage curriculum that progressively optimizes for streaming format adaptation and proactive responsiveness. To standardize the fragmented evaluation landscape, we reorganize diverse benchmarks into a unified suite covering both proactive (alert, narration) and reactive (QA) settings. Extensive experiments across 12 benchmarks demonstrate ROMA achieves state-of-the-art performance on proactive tasks while competitive in reactive settings, validating its robustness in unified real-time omni-multimodal understanding.
{
"annotation_id": "9f7f08ee-7fda-4faa-8ed8-d9bd1016fd2b",
"date_created": "2026-02-17T05:53:23.629000Z",
"date_modified": "2026-02-17T05:53:23.629000Z",
"file_hash": "603589f539dfa78722eb5da9c1506929df578e548aca452c8b660c05e3c89145",
"private": false,
"record": {
"abstract": "Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing approaches suffer from disjointed capabilities: they typically exhibit incomplete modality support or lack autonomous proactive monitoring. To address this, we present ROMA, a real-time omni-multimodal assistant for unified reactive and proactive interaction. ROMA processes continuous inputs as synchronized multimodal units, aligning dense audio with discrete video frames to handle granularity mismatches. For online decision-making, we introduce a lightweight speak head that decouples response initiation from generation to ensure precise triggering without task conflict. We train ROMA with a curated streaming dataset and a two-stage curriculum that progressively optimizes for streaming format adaptation and proactive responsiveness. To standardize the fragmented evaluation landscape, we reorganize diverse benchmarks into a unified suite covering both proactive (alert, narration) and reactive (QA) settings. Extensive experiments across 12 benchmarks demonstrate ROMA achieves state-of-the-art performance on proactive tasks while competitive in reactive settings, validating its robustness in unified real-time omni-multimodal understanding.",
"arxiv_id": "2601.10323",
"authors": [
"Xueyun Tian",
"Wei Li",
"Bingbing Xu",
"Heng Dong",
"Yuanzhuo Wang",
"Huawei Shen"
],
"categories": [
"cs.CV",
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding",
"url": "https://arxiv.org/abs/2601.10323",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e871e9df-87d9-4a5f-8951-0b7bd4d7d7bf",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}