dorsal/arxiv
View SchemaWiFo-M$^2$: Plug-and-Play Multi-Modal Sensing via Foundation Model to Empower Wireless Communications
| Authors | Haotian Zhang, Shijian Gao, Xiang Cheng |
|---|---|
| Categories | |
| ArXiv ID | 2601.09179vv1 |
| URL | https://arxiv.org/abs/2601.09179 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
The growing adoption of sensor-rich intelligent systems has boosted the use of multi-modal sensing to improve wireless communications. However, traditional methods require extensive manual design of data preprocessing, network architecture, and task-specific fine-tuning, which limits both development scalability and real-world deployment. To address this, we propose WiFo-M$^2$, a foundation model that can be easily plugged into existing deep learning-based transceivers for universal performance gains. To extract generalizable out-of-band (OOB) channel features from multi-modal sensing, we introduce ContraSoM, a contrastive pre-training strategy. Once pre-trained, WiFo-M$^2$ infers future OOB channel features from historical sensor data and strengthens feature robustness via modality-specific data augmentation. Experiments show that WiFo-M$^2$ improves performance across multiple transceiver designs and demonstrates strong generalization to unseen scenarios.
{
"annotation_id": "f6ddce48-9dd9-46e3-b94b-2c9a32698320",
"date_created": "2026-02-17T05:53:20.072000Z",
"date_modified": "2026-02-17T05:53:20.072000Z",
"file_hash": "205b08c4c1e7ec0f29623bf7503a2bda9199a94a2e8224b05103e9af7f12c982",
"private": false,
"record": {
"abstract": "The growing adoption of sensor-rich intelligent systems has boosted the use of multi-modal sensing to improve wireless communications. However, traditional methods require extensive manual design of data preprocessing, network architecture, and task-specific fine-tuning, which limits both development scalability and real-world deployment. To address this, we propose WiFo-M$^2$, a foundation model that can be easily plugged into existing deep learning-based transceivers for universal performance gains. To extract generalizable out-of-band (OOB) channel features from multi-modal sensing, we introduce ContraSoM, a contrastive pre-training strategy. Once pre-trained, WiFo-M$^2$ infers future OOB channel features from historical sensor data and strengthens feature robustness via modality-specific data augmentation. Experiments show that WiFo-M$^2$ improves performance across multiple transceiver designs and demonstrates strong generalization to unseen scenarios.",
"arxiv_id": "2601.09179",
"authors": [
"Haotian Zhang",
"Shijian Gao",
"Xiang Cheng"
],
"categories": [
"eess.SP"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "WiFo-M$^2$: Plug-and-Play Multi-Modal Sensing via Foundation Model to Empower Wireless Communications",
"url": "https://arxiv.org/abs/2601.09179",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "8cd32949-b515-4975-926f-c1dbf7aee87d",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}