dorsal/arxiv
View SchemaCosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
| Authors | Junyang Chen, Yuhang Jia, Hui Wang, Jiaming Zhou, Yaxin Han, Mengying Feng, Yong Qin |
|---|---|
| Categories | |
| ArXiv ID | 2601.05329vv1 |
| URL | https://arxiv.org/abs/2601.05329 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems suffer from complex preprocessing pipelines and a reliance on explicit external temporal alignment. Addressing these limitations, we propose CosyEdit, an end-to-end speech editing model adapted from CosyVoice through task-specific fine-tuning and an optimized inference procedure, which internalizes speech-text alignment while ensuring high consistency between the speech before and after editing. By fine-tuning on only 250 hours of supervised data from our curated GigaEdit dataset, our 400M-parameter model achieves reliable speech editing performance. Experiments on the RealEdit benchmark indicate that CosyEdit not only outperforms several billion-parameter language model baselines but also matches the performance of state-of-the-art cascade approaches. These results demonstrate that, with task-specific fine-tuning and inference optimization, robust and efficient speech editing capabilities can be unlocked from a zero-shot TTS model, yielding a novel and cost-effective end-to-end solution for high-quality speech editing.
{
"annotation_id": "d52dc75e-8edd-4e30-9ad9-bbee2c9c6681",
"date_created": "2026-02-17T05:53:04.735000Z",
"date_modified": "2026-02-17T05:53:04.735000Z",
"file_hash": "19d7785560e8743fa06f3b2b3ea37dc007e852c8342f61eb97b34afe2cb3fec5",
"private": false,
"record": {
"abstract": "Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems suffer from complex preprocessing pipelines and a reliance on explicit external temporal alignment. Addressing these limitations, we propose CosyEdit, an end-to-end speech editing model adapted from CosyVoice through task-specific fine-tuning and an optimized inference procedure, which internalizes speech-text alignment while ensuring high consistency between the speech before and after editing. By fine-tuning on only 250 hours of supervised data from our curated GigaEdit dataset, our 400M-parameter model achieves reliable speech editing performance. Experiments on the RealEdit benchmark indicate that CosyEdit not only outperforms several billion-parameter language model baselines but also matches the performance of state-of-the-art cascade approaches. These results demonstrate that, with task-specific fine-tuning and inference optimization, robust and efficient speech editing capabilities can be unlocked from a zero-shot TTS model, yielding a novel and cost-effective end-to-end solution for high-quality speech editing.",
"arxiv_id": "2601.05329",
"authors": [
"Junyang Chen",
"Yuhang Jia",
"Hui Wang",
"Jiaming Zhou",
"Yaxin Han",
"Mengying Feng",
"Yong Qin"
],
"categories": [
"cs.SD",
"eess.AS"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models",
"url": "https://arxiv.org/abs/2601.05329",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e54c4803-d921-48d0-a685-b971da136916",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}