dorsal/arxiv
View SchemaTagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
| Authors | Mingyue Huo, Yiwen Shao, Yuheng Zhang |
|---|---|
| Categories | |
| ArXiv ID | 2601.06896vv1 |
| URL | https://arxiv.org/abs/2601.06896 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via Serialized Output Training (SOT) to learn turn-taking dynamics; and (2) an interleaved time anchor mechanism that not only supports fine-grained timestamp prediction but also acts as a synchronization signal between semantic understanding and speaker tracking. Compared to previous works that primarily focus on speaker-attributed ASR or implicit diarization, TagSpeech addresses the challenge of fine-grained speaker-content alignment and explicitly models "who spoke what and when" in an end-to-end manner. Experiments on AMI and AliMeeting benchmarks demonstrate that our method achieves consistent improvements in Diarization Error Rate (DER) over strong end-to-end baselines, including Qwen-Omni and Gemini, particularly in handling complex speech overlaps. Moreover, TagSpeech employs a parameter-efficient training paradigm in which the LLM backbone is frozen and only lightweight projectors are trained, resulting in strong performance with low computational cost.
{
"annotation_id": "808d09ee-039c-4371-9d2c-66621821dc4c",
"date_created": "2026-02-17T05:53:08.653000Z",
"date_modified": "2026-02-17T05:53:08.653000Z",
"file_hash": "900fd7083e309c658b2692f3f79c5668ebf773def115f68c100cac6324635f3f",
"private": false,
"record": {
"abstract": "We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via Serialized Output Training (SOT) to learn turn-taking dynamics; and (2) an interleaved time anchor mechanism that not only supports fine-grained timestamp prediction but also acts as a synchronization signal between semantic understanding and speaker tracking. Compared to previous works that primarily focus on speaker-attributed ASR or implicit diarization, TagSpeech addresses the challenge of fine-grained speaker-content alignment and explicitly models \"who spoke what and when\" in an end-to-end manner. Experiments on AMI and AliMeeting benchmarks demonstrate that our method achieves consistent improvements in Diarization Error Rate (DER) over strong end-to-end baselines, including Qwen-Omni and Gemini, particularly in handling complex speech overlaps. Moreover, TagSpeech employs a parameter-efficient training paradigm in which the LLM backbone is frozen and only lightweight projectors are trained, resulting in strong performance with low computational cost.",
"arxiv_id": "2601.06896",
"authors": [
"Mingyue Huo",
"Yiwen Shao",
"Yuheng Zhang"
],
"categories": [
"eess.AS",
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding",
"url": "https://arxiv.org/abs/2601.06896",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "0d913088-3cba-4726-b0ac-d9ebd97030cd",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}