dorsal/arxiv
View SchemaTowards Realistic Synthetic Data for Automatic Drum Transcription
| Authors | Pierfrancesco Melucci, Paolo Merialdo, Taketo Akama |
|---|---|
| Categories | |
| ArXiv ID | 2601.09520vv1 |
| URL | https://arxiv.org/abs/2601.09520 |
| License | http://creativecommons.org/licenses/by-sa/4.0/ |
Abstract
Deep learning models define the state-of-the-art in Automatic Drum Transcription (ADT), yet their performance is contingent upon large-scale, paired audio-MIDI datasets, which are scarce. Existing workarounds that use synthetic data often introduce a significant domain gap, as they typically rely on low-fidelity SoundFont libraries that lack acoustic diversity. While high-quality one-shot samples offer a better alternative, they are not available in a standardized, large-scale format suitable for training. This paper introduces a new paradigm for ADT that circumvents the need for paired audio-MIDI training data. Our primary contribution is a semi-supervised method to automatically curate a large and diverse corpus of one-shot drum samples from unlabeled audio sources. We then use this corpus to synthesize a high-quality dataset from MIDI files alone, which we use to train a sequence-to-sequence transcription model. We evaluate our model on the ENST and MDB test sets, where it achieves new state-of-the-art results, significantly outperforming both fully supervised methods and previous synthetic-data approaches. The code for reproducing our experiments is publicly available at https://github.com/pier-maker92/ADT_STR
{
"annotation_id": "46d6c8fc-c08f-4108-8f67-7e5b93021ec3",
"date_created": "2026-02-17T05:53:20.349000Z",
"date_modified": "2026-02-17T05:53:20.349000Z",
"file_hash": "10761804b1f70f929436d56ddf064cfbe921a1614db1dede7cb67c74c7a4493d",
"private": false,
"record": {
"abstract": "Deep learning models define the state-of-the-art in Automatic Drum Transcription (ADT), yet their performance is contingent upon large-scale, paired audio-MIDI datasets, which are scarce. Existing workarounds that use synthetic data often introduce a significant domain gap, as they typically rely on low-fidelity SoundFont libraries that lack acoustic diversity. While high-quality one-shot samples offer a better alternative, they are not available in a standardized, large-scale format suitable for training. This paper introduces a new paradigm for ADT that circumvents the need for paired audio-MIDI training data. Our primary contribution is a semi-supervised method to automatically curate a large and diverse corpus of one-shot drum samples from unlabeled audio sources. We then use this corpus to synthesize a high-quality dataset from MIDI files alone, which we use to train a sequence-to-sequence transcription model. We evaluate our model on the ENST and MDB test sets, where it achieves new state-of-the-art results, significantly outperforming both fully supervised methods and previous synthetic-data approaches. The code for reproducing our experiments is publicly available at https://github.com/pier-maker92/ADT_STR",
"arxiv_id": "2601.09520",
"authors": [
"Pierfrancesco Melucci",
"Paolo Merialdo",
"Taketo Akama"
],
"categories": [
"cs.SD",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by-sa/4.0/",
"title": "Towards Realistic Synthetic Data for Automatic Drum Transcription",
"url": "https://arxiv.org/abs/2601.09520",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "83587645-3a9b-4273-b11a-5572fd089d96",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}