dorsal/arxiv
View Schema$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials
| Authors | Trisha Das, Mandis Beigi, Jacob Aptekar, Jimeng Sun |
|---|---|
| Categories | |
| ArXiv ID | 2601.06300vv1 |
| URL | https://arxiv.org/abs/2601.06300 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most commonly amended component. We introduce \textit{eligibility criteria amendment prediction}, a novel NLP task that aims to forecast whether the eligibility criteria of an initial trial protocol will undergo future amendments. To support this task, we release $\texttt{AMEND++}$, a benchmark suite comprising two datasets: $\texttt{AMEND}$, which captures eligibility-criteria version histories and amendment labels from public clinical trials, and $\verb|AMEND_LLM|$, a refined subset curated using an LLM-based denoising pipeline to isolate substantive changes. We further propose $\textit{Change-Aware Masked Language Modeling}$ (CAMLM), a revision-aware pretraining strategy that leverages historical edits to learn amendment-sensitive representations. Experiments across diverse baselines show that CAMLM consistently improves amendment prediction, enabling more robust and cost-effective clinical trial design.
{
"annotation_id": "7d89fddf-a0b5-45b6-b122-57257a77449e",
"date_created": "2026-02-17T05:53:08.674000Z",
"date_modified": "2026-02-17T05:53:08.674000Z",
"file_hash": "230f33f3ee0d5d2ab2073a012f75d4599cc1a240790cec1473ade19d89ba4cbe",
"private": false,
"record": {
"abstract": "Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most commonly amended component. We introduce \\textit{eligibility criteria amendment prediction}, a novel NLP task that aims to forecast whether the eligibility criteria of an initial trial protocol will undergo future amendments. To support this task, we release $\\texttt{AMEND++}$, a benchmark suite comprising two datasets: $\\texttt{AMEND}$, which captures eligibility-criteria version histories and amendment labels from public clinical trials, and $\\verb|AMEND_LLM|$, a refined subset curated using an LLM-based denoising pipeline to isolate substantive changes. We further propose $\\textit{Change-Aware Masked Language Modeling}$ (CAMLM), a revision-aware pretraining strategy that leverages historical edits to learn amendment-sensitive representations. Experiments across diverse baselines show that CAMLM consistently improves amendment prediction, enabling more robust and cost-effective clinical trial design.",
"arxiv_id": "2601.06300",
"authors": [
"Trisha Das",
"Mandis Beigi",
"Jacob Aptekar",
"Jimeng Sun"
],
"categories": [
"cs.CL",
"cs.AI",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "$\\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials",
"url": "https://arxiv.org/abs/2601.06300",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "289bc518-942e-4818-8c56-c157bd23c90e",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}