dorsal/arxiv
View SchemaTEMPO: A Realistic Multi-Domain Benchmark for Temporal Reasoning-Intensive Retrieval
| Authors | Abdelrahman Abdallah, Mohammed Ali, Muhammad Abdul-Mageed, Adam Jatowt |
|---|---|
| Categories | |
| ArXiv ID | 2601.09523vv1 |
| URL | https://arxiv.org/abs/2601.09523 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Existing temporal QA benchmarks focus on simple fact-seeking queries from news corpora, while reasoning-intensive retrieval benchmarks lack temporal grounding. However, real-world information needs often require reasoning about temporal evolution and synthesizing evidence across time periods. We introduce TEMPO, the first benchmark combining temporal reasoning with reasoning-intensive retrieval across 13 domains. TEMPO features: (1) 1,730 complex queries requiring deep temporal reasoning such as tracking changes, identifying trends, or comparing cross-period evidence; (2) step-wise retrieval planning with 3,976 decomposed steps and gold documents mapped to each step for multi-hop evaluation; and (3) novel temporal metrics including Temporal Coverage@k and Temporal Precision@k measuring whether results span required time periods. Evaluation of 12 retrieval systems reveals substantial challenges: the best model (DiVeR) achieves only 32.0 NDCG@10 and 71.4\% Temporal Coverage@10, demonstrating difficulty in retrieving temporally complete evidence. We believe TEMPO provides a challenging benchmark for improving temporal reasoning in retrieval and RAG systems. Our code and data are available at https://github.com/tempo-bench/Tempo. See also our official website: https://tempo-bench.github.io/.
{
"annotation_id": "7c44bdc3-9826-4a7a-9ca1-e562d31caaf3",
"date_created": "2026-02-17T05:53:19.946000Z",
"date_modified": "2026-02-17T05:53:19.946000Z",
"file_hash": "9c29a4d99397e4b6301b9800aa47289bfb61e5d49c65cf61f52c4ad4fb9e5103",
"private": false,
"record": {
"abstract": "Existing temporal QA benchmarks focus on simple fact-seeking queries from news corpora, while reasoning-intensive retrieval benchmarks lack temporal grounding. However, real-world information needs often require reasoning about temporal evolution and synthesizing evidence across time periods. We introduce TEMPO, the first benchmark combining temporal reasoning with reasoning-intensive retrieval across 13 domains. TEMPO features: (1) 1,730 complex queries requiring deep temporal reasoning such as tracking changes, identifying trends, or comparing cross-period evidence; (2) step-wise retrieval planning with 3,976 decomposed steps and gold documents mapped to each step for multi-hop evaluation; and (3) novel temporal metrics including Temporal Coverage@k and Temporal Precision@k measuring whether results span required time periods. Evaluation of 12 retrieval systems reveals substantial challenges: the best model (DiVeR) achieves only 32.0 NDCG@10 and 71.4\\% Temporal Coverage@10, demonstrating difficulty in retrieving temporally complete evidence. We believe TEMPO provides a challenging benchmark for improving temporal reasoning in retrieval and RAG systems. Our code and data are available at https://github.com/tempo-bench/Tempo. See also our official website: https://tempo-bench.github.io/.",
"arxiv_id": "2601.09523",
"authors": [
"Abdelrahman Abdallah",
"Mohammed Ali",
"Muhammad Abdul-Mageed",
"Adam Jatowt"
],
"categories": [
"cs.IR"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "TEMPO: A Realistic Multi-Domain Benchmark for Temporal Reasoning-Intensive Retrieval",
"url": "https://arxiv.org/abs/2601.09523",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "cd6b985c-a55e-49a9-ad41-edc2c25a585d",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}