dorsal/arxiv
View SchemaGanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
| Authors | Shubhashis Roy Dipta, Khairul Mahbub, Nadia Najjar |
|---|---|
| Categories | |
| ArXiv ID | 2601.06767vv1 |
| URL | https://arxiv.org/abs/2601.06767 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
We present a Bengali mathematical reasoning model called GanitLLM (named after the Bangla word for mathematics, "Ganit"), together with a new difficulty-aware Bengali math corpus and a curriculum-based GRPO pipeline. Bengali is one of the world's most widely spoken languages, yet existing LLMs either reason in English and then translate, or simply fail on multi-step Bengali math, in part because reinforcement learning recipes are tuned for high-resource languages and collapse under reward sparsity in low-resource settings. To address this, we construct Ganit, a rigorously filtered and decontaminated Bengali math dataset with automatic difficulty tags derived from the pass@k of a strong evaluator model. Building on this dataset, we propose Curriculum-GRPO, which combines multi-stage training (SFT + GRPO) with difficulty-aware sampling and verifiable rewards for format, numerical correctness, and Bengali reasoning. On Bn-MGSM and Bn-MSVAMP, GanitLLM-4B improves over its Qwen3-4B base by +8 and +7 accuracy points, respectively, while increasing the percentage of Bengali reasoning tokens from 14% to over 88% and reducing average solution length from 943 to 193 words.
{
"annotation_id": "c4242885-1fd2-4ae6-a96e-cc9a169c0fa8",
"date_created": "2026-02-17T05:53:08.626000Z",
"date_modified": "2026-02-17T05:53:08.626000Z",
"file_hash": "7a970b854661c8e2fc7f0aad0d914017d7ad4ba1c5cb5c292be7277124d23919",
"private": false,
"record": {
"abstract": "We present a Bengali mathematical reasoning model called GanitLLM (named after the Bangla word for mathematics, \"Ganit\"), together with a new difficulty-aware Bengali math corpus and a curriculum-based GRPO pipeline. Bengali is one of the world\u0027s most widely spoken languages, yet existing LLMs either reason in English and then translate, or simply fail on multi-step Bengali math, in part because reinforcement learning recipes are tuned for high-resource languages and collapse under reward sparsity in low-resource settings. To address this, we construct Ganit, a rigorously filtered and decontaminated Bengali math dataset with automatic difficulty tags derived from the pass@k of a strong evaluator model. Building on this dataset, we propose Curriculum-GRPO, which combines multi-stage training (SFT + GRPO) with difficulty-aware sampling and verifiable rewards for format, numerical correctness, and Bengali reasoning. On Bn-MGSM and Bn-MSVAMP, GanitLLM-4B improves over its Qwen3-4B base by +8 and +7 accuracy points, respectively, while increasing the percentage of Bengali reasoning tokens from 14% to over 88% and reducing average solution length from 943 to 193 words.",
"arxiv_id": "2601.06767",
"authors": [
"Shubhashis Roy Dipta",
"Khairul Mahbub",
"Nadia Najjar"
],
"categories": [
"cs.CL",
"cs.AI",
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO",
"url": "https://arxiv.org/abs/2601.06767",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "d10793e3-4959-4f9a-bb8a-bad2f66da433",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}