dorsal/arxiv
View SchemaLearning from Demonstrations via Capability-Aware Goal Sampling
| Authors | Yuanlin Duan, Yuning Wang, Wenjie Qiu, He Zhu |
|---|---|
| Categories | |
| ArXiv ID | 2601.08731vv1 |
| URL | https://arxiv.org/abs/2601.08731 |
| Journal | 39th Conference on Neural Information Processing Systems (NeurIPS 2025) |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Despite its promise, imitation learning often fails in long-horizon environments where perfect replication of demonstrations is unrealistic and small errors can accumulate catastrophically. We introduce Cago (Capability-Aware Goal Sampling), a novel learning-from-demonstrations method that mitigates the brittle dependence on expert trajectories for direct imitation. Unlike prior methods that rely on demonstrations only for policy initialization or reward shaping, Cago dynamically tracks the agent's competence along expert trajectories and uses this signal to select intermediate steps--goals that are just beyond the agent's current reach--to guide learning. This results in an adaptive curriculum that enables steady progress toward solving the full task. Empirical results demonstrate that Cago significantly improves sample efficiency and final performance across a range of sparse-reward, goal-conditioned tasks, consistently outperforming existing learning from-demonstrations baselines.
{
"annotation_id": "c4037a41-8a8c-4712-8e27-959b9e7b1c05",
"date_created": "2026-02-17T05:53:15.712000Z",
"date_modified": "2026-02-17T05:53:15.712000Z",
"file_hash": "5cef028dd1a8e87b18d2067d0324af046dbfbabfd3b9914ad0dc4618d14dc267",
"private": false,
"record": {
"abstract": "Despite its promise, imitation learning often fails in long-horizon environments where perfect replication of demonstrations is unrealistic and small errors can accumulate catastrophically. We introduce Cago (Capability-Aware Goal Sampling), a novel learning-from-demonstrations method that mitigates the brittle dependence on expert trajectories for direct imitation. Unlike prior methods that rely on demonstrations only for policy initialization or reward shaping, Cago dynamically tracks the agent\u0027s competence along expert trajectories and uses this signal to select intermediate steps--goals that are just beyond the agent\u0027s current reach--to guide learning. This results in an adaptive curriculum that enables steady progress toward solving the full task. Empirical results demonstrate that Cago significantly improves sample efficiency and final performance across a range of sparse-reward, goal-conditioned tasks, consistently outperforming existing learning from-demonstrations baselines.",
"arxiv_id": "2601.08731",
"authors": [
"Yuanlin Duan",
"Yuning Wang",
"Wenjie Qiu",
"He Zhu"
],
"categories": [
"cs.AI"
],
"journal_ref": "39th Conference on Neural Information Processing Systems (NeurIPS 2025)",
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Learning from Demonstrations via Capability-Aware Goal Sampling",
"url": "https://arxiv.org/abs/2601.08731",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "80cde97f-e1a9-4ce8-83d5-c01ff1778200",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}