dorsal/arxiv
View SchemaPALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
| Authors | Yuanzhe Liu, Jingyuan Zhu, Yuchen Mo, Gen Li, Xu Cao, Jin Jin, Yifan Shen, Zhengyuan Li, Tianjiao Yu, Wenzhen Yuan, Fangqiang Ding, Ismini Lourentzou |
|---|---|
| Categories | |
| ArXiv ID | 2601.07060vv1 |
| URL | https://arxiv.org/abs/2601.07060 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify task-relevant interaction cues or track progress within a subtask, leading to critical execution errors such as repeated actions, missed steps, and premature termination. To address these challenges, we introduce PALM, a VLA framework that structures policy learning around interaction-centric affordance reasoning and subtask progress cues. PALM distills complementary affordance representations that capture object relevance, contact geometry, spatial placements, and motion dynamics, and serve as task-relevant anchors for visuomotor control. To further stabilize long-horizon execution, PALM predicts continuous within-subtask progress, enabling seamless subtask transitions. Across extensive simulation and real-world experiments, PALM consistently outperforms baselines, achieving a 91.8% success rate on LIBERO-LONG, a 12.5% improvement in average length on CALVIN ABC->D, and a 2x improvement over real-world baselines across three long-horizon generalization settings.
{
"annotation_id": "31c85931-5ad3-4c9d-94f2-3a2563179979",
"date_created": "2026-02-17T05:53:08.879000Z",
"date_modified": "2026-02-17T05:53:08.879000Z",
"file_hash": "51ecb867dda40458aedb0100944fa469f481e2be0fafa7bd296b37879ddd6f39",
"private": false,
"record": {
"abstract": "Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify task-relevant interaction cues or track progress within a subtask, leading to critical execution errors such as repeated actions, missed steps, and premature termination. To address these challenges, we introduce PALM, a VLA framework that structures policy learning around interaction-centric affordance reasoning and subtask progress cues. PALM distills complementary affordance representations that capture object relevance, contact geometry, spatial placements, and motion dynamics, and serve as task-relevant anchors for visuomotor control. To further stabilize long-horizon execution, PALM predicts continuous within-subtask progress, enabling seamless subtask transitions. Across extensive simulation and real-world experiments, PALM consistently outperforms baselines, achieving a 91.8% success rate on LIBERO-LONG, a 12.5% improvement in average length on CALVIN ABC-\u003eD, and a 2x improvement over real-world baselines across three long-horizon generalization settings.",
"arxiv_id": "2601.07060",
"authors": [
"Yuanzhe Liu",
"Jingyuan Zhu",
"Yuchen Mo",
"Gen Li",
"Xu Cao",
"Jin Jin",
"Yifan Shen",
"Zhengyuan Li",
"Tianjiao Yu",
"Wenzhen Yuan",
"Fangqiang Ding",
"Ismini Lourentzou"
],
"categories": [
"cs.RO"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation",
"url": "https://arxiv.org/abs/2601.07060",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "6a5ce10c-0b59-41de-ae73-bf4ad90f36eb",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}