dorsal/arxiv
View SchemaThe Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios
| Authors | Daocheng Fu, Jianbiao Mei, Rong Wu, Xuemeng Yang, Jia Xu, Ding Wang, Pinlong Cai, Yong Liu, Licheng Wen, Botian Shi |
|---|---|
| Categories | |
| ArXiv ID | 2601.08173vv1 |
| URL | https://arxiv.org/abs/2601.08173 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
The rapid evolution of Multi-modal Large Language Models (MLLMs) has advanced workflow automation; however, existing research mainly targets performance upper bounds in static environments, overlooking robustness for stochastic real-world deployment. We identify three key challenges: dynamic task scheduling, active exploration under uncertainty, and continuous learning from experience. To bridge this gap, we introduce \method{}, a dynamic evaluation environment that simulates a "trainee" agent continuously exploring a novel setting. Unlike traditional benchmarks, \method{} evaluates agents along three dimensions: (1) context-aware scheduling for streaming tasks with varying priorities; (2) prudent information acquisition to reduce hallucination via active exploration; and (3) continuous evolution by distilling generalized strategies from rule-based, dynamically generated tasks. Experiments show that cutting-edge agents have significant deficiencies in dynamic environments, especially in active exploration and continual learning. Our work establishes a framework for assessing agent reliability, shifting evaluation from static tests to realistic, production-oriented scenarios. Our codes are available at https://github.com/KnowledgeXLab/EvoEnv
{
"annotation_id": "874ae857-acab-41fc-a41e-84d8b9917fbc",
"date_created": "2026-02-17T05:53:15.942000Z",
"date_modified": "2026-02-17T05:53:15.942000Z",
"file_hash": "9869200dd2385c9a4cf6b7922af20b5ad771b13bc0692c21463f38e8e4056eb4",
"private": false,
"record": {
"abstract": "The rapid evolution of Multi-modal Large Language Models (MLLMs) has advanced workflow automation; however, existing research mainly targets performance upper bounds in static environments, overlooking robustness for stochastic real-world deployment. We identify three key challenges: dynamic task scheduling, active exploration under uncertainty, and continuous learning from experience. To bridge this gap, we introduce \\method{}, a dynamic evaluation environment that simulates a \"trainee\" agent continuously exploring a novel setting. Unlike traditional benchmarks, \\method{} evaluates agents along three dimensions: (1) context-aware scheduling for streaming tasks with varying priorities; (2) prudent information acquisition to reduce hallucination via active exploration; and (3) continuous evolution by distilling generalized strategies from rule-based, dynamically generated tasks. Experiments show that cutting-edge agents have significant deficiencies in dynamic environments, especially in active exploration and continual learning. Our work establishes a framework for assessing agent reliability, shifting evaluation from static tests to realistic, production-oriented scenarios. Our codes are available at https://github.com/KnowledgeXLab/EvoEnv",
"arxiv_id": "2601.08173",
"authors": [
"Daocheng Fu",
"Jianbiao Mei",
"Rong Wu",
"Xuemeng Yang",
"Jia Xu",
"Ding Wang",
"Pinlong Cai",
"Yong Liu",
"Licheng Wen",
"Botian Shi"
],
"categories": [
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "The Agent\u0027s First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios",
"url": "https://arxiv.org/abs/2601.08173",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "83b22dc6-5c05-4169-adf8-9366b6933ace",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}