dorsal/arxiv
View SchemaFOCAL: A Novel Benchmarking Technique for Multi-modal Agents
| Authors | Aditya Choudhary, Anupam Purwar |
|---|---|
| Categories | |
| ArXiv ID | 2601.07367vv1 |
| URL | https://arxiv.org/abs/2601.07367 |
| License | http://creativecommons.org/licenses/by-nc-nd/4.0/ |
Abstract
With the recent advancements in reasoning capabilities, tool calling using MCP servers and Audio Language Models (ALMs), development and integration of multi-modal agents (with voice and text support) has come to the industry forefront. Cascading pipelines for voice agents still play a central role in the industry owing to their superior reasoning capabilities facilitated by LLMs. Although, cascading pipelines often present error propagation through the pipeline. We propose a framework, FOCAL to benchmark end-to-end reasoning, component-wise error propagation and error analysis for automated as well as human-assisted testing of multi-modal agents (voice to voice + text input). We also share two novel metrics viz. Reasoning and Semantic scores to evaluate efficacy of the agent in having meaningful conversations in voice mode.
{
"annotation_id": "fd0e0759-7040-4d73-930e-bc92a783eadc",
"date_created": "2026-02-17T05:53:11.818000Z",
"date_modified": "2026-02-17T05:53:11.818000Z",
"file_hash": "fe4961172c8e854125c2cdaa1dacac420ea6e3d000dcf9421eebe3cf2410e624",
"private": false,
"record": {
"abstract": "With the recent advancements in reasoning capabilities, tool calling using MCP servers and Audio Language Models (ALMs), development and integration of multi-modal agents (with voice and text support) has come to the industry forefront. Cascading pipelines for voice agents still play a central role in the industry owing to their superior reasoning capabilities facilitated by LLMs. Although, cascading pipelines often present error propagation through the pipeline. We propose a framework, FOCAL to benchmark end-to-end reasoning, component-wise error propagation and error analysis for automated as well as human-assisted testing of multi-modal agents (voice to voice + text input). We also share two novel metrics viz. Reasoning and Semantic scores to evaluate efficacy of the agent in having meaningful conversations in voice mode.",
"arxiv_id": "2601.07367",
"authors": [
"Aditya Choudhary",
"Anupam Purwar"
],
"categories": [
"cs.SD"
],
"license": "http://creativecommons.org/licenses/by-nc-nd/4.0/",
"title": "FOCAL: A Novel Benchmarking Technique for Multi-modal Agents",
"url": "https://arxiv.org/abs/2601.07367",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "c272c839-4603-4c79-bee7-ca52fba51399",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}