dorsal/arxiv
View SchemaPrompt-Based Clarity Evaluation and Topic Detection in Political Question Answering
| Authors | Lavanya Prahallad, Sai Utkarsh Choudarypally, Pragna Prahallad, Pranathi Prahallad |
|---|---|
| Categories | |
| ArXiv ID | 2601.08176vv1 |
| URL | https://arxiv.org/abs/2601.08176 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Automatic evaluation of large language model (LLM) responses requires not only factual correctness but also clarity, particularly in political question-answering. While recent datasets provide human annotations for clarity and evasion, the impact of prompt design on automatic clarity evaluation remains underexplored. In this paper, we study prompt-based clarity evaluation using the CLARITY dataset from the SemEval 2026 shared task. We compare a GPT-3.5 baseline provided with the dataset against GPT-5.2 evaluated under three prompting strategies: simple prompting, chain-of-thought prompting, and chain-of-thought with few-shot examples. Model predictions are evaluated against human annotations using accuracy and class-wise metrics for clarity and evasion, along with hierarchical exact match. Results show that GPT-5.2 consistently outperforms the GPT-3.5 baseline on clarity prediction, with accuracy improving from 56 percent to 63 percent under chain-of-thought with few-shot prompting. Chain-of-thought prompting yields the highest evasion accuracy at 34 percent, though improvements are less stable across fine-grained evasion categories. We further evaluate topic identification and find that reasoning-based prompting improves accuracy from 60 percent to 74 percent relative to human annotations. Overall, our findings indicate that prompt design reliably improves high-level clarity evaluation, while fine-grained evasion and topic detection remain challenging despite structured reasoning prompts.
{
"annotation_id": "98f004e2-44f7-4745-89f4-e45e9f6b937c",
"date_created": "2026-02-17T05:53:15.943000Z",
"date_modified": "2026-02-17T05:53:15.943000Z",
"file_hash": "f6cd677cf9bf14a26ea42b102a409065ef2c681bc55f3df172b61f4072f3fbff",
"private": false,
"record": {
"abstract": "Automatic evaluation of large language model (LLM) responses requires not only factual correctness but also clarity, particularly in political question-answering. While recent datasets provide human annotations for clarity and evasion, the impact of prompt design on automatic clarity evaluation remains underexplored. In this paper, we study prompt-based clarity evaluation using the CLARITY dataset from the SemEval 2026 shared task. We compare a GPT-3.5 baseline provided with the dataset against GPT-5.2 evaluated under three prompting strategies: simple prompting, chain-of-thought prompting, and chain-of-thought with few-shot examples. Model predictions are evaluated against human annotations using accuracy and class-wise metrics for clarity and evasion, along with hierarchical exact match. Results show that GPT-5.2 consistently outperforms the GPT-3.5 baseline on clarity prediction, with accuracy improving from 56 percent to 63 percent under chain-of-thought with few-shot prompting. Chain-of-thought prompting yields the highest evasion accuracy at 34 percent, though improvements are less stable across fine-grained evasion categories. We further evaluate topic identification and find that reasoning-based prompting improves accuracy from 60 percent to 74 percent relative to human annotations. Overall, our findings indicate that prompt design reliably improves high-level clarity evaluation, while fine-grained evasion and topic detection remain challenging despite structured reasoning prompts.",
"arxiv_id": "2601.08176",
"authors": [
"Lavanya Prahallad",
"Sai Utkarsh Choudarypally",
"Pragna Prahallad",
"Pranathi Prahallad"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Prompt-Based Clarity Evaluation and Topic Detection in Political Question Answering",
"url": "https://arxiv.org/abs/2601.08176",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "8d501c88-5e38-49bd-823e-0f09d2ac93ed",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}