dorsal/arxiv
View SchemaImproving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
| Authors | Ewelina Gajewska, Katarzyna Budzynska, Jarosław A Chudziak |
|---|---|
| Categories | |
| ArXiv ID | 2601.09342vv2 |
| URL | https://arxiv.org/abs/2601.09342 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
This work proposes a contextualised detection framework for implicitly hateful speech, implemented as a multi-agent system comprising a central Moderator Agent and dynamically constructed Community Agents representing specific demographic groups. Our approach explicitly integrates socio-cultural context from publicly available knowledge sources, enabling identity-aware moderation that surpasses state-of-the-art prompting methods (zero-shot prompting, few-shot prompting, chain-of-thought prompting) and alternative approaches on a challenging ToxiGen dataset. We enhance the technical rigour of performance evaluation by incorporating balanced accuracy as a central metric of classification fairness that accounts for the trade-off between true positive and true negative rates. We demonstrate that our community-driven consultative framework significantly improves both classification accuracy and fairness across all target groups.
{
"annotation_id": "82ecf1f2-ae60-41dc-84c4-abebf16730e2",
"date_created": "2026-02-17T05:53:20.114000Z",
"date_modified": "2026-02-17T05:53:20.114000Z",
"file_hash": "05d72d8a4c6877fb43c5d2a92d2a021bc8846165b369549bc33b636a51597388",
"private": false,
"record": {
"abstract": "This work proposes a contextualised detection framework for implicitly hateful speech, implemented as a multi-agent system comprising a central Moderator Agent and dynamically constructed Community Agents representing specific demographic groups. Our approach explicitly integrates socio-cultural context from publicly available knowledge sources, enabling identity-aware moderation that surpasses state-of-the-art prompting methods (zero-shot prompting, few-shot prompting, chain-of-thought prompting) and alternative approaches on a challenging ToxiGen dataset. We enhance the technical rigour of performance evaluation by incorporating balanced accuracy as a central metric of classification fairness that accounts for the trade-off between true positive and true negative rates. We demonstrate that our community-driven consultative framework significantly improves both classification accuracy and fairness across all target groups.",
"arxiv_id": "2601.09342",
"authors": [
"Ewelina Gajewska",
"Katarzyna Budzynska",
"Jaros\u0142aw A Chudziak"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework",
"url": "https://arxiv.org/abs/2601.09342",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "1d9f5549-9d15-4a63-9902-294bbc7674e6",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}