dorsal/arxiv
View SchemaImproving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework
| Authors | Ewelina Gajewska, Katarzyna Budzynska, Jarosław A Chudziak |
|---|---|
| Categories | |
| ArXiv ID | 2601.09342vv1 |
| URL | https://arxiv.org/abs/2601.09342 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
This work proposes a contextualised detection framework for implicitly hateful speech, implemented as a multi-agent system comprising a central Moderator Agent and dynamically constructed Community Agents representing specific demographic groups. Our approach explicitly integrates socio-cultural context from publicly available knowledge sources, enabling identity-aware moderation that surpasses state-of-the-art prompting methods (zero-shot prompting, few-shot prompting, chain-of-thought prompting) and alternative approaches on a challenging ToxiGen dataset. We enhance the technical rigour of performance evaluation by incorporating balanced accuracy as a central metric of classification fairness that accounts for the trade-off between true positive and true negative rates. We demonstrate that our community-driven consultative framework significantly improves both classification accuracy and fairness across all target groups.
{
"annotation_id": "04eceb95-fb8c-4409-985d-d684bb297e55",
"date_created": "2026-02-17T05:53:20.113000Z",
"date_modified": "2026-02-17T05:53:20.113000Z",
"file_hash": "0502a207539d5cf9527dd2b8c719714299270792e697a036c6171275c796fee7",
"private": false,
"record": {
"abstract": "This work proposes a contextualised detection framework for implicitly hateful speech, implemented as a multi-agent system comprising a central Moderator Agent and dynamically constructed Community Agents representing specific demographic groups. Our approach explicitly integrates socio-cultural context from publicly available knowledge sources, enabling identity-aware moderation that surpasses state-of-the-art prompting methods (zero-shot prompting, few-shot prompting, chain-of-thought prompting) and alternative approaches on a challenging ToxiGen dataset. We enhance the technical rigour of performance evaluation by incorporating balanced accuracy as a central metric of classification fairness that accounts for the trade-off between true positive and true negative rates. We demonstrate that our community-driven consultative framework significantly improves both classification accuracy and fairness across all target groups.",
"arxiv_id": "2601.09342",
"authors": [
"Ewelina Gajewska",
"Katarzyna Budzynska",
"Jaros\u0142aw A Chudziak"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework",
"url": "https://arxiv.org/abs/2601.09342",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "7bb8d8eb-c473-401e-9036-08186163256d",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}