dorsal/arxiv
View SchemaA Multi-Stage Workflow for the Review of Marketing Content with Reasoning Large Language Models
| Authors | Alberto Purpura, Emily Chen, Swapnil Shinde |
|---|---|
| Categories | |
| ArXiv ID | 2601.06054vv1 |
| URL | https://arxiv.org/abs/2601.06054 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Reasoning Large Language Models (LLMs) have shown promising results when tasked with solving complex problems. In this paper, we propose and evaluate a multi-stage workflow that leverages the capabilities of fine-tuned reasoning LLMs to assist in the review process of marketing content, making sure they comply with a given list of requirements. The contributions of this paper are the following: (i) we present a novel approach -- that does not rely on any external knowledge representation -- for the automatic identification of compliance issues in textual content; (ii) compare the effectiveness of different fine-tuning strategies like Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO) in training models to solve this problem; (iii) we evaluate the effectiveness of training small LLMs to generate reasoning tokens before providing their final response; (iv) we evaluate how the choice and combinations of different reward functions affects the performance of a model trained with GRPO.
{
"annotation_id": "60f144e7-ff8e-4416-9253-ac61261004dd",
"date_created": "2026-02-17T05:53:04.132000Z",
"date_modified": "2026-02-17T05:53:04.132000Z",
"file_hash": "11fb9ecefcaf51ded7110719eb959870d05bbe8e2a59102b1b11ab679ddb1c4a",
"private": false,
"record": {
"abstract": "Reasoning Large Language Models (LLMs) have shown promising results when tasked with solving complex problems. In this paper, we propose and evaluate a multi-stage workflow that leverages the capabilities of fine-tuned reasoning LLMs to assist in the review process of marketing content, making sure they comply with a given list of requirements. The contributions of this paper are the following: (i) we present a novel approach -- that does not rely on any external knowledge representation -- for the automatic identification of compliance issues in textual content; (ii) compare the effectiveness of different fine-tuning strategies like Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO) in training models to solve this problem; (iii) we evaluate the effectiveness of training small LLMs to generate reasoning tokens before providing their final response; (iv) we evaluate how the choice and combinations of different reward functions affects the performance of a model trained with GRPO.",
"arxiv_id": "2601.06054",
"authors": [
"Alberto Purpura",
"Emily Chen",
"Swapnil Shinde"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "A Multi-Stage Workflow for the Review of Marketing Content with Reasoning Large Language Models",
"url": "https://arxiv.org/abs/2601.06054",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "f30deed4-35fa-4695-8001-62f0e3fd933c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}