dorsal/arxiv
View SchemaForensicFormer: Hierarchical Multi-Scale Reasoning for Cross-Domain Image Forgery Detection
| Authors | Hema Hariharan Samson |
|---|---|
| Categories | |
| ArXiv ID | 2601.08873vv1 |
| URL | https://arxiv.org/abs/2601.08873 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
The proliferation of AI-generated imagery and sophisticated editing tools has rendered traditional forensic methods ineffective for cross-domain forgery detection. We present ForensicFormer, a hierarchical multi-scale framework that unifies low-level artifact detection, mid-level boundary analysis, and high-level semantic reasoning via cross-attention transformers. Unlike prior single-paradigm approaches, which achieve <75% accuracy on out-of-distribution datasets, our method maintains 86.8% average accuracy across seven diverse test sets, spanning traditional manipulations, GAN-generated images, and diffusion model outputs - a significant improvement over state-of-the-art universal detectors. We demonstrate superior robustness to JPEG compression (83% accuracy at Q=70 vs. 66% for baselines) and provide pixel-level forgery localization with a 0.76 F1-score. Extensive ablation studies validate that each hierarchical component contributes 4-10% accuracy improvement, and qualitative analysis reveals interpretable forensic features aligned with human expert reasoning. Our work bridges classical image forensics and modern deep learning, offering a practical solution for real-world deployment where manipulation techniques are unknown a priori.
{
"annotation_id": "b5f23b53-6294-4f7a-b532-8769b416f463",
"date_created": "2026-02-17T05:53:19.290000Z",
"date_modified": "2026-02-17T05:53:19.290000Z",
"file_hash": "f176bf2bc63b4e059e1ffbcd7f806fe6a72657793b19fc94b8521920b01ed9d3",
"private": false,
"record": {
"abstract": "The proliferation of AI-generated imagery and sophisticated editing tools has rendered traditional forensic methods ineffective for cross-domain forgery detection. We present ForensicFormer, a hierarchical multi-scale framework that unifies low-level artifact detection, mid-level boundary analysis, and high-level semantic reasoning via cross-attention transformers. Unlike prior single-paradigm approaches, which achieve \u003c75% accuracy on out-of-distribution datasets, our method maintains 86.8% average accuracy across seven diverse test sets, spanning traditional manipulations, GAN-generated images, and diffusion model outputs - a significant improvement over state-of-the-art universal detectors. We demonstrate superior robustness to JPEG compression (83% accuracy at Q=70 vs. 66% for baselines) and provide pixel-level forgery localization with a 0.76 F1-score. Extensive ablation studies validate that each hierarchical component contributes 4-10% accuracy improvement, and qualitative analysis reveals interpretable forensic features aligned with human expert reasoning. Our work bridges classical image forensics and modern deep learning, offering a practical solution for real-world deployment where manipulation techniques are unknown a priori.",
"arxiv_id": "2601.08873",
"authors": [
"Hema Hariharan Samson"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.LG",
"cs.MM"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "ForensicFormer: Hierarchical Multi-Scale Reasoning for Cross-Domain Image Forgery Detection",
"url": "https://arxiv.org/abs/2601.08873",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "b1d7568e-130b-468e-994d-00d3d6e67518",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}