AIForge-Doc: A Benchmark for Detecting AI-Forged Tampering in Financial and Form Documents
- URL: http://arxiv.org/abs/2602.20569v1
- Date: Tue, 24 Feb 2026 05:37:35 GMT
- Title: AIForge-Doc: A Benchmark for Detecting AI-Forged Tampering in Financial and Form Documents
- Authors: Jiaqi Wu, Yuchen Zhou, Muduo Xu, Zisheng Liang, Simiao Ren, Jiayu Xue, Meige Yang, Siying Chen, Jingheng Huan,
- Abstract summary: We present AIForge-Doc, the first dedicated benchmark targeting exclusively diffusion-model-based inpainting in financial and form documents with pixel-level annotation.<n>We benchmark three representative detectors -- TruFor, DocTamper, and a zero-shot GPT-4o judge -- and find that all existing methods degrade substantially.
- Score: 7.014776899553499
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: We present AIForge-Doc, the first dedicated benchmark targeting exclusively diffusion-model-based inpainting in financial and form documents with pixel-level annotation. Existing document forgery datasets rely on traditional digital editing tools (e.g., Adobe Photoshop, GIMP), creating a critical gap: state-of-the-art detectors are blind to the rapidly growing threat of AI-forged document fraud. AIForge-Doc addresses this gap by systematically forging numeric fields in real-world receipt and form images using two AI inpainting APIs -- Gemini 2.5 Flash Image and Ideogram v2 Edit -- yielding 4,061 forged images from four public document datasets (CORD, WildReceipt, SROIE, XFUND) across nine languages, annotated with pixel-precise tampered-region masks in DocTamper-compatible format. We benchmark three representative detectors -- TruFor, DocTamper, and a zero-shot GPT-4o judge -- and find that all existing methods degrade substantially: TruFor achieves AUC=0.751 (zero-shot, out-of-distribution) vs. AUC=0.96 on NIST16; DocTamper achieves AUC=0.563 vs. AUC=0.98 in-distribution, with pixel-level IoU=0.020; GPT-4o achieves only 0.509 -- essentially at chance -- confirming that AI-forged values are indistinguishable to automated detectors and VLMs. These results demonstrate that AIForge-Doc represents a qualitatively new and unsolved challenge for document forensics.
Related papers
- DOCFORGE-BENCH: A Comprehensive Benchmark for Document Forgery Detection and Analysis [7.0914801556869]
We present DOCFORGE-BENCH, the first unified zero-shot benchmark for document forgery detection.<n>We evaluate 14 methods across eight datasets spanning text tampering, receipt forgery, and identity document manipulation.<n>Our central finding is a pervasive calibration failure invisible under single-threshold protocols.
arXiv Detail & Related papers (2026-03-02T04:26:57Z) - Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting [46.102790941920865]
We present Dolphin-v2, a two-stage document image parsing model that substantially improves upon the original Dolphin.<n>In the first stage, Dolphin-v2 jointly performs document type classification (digital-born versus photographed) alongside layout analysis.<n>In the second stage, we employ a hybrid parsing strategy: photographed documents are parsed holistically as complete pages to handle geometric distortions, while digital-born documents undergo element-wise parallel parsing guided by the detected layout anchors.
arXiv Detail & Related papers (2026-02-05T07:09:57Z) - EdgeDoc: Hybrid CNN-Transformer Model for Accurate Forgery Detection and Localization in ID Documents [6.690084812573466]
EdgeDoc is a novel approach for the detection and localization of document forgeries.<n>Our architecture combines a lightweight convolutional transformer with auxiliary noiseprint features extracted from the images.
arXiv Detail & Related papers (2025-08-22T10:45:14Z) - CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AI [58.35348718345307]
Current efforts to distinguish between real and AI-generated images may lack generalization.<n>We propose a novel framework, Co-Spy, that first enhances existing semantic features.<n>We also create Co-Spy-Bench, a comprehensive dataset comprising 5 real image datasets and 22 state-of-the-art generative models.
arXiv Detail & Related papers (2025-03-24T01:59:29Z) - Zero-Shot Detection of AI-Generated Images [54.01282123570917]
We propose a zero-shot entropy-based detector (ZED) to detect AI-generated images.
Inspired by recent works on machine-generated text detection, our idea is to measure how surprising the image under analysis is compared to a model of real images.
ZED achieves an average improvement of more than 3% over the SoTA in terms of accuracy.
arXiv Detail & Related papers (2024-09-24T08:46:13Z) - UNIT: Unifying Image and Text Recognition in One Vision Encoder [51.140564856352825]
UNIT is a novel training framework aimed at UNifying Image and Text recognition within a single model.
We show that UNIT significantly outperforms existing methods on document-related tasks.
Notably, UNIT retains the original vision encoder architecture, making it cost-free in terms of inference and deployment.
arXiv Detail & Related papers (2024-09-06T08:02:43Z) - Watermark Text Pattern Spotting in Document Images [3.6298655794854464]
In the wild, writing can come in various fonts, sizes and forms, making generic recognition a very difficult problem.
We propose a novel benchmark (K-Watermark) containing 65,447 data samples generated using Wrender.
A validity study using humans raters yields an authenticity score of 0.51 against pre-generated watermarked documents.
arXiv Detail & Related papers (2024-01-10T14:02:45Z) - DocMAE: Document Image Rectification via Self-supervised Representation
Learning [144.44748607192147]
We present DocMAE, a novel self-supervised framework for document image rectification.
We first mask random patches of the background-excluded document images and then reconstruct the missing pixels.
With such a self-supervised learning approach, the network is encouraged to learn the intrinsic structure of deformed documents.
arXiv Detail & Related papers (2023-04-20T14:27:15Z) - Deep Unrestricted Document Image Rectification [110.61517455253308]
We present DocTr++, a novel unified framework for document image rectification.
We upgrade the original architecture by adopting a hierarchical encoder-decoder structure for multi-scale representation extraction and parsing.
We contribute a real-world test set and metrics applicable for evaluating the rectification quality.
arXiv Detail & Related papers (2023-04-18T08:00:54Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.