論文の概要: When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents
- arxiv url: http://arxiv.org/abs/2604.25213v1
- Date: Tue, 28 Apr 2026 04:33:26 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-29 16:49:17.715744
- Title: When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents
- Title(参考訳): GPT-Image-2は偽造文書を認識できない
- Authors: Jiaqi Wu, Yuchen Zhou, Dennis Tsang Ng, Xingyu Shen, Kidus Zewde, Ankit Raj, Tommy Duong, Simiao Ren,
- Abstract要約: OpenAIのGPT-Image-2は、本物とAIが編集した文書イメージの視覚的境界を効果的に消した。
AIForge-Doc v2は、3,066のGPT-Image-2ドキュメントフォージェリーのペアデータセットです。
- 参考スコア(独自算出の注目度): 8.591475917096318
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: OpenAI's GPT-Image-2 has effectively erased the visual boundary between authentic and AI-edited document images: a single number on a receipt can be replaced in under a second for a few cents. We release AIForge-Doc v2, a paired dataset of 3,066 GPT-Image-2 document forgeries with pixel-precise masks in DocTamper-compatible format, and benchmark four lines of defence: human inspectors (N=120, n=365 pair-votes via the public 2AFC site CanUSpotAI.com), TruFor (generic forensic), DocTamper (qcf-568, document-specific), and the same GPT-Image-2 model as a zero-shot self-judge -- asked, to avoid the trivial "image is mostly real" reading, whether any region was generated or edited by an AI image model. Human 2AFC accuracy is 0.501, indistinguishable from chance: even side-by-side, inspectors cannot tell GPT-Image-2 receipt forgeries from authentic counterparts. The three computational judges sit only modestly above (TruFor 0.599, DocTamper 0.585, self-judge 0.532). The self-judge fails consistently, not by chance: across five prompt strategies and four policies for handling ambiguous responses, AUC never rises above 0.59. To rule out the possibility that the two forensic detectors are broken on our source domain rather than blind to AI inpainting, we calibrate each on a same-domain traditional-tampering set built for its training distribution: TruFor reaches AUC 0.962 on cross-camera splicing of our dataset, DocTamper reaches 0.852 on cross-document OCR-token splicing with two-pass JPEG re-encoding. Both retain near-published performance on traditional tampering; switching to GPT-Image-2 inpainting drops AUC by 0.27-0.36 (0.962->0.599 TruFor; 0.852->0.585 DocTamper), isolating a detection gap specific to GPT-Image-2 inpainting. We release the dataset, pipeline, four-judge protocol, and calibration sets.
- Abstract(参考訳): OpenAIのGPT-Image-2は、本物とAIが編集した文書イメージの視覚的境界を効果的に消去した。
AIForge-Doc v2は、3,066のGPT-Image-2ドキュメントフォージェリのペアデータセットでDocTamper互換のフォーマットで、人間のインスペクタ(N=120, n=365ペアボイト)、公開2AFCサイトCanUSpotAI.com、TruFor(ジェネリック・フォレンジック)、DocTamper(qcf-568、ドキュメント特化)、およびゼロショットのセルフジャッジと同じGPT-Image-2モデルである。
人間2AFCの精度は0.501であり、偶然とは区別できない。
3人の審査員はわずかに上向き(TruFor 0.599, DocTamper 0.585, self-judge 0.532)である。
5つの迅速な戦略と曖昧な対応のための4つのポリシーで、AUCは0.59を超えない。
TruForはデータセットのクロスカメラスプリシングでAUC 0.962に達し、DocTamperはクロスドキュメントのOCR-tokenスプリシングで0.852に達した。
GPT-Image-2の塗装は0.27-0.36(0.962->0.599 TruFor; 0.852->0.585 DocTamper)に切り替え、GPT-Image-2の塗装に特有の検出ギャップを分離する。
データセット、パイプライン、4judgeプロトコル、キャリブレーションセットをリリースしています。
関連論文リスト
- GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment [5.547870160341159]
OpenAIによるGPT-image-2は、AI生成画像の分岐点である。
このデータセットは、公開されたTwitter/XポストからソースされたGPT-image-2生成イメージの最初のデータセットである。
データセットを4つの分析で特徴づける。
論文 参考訳(メタデータ) (2026-04-28T08:35:51Z) - DOCFORGE-BENCH: A Comprehensive 0-shot Benchmark for Document Forgery Detection and Analysis [6.329569532501952]
文書偽造検出のための最初の統一ゼロショットベンチマークであるDOCFORGE-BENCHを提案する。
テキスト改ざん、レシート偽造、ID文書操作にまたがる8つのデータセットにまたがる14の手法を評価する。
私たちの中心的な発見は、シングルスレッドプロトコルでは見えない、広範囲にわたるキャリブレーション障害です。
論文 参考訳(メタデータ) (2026-03-02T04:26:57Z) - AIForge-Doc: A Benchmark for Detecting AI-Forged Tampering in Financial and Form Documents [7.014776899553499]
我々は,ファイナンシャルおよびフォーム文書にピクセルレベルのアノテーションを付加した拡散モデルベースの塗り絵のみを対象とする,最初の専用ベンチマークであるAIForge-Docを紹介する。
TruFor、DocTamper、ゼロショットGPT-4oの3つの代表検出器をベンチマークした結果、既存のメソッドはすべて大幅に劣化していることがわかった。
論文 参考訳(メタデータ) (2026-02-24T05:37:35Z) - Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting [46.102790941920865]
2段階の文書画像解析モデルであるDolphin-v2を提案する。
第1段階では、Dolphin-v2 はレイアウト解析とともに文書型分類(デジタル生まれか写真か)を共同で行う。
第2段階では、撮影された文書は、幾何学的歪みを処理するために全ページとして一様に解析されるのに対し、デジタル生まれの文書は、検出されたレイアウトアンカーによって案内される要素的並列解析を行う。
論文 参考訳(メタデータ) (2026-02-05T07:09:57Z) - Impact of Labeling Inaccuracy and Image Noise on Tooth Segmentation in Panoramic Radiographs using Federated, Centralized and Local Learning [46.232038247686745]
フェデレートラーニング(FL)は、歯科診断AIにおけるプライバシー制約、不均一なデータ品質、一貫性のないラベル付けを緩和する。
複数のデータ破損シナリオを対象としたパノラマX線撮影において,FLと集中学習(CL)と局所学習(LL)を比較した。
論文 参考訳(メタデータ) (2025-09-08T11:07:47Z) - A Vessel Bifurcation Landmark Pair Dataset for Abdominal CT Deformable Image Registration (DIR) Validation [1.906989881803579]
変形可能な画像登録(DIR)は多くの診断および治療タスクにおいて実現可能な技術である。
このデータセットは腹部DIR検証のための第一種である。
ランドマークペアの数、精度、分布により、現在利用可能なもの以上の精度でDIRアルゴリズムの堅牢な検証が可能になる。
論文 参考訳(メタデータ) (2025-01-15T21:28:47Z) - Deep Unrestricted Document Image Rectification [110.61517455253308]
文書画像修正のための新しい統合フレームワークDocTr++を提案する。
我々は,階層型エンコーダデコーダ構造を多スケール表現抽出・解析に適用することにより,元のアーキテクチャをアップグレードする。
実際のテストセットとメトリクスをコントリビュートして、修正品質を評価します。
論文 参考訳(メタデータ) (2023-04-18T08:00:54Z) - DiT: Self-supervised Pre-training for Document Image Transformer [85.78807512344463]
自己教師付き文書画像変換モデルであるDiTを提案する。
さまざまなビジョンベースのDocument AIタスクでは,バックボーンネットワークとしてDiTを活用しています。
実験結果から, 自己教師付き事前訓練型DiTモデルにより, 新たな最先端結果が得られることが示された。
論文 参考訳(メタデータ) (2022-03-04T15:34:46Z) - DocTr: Document Image Transformer for Geometric Unwarping and
Illumination Correction [99.09177377916369]
文書画像の幾何学的および照明歪みに対処する文書画像変換器(DocTr)を提案する。
DocTrは20.02%のキャラクタエラー率(CER)を実現しています。
論文 参考訳(メタデータ) (2021-10-25T13:27:10Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。