論文の概要: DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text
- arxiv url: http://arxiv.org/abs/2610.00883v1
- Date: Thu, 01 Oct 2026 00:57:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:23.835487
- Title: DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text
- Title(参考訳): DeBERTa-ConPara:AI生成テキストのアタック・アウェアとデプロイ・リアリスティック検出
- Abstract要約: DeBERTa-ConParaは,HC3 Plus,M4,MAGE,RAIDで学習した文脈変換器エンコーダと,攻撃対応Unicode前処理を組み合わせたデプロイメント指向検出器である。
我々の中心的な発見は、前処理が適用された場所によって反対方向に作用することである。
2つの配置の因子的変化は、正規化推論による生のトレーニングを最高の設定として独立に識別する。
12種類の攻撃クラスのうち、ホモグリーフとゼロ幅空間の挿入は11.05%から1.12%から96.98%に増加した。
- 参考スコア(独自算出の注目度): 42.255101883505056
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Robust detection of AI-generated text under deployment conditions is challenging: distribution shifts across domains and generators, adversarial perturbations of the input surface, and the absence of target-domain labels for threshold calibration all degrade detectors that perform well in-domain. We present DeBERTa-ConPara, a deployment-oriented detector combining attack-aware Unicode preprocessing with a contextual transformer encoder trained over HC3 Plus, M4, MAGE and RAID. Our central finding is that preprocessing acts in opposite directions depending on where it is applied: normalising the training corpus deduplicates it, collapsing 35.4% of RAID rows into copies of their clean siblings and deleting the adversarial supervision, whereas normalising at inference is an effective defence. A factorial varying the two placements independently identifies raw training with normalised inference as the best configuration, reaching 99.61% AUROC, 99.01% TPR@5% FPR and 96.57% TPR@1% FPR on the official RAID hidden test, alongside 93.14% average balanced accuracy across HC3 Plus and MAGE under a fixed threshold. The gain is confined to two of twelve attack classes: homoglyph and zero-width-space insertion rise from 11.05% and 1.12% to 96.98%. The same signature reproduces in a zero-shot detector of different architecture, showing the effect belongs to the attacks rather than to our model. We additionally report two negative results: semantic-invariance augmentation through paraphrasing and supervised contrastive learning (ConPara) does not improve the best configuration, and the handcrafted feature-fusion branch is inert in distribution and harmful outside it.
- Abstract(参考訳): ドメインとジェネレータ間の分散シフト、入力表面の逆方向の摂動、しきい値のキャリブレーションのためのターゲットドメインラベルの欠如。
DeBERTa-ConParaは,HC3 Plus,M4,MAGE,RAIDで学習した文脈変換器エンコーダと,攻撃対応Unicode前処理を組み合わせたデプロイメント指向検出器である。
トレーニングコーパスの正規化は、それらを重複させ、RAID行の35.4%をクリーンな兄弟姉妹のコピーに分解し、敵の監督を削除し、一方、推論における正規化は効果的な防御である。
2つの配置の因子的変化は、正規化推論による生のトレーニングを最高の構成として識別し、公式のRAID隠蔽試験では99.61% AUROC、99.01% TPR@5% FPR、96.57% TPR@1% FPRに達し、HC3 PlusとMAGEの平均バランス精度は93.14%である。
12種類の攻撃クラスのうち、ホモグリーフとゼロ幅空間の挿入は11.05%から1.12%から96.98%に増加した。
同じシグネチャは異なるアーキテクチャのゼロショット検出器で再現され、その効果は我々のモデルではなく攻撃に属する。
パラフレージングと教師付きコントラスト学習(ConPara)による意味的不変性の増大は、最良の構成を改善せず、手作りのフィーチャーフュージョンブランチは、その外部で不活性であり有害である。
関連論文リスト
- Amplifying, Not Learning: Fine-Tuned AI Text Detectors Amplify a Pretrained Direction [51.56484100374058]
テキスト検出器は、事前訓練された典型軸を増幅する。
タスク監督前の生エンコーダでは、3つのアーキテクチャでNYT-vs-HC3 AUROC 0.806/0.944/0.834を達成する。
RoBERTaベースでは、生のプロジェクションは微調整を超えるが、RoBERTaベースでは、フル微調整は、試験された流線型人口の双方で生よりも識別を小さくする。
論文 参考訳(メタデータ) (2026-05-20T19:08:38Z) - GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives [48.545980031973556]
GAMBITは、インポスタ検出器を評価するための3つの評価モードと2つの独立したスコアを持つベンチマークである。
ベンチマークには、240の共進化型インポスタ戦略にまたがる27,804のラベル付きインスタンスのデータセットが付属している。
論文 参考訳(メタデータ) (2026-05-09T16:07:23Z) - Feature-Augmented Transformers for Robust AI-Text Detection Across Domains and Generators [0.10923877073891443]
我々はHC3 PLUSで変圧器ベースの検出器を訓練し、ホールドアウト検証におけるバランスの取れた精度を最大化することにより、単一判定閾値を校正する。
HC3 PLUS の領域内、マルチドメインのマルチジェネレータ M4 ベンチマークへのクロスデータセット転送、および外部 AI-Text-Detection-Pile 上での評価を行う。
我々の最良のモデル(DeBERTa-v3-base+FeatAttn)はM4上で85.9%のバランスの取れた精度を達成する。
論文 参考訳(メタデータ) (2026-05-05T16:52:26Z) - Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations [2.7620215077666557]
現代の検出器は敵の攻撃に弱いことで知られており、パラフレーズは効果的な回避技術として際立っている。
本稿では,まず,標準的な対人訓練の限界を定量化することにより,対人的堅牢性の比較研究を行う。
次に、新しい、はるかに回復力のある検出フレームワークを紹介します。
論文 参考訳(メタデータ) (2025-09-22T13:03:53Z) - SEAL: Steerable Reasoning Calibration of Large Language Models for Free [58.931194824519935]
大規模言語モデル(LLM)は、拡張チェーン・オブ・ソート(CoT)推論機構を通じて複雑な推論タスクに魅力的な機能を示した。
最近の研究では、CoT推論トレースにかなりの冗長性が示されており、これはモデル性能に悪影響を及ぼす。
我々は,CoTプロセスをシームレスに校正し,高い効率性を示しながら精度を向上する,トレーニング不要なアプローチであるSEALを紹介した。
論文 参考訳(メタデータ) (2025-04-07T02:42:07Z) - ODDR: Outlier Detection & Dimension Reduction Based Defense Against Adversarial Patches [4.4100683691177816]
敵対的攻撃は、機械学習モデルの信頼性の高いデプロイに重大な課題をもたらす。
パッチベースの敵攻撃に対処するための総合的な防御戦略である外乱検出・次元削減(ODDR)を提案する。
提案手法は,逆パッチに対応する入力特徴を外れ値として同定できるという観測に基づいている。
論文 参考訳(メタデータ) (2023-11-20T11:08:06Z) - Detection and Mitigation of Byzantine Attacks in Distributed Training [24.951227624475443]
ワーカノードの異常なビザンチン挙動は、トレーニングを脱線させ、推論の品質を損なう可能性がある。
最近の研究は、幅広い攻撃モデルを検討し、歪んだ勾配を補正するために頑健な集約と/または計算冗長性を探究している。
本研究では、強力な攻撃モデルについて検討する:$q$ omniscient adversaries with full knowledge of the defense protocol that can change from iteration to iteration to weak one: $q$ randomly selected adversaries with limited collusion abilities。
論文 参考訳(メタデータ) (2022-08-17T05:49:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。