論文の概要: Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law
- arxiv url: http://arxiv.org/abs/2608.09393v1
- Date: Mon, 10 Aug 2026 10:20:13 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-11 19:16:37.205945
- Title: Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law
- Title(参考訳): 法律RAGにおける時間的過ち--フランス税法のためのバージョン付きコーパスベンチマーク
- Authors: Rose Cymbler, Daniel Guez, Laurent Fabre,
- Abstract要約: 我々は,現在施行中の法律記事の体系的検索と引用という,時間的誤動作を特定し,定量化する。
われわれはFiscalQA Proを導入し、フランスの税法典の32,436項目のコーパスと、すべてモデルに固執した時間的推論トラックを組み合わせている。
オラクルなしのマルチバージョンインデックス上のエンドツーエンドレトリバーは、98.3%という厳密な値に達した。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treats the corpus as static; we argue legal question answering is a temporally-indexed retrieval problem. We introduce FiscalQA Pro, pairing a versioned corpus of 32,436 article-versions of the French tax code (93 years, 1938-2031) with an all-model-hard temporal-reasoning track: 209 scored, expert-reviewed questions across 33 CGI articles (221 released; twelve flagged out of the answerable scope). At selection time, no evaluated model recovered its date-applicable answer closed-book in any of four sampling draws, and the currently in-force text lacks the gold value for all but one of the scored questions. Answers are scored deterministically via atomic ground-truth "nuggets" (regex and numeric-with-tolerance), never LLM-as-judge: an LLM judge would inherit the temporal bias it is meant to score. Across eleven models (five frontier closed-API systems plus Gemini 2.5 Pro as a substitute entry, and five open-weight), parametric knowledge yields 3.0% mean strict accuracy and RAG over a static current-version corpus 2.7%. Static RAG retrieves the date-applicable version 0% of the time, confidently citing a real but inapplicable version. Our end-to-end retriever over a multi-version index, with no oracle, reaches 98.3% mean strict; an oracle-article ablation reaches 99.1%, locating the residual gap in first-stage recall, not version selection. We additionally release a version-aware jurisprudence dataset of 69,208 citation links, together with the corpus, benchmark, model responses, and pipeline code.
- Abstract(参考訳): 適用可能なバージョンが早期または将来のものである場合に、現在実施中の法的記事の体系的な検索と引用を行う。
標準法的なRAGはコーパスを静的として扱い、法的な質問応答は時間的にインデックス付けされた検索問題であると主張する。
本稿では、フランス税法(1938-2031年93年)の32,436条記事変換コーパスを、CGI33項目(221件、回答可能な範囲外12件)で209点のスコア付き専門家レビュー付き質問トラックと組み合わせて、FiscalQA Proを紹介した。
選別時に、4つのサンプリングドローのいずれかで日付適用可能な回答のクローズブックを復元する評価モデルはなく、現在使われているインフォーステキストは、評価された質問のうち1つを除いて、すべてに金の価値を欠いている。
答えはアトミック・グラウンド・トゥルース・ナゲット(英語版) (regex and numeric-with-tolerance) を通じて決定的に採点され、LLM-as-judge: LLM判事は、スコアするはずの時間バイアスを継承する。
11モデル(5つのフロンティアクローズドAPIシステムと5つのGemini 2.5 Pro、および5つのオープンウェイト)のパラメトリック知識は、静的な電流変換コーパス2.7%よりも3.0%高い精度とRAGをもたらす。
静的RAGは、日時適用可能なバージョン0%を取得し、本物だが適用できないバージョンを確実に引用する。
オラクルを含まないマルチバージョンインデックス上のエンド・ツー・エンド・レトリバーは、98.3%の厳密な平均値に達し、オラクル・アーティクル・アブレーションは99.1%に達し、バージョン選択ではなく、ファーストステージリコールにおける残差を突き止める。
さらに、69,208個の引用リンクと、コーパス、ベンチマーク、モデル応答、パイプラインコードを組み合わせた、バージョン対応のジャリスプルーデンスデータセットをリリースする。
関連論文リスト
- Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding [3.285110528753401]
法律、税法、ソフトウェア文書などの進化した文書は、修正され、置き換えられ、時には時間とともに復活する。
我々は、バングラデシュ政府によって発行された644の公式税関の計3,050対のQAに関する専門家が検証したベンチマークであるTIDEを提示する。
論文 参考訳(メタデータ) (2026-08-09T06:17:41Z) - When Memory Updates but Behavior Does Not: Repairing Implicit Stale Dependencies in Personalized Agent Responses [1.5770673503612611]
メモリ拡張されたエージェントは、ユーザの格納された状態が時代遅れであることを知ることができ、なおも古い値を計画する。
私たちは1つの構造的コントリビュータを特定します。 ドラフトアンコールによる検証は、応答が何を言っているかをチェックします。
論文 参考訳(メタデータ) (2026-08-03T02:49:08Z) - Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks [40.92390378341581]
APIの廃止やコード再構成といった技術的なコーパスの時間的変化は、既存のベンチマークを不安定にすることができる。
我々は2024年10月から2025年10月にかけて、FreshStackの2つの独立したコーパススナップショットを調査し、LangChainに関する質問に答える。
論文 参考訳(メタデータ) (2026-03-04T19:18:11Z) - Recon, Answer, Verify: Agents in Search of Truth [36.56689822791777]
Politi Fact Only (PFO)は、politifact.comの2,982件の政治的主張のベンチマークデータセットである。
すべてのポストクレーム分析とアノテーションキューが手作業で削除された。
本稿では,質問生成,回答生成,ラベル生成という3つのエージェントからなるエージェントフレームワークであるRAVを提案する。
論文 参考訳(メタデータ) (2025-07-04T15:44:28Z) - LEXam: Benchmarking Legal Reasoning on 340 Law Exams [76.3521146499006]
textscLEXamは,法科116科の法科試験を対象とする340件の法科試験を対象とする,新しいベンチマークである。
このデータセットは、英語とドイツ語で4,886の法試験質問で構成されており、その中には2,841の長文のオープンエンド質問と2,045の多重選択質問が含まれている。
この結果から,モデル間の差分化におけるデータセットの有効性が示唆された。
論文 参考訳(メタデータ) (2025-05-19T08:48:12Z) - IfQA: A Dataset for Open-domain Question Answering under Counterfactual
Presuppositions [54.23087908182134]
本稿では,QA(FifQA)と呼ばれる,最初の大規模対実的オープンドメイン質問応答(QA)ベンチマークを紹介する。
IfQAデータセットには3,800以上の質問が含まれている。
IfQAベンチマークによって引き起こされるユニークな課題は、検索と対実的推論の両方に関して、オープンドメインのQA研究を促進することである。
論文 参考訳(メタデータ) (2023-05-23T12:43:19Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。