論文の概要: Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing
- arxiv url: http://arxiv.org/abs/2607.04572v1
- Date: Mon, 06 Jul 2026 00:50:27 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-07 22:26:29.982175
- Title: Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing
- Title(参考訳): LLMを用いた学習指導者における回答型推論の検出
- Authors: Bonan Shen, Dingyan Shang, Youting Wang, Tao Ning,
- Abstract要約: 本研究では,個人回答情報によって教師説明が解答駆動的になるかどうかを考察する。
我々は,質問専用,正解キー,誤解キーの3つの学習文脈下で1000 GSM8Kテスト問題を評価する。
Qwen2.5-3B-Instructでは、回答キーアクセスは中央値のTRACE AUCを0.375から0.900に上昇させ、1000件中997件で最初の10%プレフィックスで金の回答を利用できる。
- 参考スコア(独自算出の注目度): 0.25999037208435705
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Large language model (LLM) tutors often produce fluent step-by-step explanations, but a correct and pedagogically formatted response does not guarantee that the answer was derived from the student-facing problem. In realistic tutoring systems, the model may also have access to teacher notes, answer keys, rubrics, or retrieved solution artifacts. We study whether such private answer information can make tutor explanations answer-driven: the final answer is behaviorally available before the written explanation has justified it. Using Truncated Reasoning AUC Evaluation (TRACE), which probes how early a chain-of-thought prefix can pass a verifier, we evaluate 1000 GSM8K test problems under three paired tutoring contexts: question-only, correct answer-key, and wrong answer-key. At fixed fractions of each generated explanation, we force the model to answer immediately and verify the response against the gold numeric answer. With Qwen2.5-3B-Instruct, answer-key access raises median TRACE AUC from 0.375 to 0.900 and makes the gold answer available at the first 10% prefix in 997 of 1000 cases. The effect remains strong on the 746 examples where both question-only and answer-key explanations end with the correct answer. These results support truncated CoT auditing as a lightweight process-level diagnostic for answer-driven reasoning in math tutoring explanations.
- Abstract(参考訳): 大規模言語モデル (LLM) のチューターは、しばしばステップ・バイ・ステップの説明を流用するが、正確で教育的に形式化された応答は、その答えが学生が直面する問題に由来することを保証しない。
現実的なチューターシステムでは、このモデルは教師のノート、回答キー、ルーリック、あるいは検索されたソリューションアーティファクトへのアクセスも可能である。
筆者らは,このようなプライベートな回答情報によって,教師による説明が解答に結びつくかどうかを考察した。
そこで,Truncated Reasoning AUC Evaluation (TRACE) を用いて,1000 GSM8K検定問題を,質問専用,正解キー,誤解キーの3つの文脈で検証した。
生成された各説明の固定分数において、モデルに即座に答えさせ、金の数値解に対する応答を検証する。
Qwen2.5-3B-Instructでは、回答キーアクセスは中央値のTRACE AUCを0.375から0.900に上昇させ、1000件中997件で最初の10%プレフィックスで金の回答を利用できる。
この効果は、質問のみと回答キーの両方が正しい回答で終わる746の例に強く残っている。
これらの結果は、数学の授業説明において、回答駆動推論のための軽量なプロセスレベル診断として、truncated CoT監査をサポートする。
関連論文リスト
- Localizing and Mitigating Errors in Long-form Question Answering [79.63372684264921]
LFQA(Long-form Question answering)は、複雑な質問に対して徹底的で深い回答を提供し、理解を深めることを目的としている。
この研究は、人書きおよびモデル生成LFQA回答の局所的エラーアノテーションを備えた最初の幻覚データセットであるHaluQuestQAを紹介する。
論文 参考訳(メタデータ) (2024-07-16T17:23:16Z) - Learn to Explain: Multimodal Reasoning via Thought Chains for Science
Question Answering [124.16250115608604]
本稿では,SQA(Science Question Answering)について紹介する。SQA(Science Question Answering)は,21万のマルチモーダルな複数選択質問と多様な科学トピックと,それに対応する講義や説明による回答の注釈からなる新しいベンチマークである。
また,SQAでは,数ショットのGPT-3では1.20%,微調整のUnifiedQAでは3.99%の改善が見られた。
我々の分析は、人間に似た言語モデルは、より少ないデータから学習し、わずか40%のデータで同じパフォーマンスを達成するのに、説明の恩恵を受けることを示している。
論文 参考訳(メタデータ) (2022-09-20T07:04:24Z) - Teaching language models to support answers with verified quotes [12.296242080730831]
オープンブック”QAモデルをトレーニングし、その一方で、その主張に関する具体的な証拠を引用しています。
2800億のパラメータモデルであるGopherCiteは、高品質なサポートエビデンスで回答を生成し、不確実な場合には回答を控えることができます。
論文 参考訳(メタデータ) (2022-03-21T17:26:29Z) - Reinforcement Learning from Reformulations in Conversational Question
Answering over Knowledge Graphs [28.507683735633464]
本研究では,質問や修正の会話の流れから学習できる強化学習モデル「ConQUER」を提案する。
実験では、CONQUERが騒々しい報酬信号から会話の質問に答えることに成功したことが示されています。
論文 参考訳(メタデータ) (2021-05-11T08:08:35Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。