論文の概要: Quantization Thresholds Replicate, Failure Modes Do Not: A Three-Model Study of Agentic Tool Use in Polish from 8-bit to 2-bit
- arxiv url: http://arxiv.org/abs/2609.32042v1
- Date: Fri, 25 Sep 2026 22:13:11 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-08 03:36:12.06544
- Title: Quantization Thresholds Replicate, Failure Modes Do Not: A Three-Model Study of Agentic Tool Use in Polish from 8-bit to 2-bit
- Title(参考訳): 量子化閾値は重複し、失敗モードはしない:8ビットから2ビットまでのポーランドにおけるエージェントツールの3モデル研究
- Abstract要約: PolAgentBenchはポーランドのプロンプトと英語のツールスキーマを用いた決定論的ベンチマークである。
ビエリク-11B-v3.0とその熟成された蒸留子ビエリク-ミニトロン-7B-v3.0はモデル圧縮を分離する。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/publicdomain/zero/1.0/
- Abstract: We ask how GGUF quantization affects agentic tool use in Polish and whether the effects generalize across models. We introduce PolAgentBench, a deterministic benchmark with Polish prompts and English tool schemas: a 67-task main suite (15 adversarial probes, 52 hard-tier tasks) and a 46-task arithmetic isolation ladder. Three models span two axes of variation: Bielik-11B-v3.0 and its pruned, distilled child Bielik-Minitron-7B-v3.0 isolate model compression, and Llama-PLLuM-8B adds a change of pretraining family. Each is measured at six precisions, Q8_0 to Q2_K. Only the collapse threshold replicates. (1) All three models fall off a cliff between 3-bit and 2-bit (11B 0.716 to 0.045, 7B 0.463 to 0.149, PLLuM 0.224 to 0.015; paired McNemar p < 0.001 in each), across a fourfold capability spread and both axes. (2) Failure modes do not replicate: at 2-bit the 7B fails long (median 9.1k tokens, 4 steps) while the 11B mostly answers at the first step with a confabulated final answer (37 of 64 failures); PLLuM fails on content across precisions (71.8-92.0% of steps parse). (3) On the arithmetic ladder the unscaffolded rung is a floor, left standing by a rerun that states the no-tool rule; four explicit calls lift the 8-bit 11B from 1/10 to 9/10 and the 7B from 0/10 to 7/10 after format-only failures with the gold value are forgiven, an exploratory effect with eight distinct baseline inputs that does not survive multiplicity correction, while the order-trap arm separates the models at 8-bit (11B 6/6, 7B 0/6). (4) The Polish-versus-English gap is associated with degradation or with task family. We document four artifacts that shaped our conclusions (rounding-hostile gold values, strict answer typing, a no-tool rule the prompt never stated, priority-ordered failure labels), report affected results in strict and corrected form, and release the benchmark, trajectories and commit-stamped artifacts.
- Abstract(参考訳): GGUFの量子化がポーランドにおけるエージェントツールの使用にどのように影響するか、モデル全体で効果が一般化されるかどうかを問う。
ポーランドのプロンプトと英語のツールスキーマを用いた決定論的ベンチマークであるPolAgentBenchを紹介した。
Bielik-11B-v3.0とその熟成された蒸留子Bielik-Minitron-7B-v3.0はモデル圧縮を分離し、Llama-PLLuM-8Bはプレトレーニングファミリーを変更する。
それぞれ6つの精度、Q8_0〜Q2_Kで測定する。
崩壊しきい値のみが複製される。
1) 3モデルとも3ビットから2ビットの崖(11B 0.716 - 0.045, 7B 0.463 - 0.149, PLLuM 0.224 - 0.015, それぞれMcNemar p < 0.001)から4倍の速さで落下する。
2) 障害モードは複製されない: 2ビットで7Bは長めに失敗する(中間9.1kトークン、4ステップ)が、11Bは最終回答(64回のうち37回)で最初のステップでほとんど答える。
算術ラグにおいて、非スキャフォールドのラングは、ノーツールルールを記述したリランによって左に立つフロアであり、4つの明示的なコールが1/10から9/10までの8ビット11Bと、ゴールド値のフォーマットのみの失敗の後0/10から7/10までの7Bを持ち上げ、マルチプライオリティ補正に耐えられない8つの異なるベースライン入力を持つ探索効果が許される。
(4)ポーランド語と英語のギャップは、劣化またはタスクファミリーと関連している。
結果が厳密で修正された形で報告され、ベンチマーク、トラジェクトリ、コミットスタンプされたアーティファクトがリリースされます。
関連論文リスト
- Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge [63.752287291668786]
テスト時間計算のスケールは、言語モデル推論を改善する強力な方法である。
自己閉じ込めは、既に現在の状態から到達可能な解に確率質量を大半を凝縮することを示す。
選択的なクエリフレームワークであるFlyByを導入し、まず4Bと8Bの変種をトレーニングする。
論文 参考訳(メタデータ) (2026-09-28T05:07:24Z) - On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability [57.03393882933515]
Qwen3.8-Flash-Nextは125Bパラメータ、トークンあたり6Bアクティベートされ、加速器から保持されるn-gram埋め込みテーブルの51Bパラメータを持つスパースミックス・オブ・エキスパートモデルである。
トレーニング前の14のベンチマークでは、このモデルは前モデルの397B-A17Bを8点、残りを少なくとも2.6ポイント、アクティベートされたパラメータが1/3、トレーニングトークンが1/3、トレーニング用FLOPが1/9であった。
論文 参考訳(メタデータ) (2026-08-31T06:35:07Z) - MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models [49.47300047353726]
MemToCは、実行時ツールによるポストツール-リターン仲裁のベンチマークである。
MemToCは、542の質制御された事実質問から構築された6,504回の評価エピソードで構成されている。
オープンウェイトな7-9Bモデルでは、ツールが強く支配的なクローズドブックの回答を返す。
論文 参考訳(メタデータ) (2026-08-26T18:22:03Z) - Joint Optimization of Tool Creation and Use for Large Language Model Agents [69.07066297088187]
SMITH (-grounded Multi-task Iterative Tool Honing) は、単一のポリシー内でツールの作成とツールの使用を訓練する。
SMITHで訓練された4B Qwen3は、正確に検証された13の手続き的推論タスクで、保持されたタスクで平均79.8のマクロ平均精度に達する。
また、TabMWP-Hardでは40.4、ドメイン外のGQAでは42.6に到達している。
論文 参考訳(メタデータ) (2026-08-25T13:59:34Z) - Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs [0.23438564092609357]
プロアクティブ干渉(英: Proactive interference, PI)は、大規模言語モデルにおける文書化された障害モードであり、繰り返し上書きされた値の検索は、事前上書きが蓄積されるにつれて劣化する。
PTQは現在、オープンウェイトモデルのデフォルトのデプロイメントパスとなっているが、この障害モードへの影響はテストされていない。
アーキテクチャ的に異なる3つの命令調整モデルの3つの精度レベル(FP16, INT8, INT4/NF4, ビット/バイト)を評価する。
論文 参考訳(メタデータ) (2026-08-19T06:17:13Z) - Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents [22.30468215041597]
ツール使用エージェントは別の言語で同じタスクを与えられるが、それでも同じステップを踏むだろうか?
アクションポリシーを,8つのモデル,6つの並列ベンチマーク,41の言語で測定対象とする。
論文 参考訳(メタデータ) (2026-08-11T16:18:34Z) - Wrong Before Right: Late Rescue and Interface Failure in Aligned Language Models [2.1291393112295935]
コーディネート言語モデル内での正確性について検討する。
極性制御された最小対上の層間差差分法(DiD)トラジェクトリを用いて、誤差を同定する。
このことは17モデル、3ファミリー、64倍スケール(0.5B-32B)のパッチスコープ型活性化移植による検証である。
論文 参考訳(メタデータ) (2026-07-06T03:51:24Z) - Automated Proving of Shannon-Type Entropy Inequalities via Fine-Tuned Language Models and Guided Tree Search [50.16356451328644]
シャノン型エントロピーの不等式を証明することは情報理論の基本的な課題である。
我々は,原子実証のステップを微調整した小規模大規模言語モデルがこのプロセスを自動化することができるか検討する。
GPT-5.5は0ショットプロンプトで1.7%のサンプルを解き、Psitipは33.3%のサンプルを解いた。
論文 参考訳(メタデータ) (2026-06-04T05:43:12Z) - Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning [51.88950852117154]
Chunk-Level Guided Generationは、既製の大規模言語モデルをプロセススコアラとして使用する、トレーニング不要の代替手段である。
本研究では,系統的な長さバイアスのため,大モデル確率の可変長推論ステップが信頼できないことを示す。
Chunk-Level Guided Generation は PRM guided search よりもかなり短い推論トレースを生成する。
論文 参考訳(メタデータ) (2026-06-01T04:43:36Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。