論文の概要: Fine-Tuning Large Language Models for Quantum Reasoning
- arxiv url: http://arxiv.org/abs/2606.21974v1
- Date: Sat, 20 Jun 2026 10:06:29 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-25 23:26:56.110493
- Title: Fine-Tuning Large Language Models for Quantum Reasoning
- Title(参考訳): 量子推論のための微調整大言語モデル
- Authors: Katherine Ip, Casey R. Myers, Udaya Parampalli, James Quach, Peiyong Wang,
- Abstract要約: 大規模言語モデル(LLM)は、自然言語モデリングやテキスト生成以外の能力を示す。
それらの推論能力の最近の進歩は、複雑な科学的タスクにLLMを適用することへの関心を喚起している。
- 参考スコア(独自算出の注目度): 1.1417805445492082
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Large language models (LLMs) exhibit abilities beyond natural language modelling and text generation. Recent advances in their reasoning capabilities have spurred interest in applying LLMs to complex scientific tasks requiring deep domain expertise and sophisticated reasoning. Quantum computing, as a highly specialised field with significant knowledge barriers and hardware constraints, could greatly benefit from such advancements. However, a key open question that first must be answered is: How can we develop fine-tuning pipelines that instil genuine quantum reasoning in LLMs, rather than task-specific pattern matching? We study this question through quantum circuit simulation as a training objective, where the model must predict the measurement probability distribution resulting from a sequence of quantum gate operations. We propose and compare two fine-tuning pipelines: (1) Supervised Fine-Tuning (SFT) on explicit gate-by-gate state-vector simulation traces, and (2) a two-stage SFT+Group Relative Policy Optimisation (GRPO) approach that sequentially applies SFT followed by GRPO with verifiable rewards. Our findings show that SFT achieves near-perfect in-distribution and gate-count extrapolation accuracy, significantly outperforming both the base model and the GPT-OSS-120B baseline. SFT+GRPO trades some in-distribution precision for better generalisation to larger qubit systems that SFT alone cannot handle. Both pipelines significantly outperform the baselines, demonstrating that targeted fine-tuning on explicit reasoning traces is an effective strategy for advancing quantum reasoning in LLMs.
- Abstract(参考訳): 大規模言語モデル(LLM)は、自然言語モデリングやテキスト生成以外の能力を示す。
それらの推論能力の最近の進歩は、深いドメインの専門知識と洗練された推論を必要とする複雑な科学的タスクにLLMを適用することへの関心を喚起している。
量子コンピューティングは、重要な知識障壁とハードウェア制約を持つ高度に専門化された分野であり、そのような進歩の恩恵を受けることができる。
タスク固有のパターンマッチングではなく、LLMに真の量子推論を取り入れた微調整パイプラインをどのように開発すればよいのか?
本研究では,量子回路シミュレーションをトレーニング対象とし,量子ゲート演算の列から得られる測定確率分布の予測を行う。
本研究では,(1)ゲートバイゲート状態ベクトルシミュレーショントレース上でのSFT (Supervised Fine-Tuning) と(2)SFT+Group Relative Policy Optimisation (GRPO) の2段階のアプローチを提案する。
以上の結果から, SFTは, ほぼ完全な分配精度とゲート数外挿精度を達成し, ベースモデルとGPT-OSS-120Bベースラインの両方において有意に優れていた。
SFT+GRPOは、SFT単独では扱えないより大きな量子ビット系へのより良い一般化のために、いくつかの分配精度を交換する。
どちらのパイプラインもベースラインを著しく上回り、明示的な推論トレースに基づく微調整がLLMの量子推論を前進させる効果的な戦略であることを示す。
関連論文リスト
- Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning [20.515599491717442]
マルチモーダル推論モデル学習のためのtextbfMetis-RISE (textbfRL textbfSFT textbfEnhances) を提案する。
論文 参考訳(メタデータ) (2025-06-16T02:56:13Z) - Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections [65.36449542323277]
本稿では,Large Language Model (LLM) 後の学習において,SFT(Supervised Fine-Tuning) と優先学習を統合した理論フレームワークを提案する。
そこで本研究では,学習率の簡易かつ効果的な削減手法を提案する。
論文 参考訳(メタデータ) (2025-06-15T05:42:29Z) - RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models [53.571195477043496]
本稿では,RoSTE (Rotated Straight-Through-Estimator) というアルゴリズムを提案する。
RoSTEは、量子化を意識した微調整(QA-SFT)と適応的な回転戦略を組み合わせることで、アクティベーションアウトリーを減少させる。
その結果, 予測誤差は収束重みの量子化誤差と直接比例し, 最適化された回転構成により効果的に管理できることが判明した。
論文 参考訳(メタデータ) (2025-02-13T06:44:33Z) - Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process [19.986235452236272]
Supervised Fine-Tuning (SFT) と Preference Optimization (PO) は、言語モデル(LM)を事前学習後の人間の好みに合わせるための重要なプロセスである。
Intuitive Fine-Tuning (IFT)を導入し,SFTとPOをひとつのプロセスに統合する。
IFT は SFT やいくつかの典型的な PO メソッドと相容れないか、それ以上に優れている。
論文 参考訳(メタデータ) (2024-05-20T08:23:28Z) - Realization of arbitrary doubly-controlled quantum phase gates [62.997667081978825]
本稿では,最適化問題における短期量子優位性の提案に着想を得た高忠実度ゲートセットを提案する。
3つのトランペット四重項のコヒーレントな多レベル制御を編成することにより、自然な3量子ビット計算ベースで作用する決定論的連続角量子位相ゲートの族を合成する。
論文 参考訳(メタデータ) (2021-08-03T17:49:09Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。