Fugu-MT 論文翻訳(概要): Reasoning with Preference Constraints: A Benchmark for Language Models in Many-to-One Matching Markets

論文の概要: Reasoning with Preference Constraints: A Benchmark for Language Models in Many-to-One Matching Markets

arxiv url: http://arxiv.org/abs/2509.13131v1
Date: Tue, 16 Sep 2025 14:48:46 GMT
ステータス: 翻訳完了
システム内更新日: 2025-09-17 17:50:53.127922
Title: Reasoning with Preference Constraints: A Benchmark for Language Models in Many-to-One Matching Markets
Title（参考訳）: 嗜好制約による推論: 複数対1のマッチング市場における言語モデルのベンチマーク
Authors: Marylou Fauchard, Florian Carichon, Margarida Carvalho, Golnoosh Farnadi,
Abstract要約: 大規模言語モデル (LLM) は、最適化を含む複雑な数学的タスクにおいて強い性能を示している。優先的かつ構造的な制約の下で推論を必要とする問題にLLMを適用することは、まだ未定である。我々は,大学入学問題の369件の新たなベンチマークを用いて,実用性,安定性,最適性といった重要な次元にわたるLSMを評価する。
参考スコア（独自算出の注目度）: 13.111181135818184
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Recent advances in reasoning with large language models (LLMs) have demonstrated strong performance on complex mathematical tasks, including combinatorial optimization. Techniques such as Chain-of-Thought and In-Context Learning have further enhanced this capability, making LLMs both powerful and accessible tools for a wide range of users, including non-experts. However, applying LLMs to matching problems, which require reasoning under preferential and structural constraints, remains underexplored. To address this gap, we introduce a novel benchmark of 369 instances of the College Admission Problem, a canonical example of a matching problem with preferences, to evaluate LLMs across key dimensions: feasibility, stability, and optimality. We employ this benchmark to assess the performance of several open-weight LLMs. Our results first reveal that while LLMs can satisfy certain constraints, they struggle to meet all evaluation criteria consistently. They also show that reasoning LLMs, like QwQ and GPT-oss, significantly outperform traditional models such as Llama, Qwen or Mistral, defined here as models used without any dedicated reasoning mechanisms. Moreover, we observed that LLMs reacted differently to the various prompting strategies tested, which include Chain-of-Thought, In-Context Learning and role-based prompting, with no prompt consistently offering the best performance. Finally, we report the performances from iterative prompting with auto-generated feedback and show that they are not monotonic; they can peak early and then significantly decline in later attempts. Overall, this work offers a new perspective on model reasoning performance and the effectiveness of prompting strategies in combinatorial optimization problems with preferential constraints.
Abstract（参考訳）: 大規模言語モデル(LLM)を用いた推論の最近の進歩は、組合せ最適化を含む複雑な数学的タスクにおいて強い性能を示している。 Chain-of-ThoughtやIn-Context Learningといったテクニックにより、この機能がさらに強化され、LLMは、非専門家を含む幅広いユーザに対して、強力でアクセスしやすいツールとなる。しかし、優先的かつ構造的な制約の下での推論を必要とするマッチング問題にLLMを適用することは、まだ未定である。このギャップに対処するため,大学入学問題(College Admission Problem)の369件の新たなベンチマークを導入する。オープンウェイトLLMの性能評価には,このベンチマークを用いている。その結果, LLMは一定の制約を満たすことができるが, 全ての評価基準を一貫して満たすのに苦慮していることが明らかとなった。また、QwQ や GPT-oss のような推論 LLM は、Llama や Qwen や Mistral といった従来のモデルよりも大幅に優れており、特別な推論機構を持たないモデルとしてここで定義されている。さらに,LLMは,Chain-of-Thought,In-Context Learning,ロールベースのプロンプトなど,テスト対象のプロンプト戦略と異なる反応を示した。最後に、自動生成フィードバックによる反復的フィードバックによるパフォーマンスを報告し、モノトニックではないことを示す。全体として、本研究はモデル推論性能の新しい視点と、優先的な制約を伴う組合せ最適化問題における戦略の促進効果を提供する。

論文の概要: Reasoning with Preference Constraints: A Benchmark for Language Models in Many-to-One Matching Markets

関連論文リスト