Related papers: ASP-Bench: From Natural Language to Logic Programs

ASP-Bench: From Natural Language to Logic Programs

URL: http://arxiv.org/abs/2602.01171v1
Date: Sun, 01 Feb 2026 11:48:36 GMT
Title: ASP-Bench: From Natural Language to Logic Programs
Authors: Stefan Szeider,
Abstract summary: We present a benchmark comprising 128 natural language problem instances, 64 base problems with easy and hard variants.<n>It evaluates systems that translate natural-language problems into Answer Set Programs (ASPs), a prominent form of logic programming.
Score: 27.126691338850254
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Automating the translation of natural-language specifications into logic programs is a challenging task that affects neurosymbolic engineering. We present ASP-Bench, a benchmark comprising 128 natural language problem instances, 64 base problems with easy and hard variants. It evaluates systems that translate natural-language problems into Answer Set Programs (ASPs), a prominent form of logic programming. It provides systematic coverage of ASP features, including choice rules, aggregates, and optimization. Each problem includes reference validators that check whether solutions satisfy the problem specification. We characterize problems along seven largely independent reasoning aspects (optimization, temporal reasoning, default logic, resource allocation, recursion, spatial reasoning, and quantitative complexity), providing a multidimensional view of modeling difficulty. We test the benchmark using an agentic approach based on the ReAct (Reason and Act) framework, which achieves full saturation, demonstrating that feedback-driven iterative refinement with solver feedback provides a reliable and robust approach for modeling natural language in ASP. Our analysis across multiple agent runs enables us to gain insights into what determines a problem's modeling hardness.

Related papers

Bridging Natural Language and ASP: A Hybrid Approach Using LLMs and AMR Parsing [0.14658400971135646]
This paper proposes a novel method of translating unconstrained English into ASP programs for logic puzzles.<n>Everything from ASP rules, facts, and constraints is generated to fully represent and solve the desired problem.
arXiv Detail & Related papers (2025-11-11T19:25:44Z)
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models? [59.970391602080205]
Despite multilingual training, LRMs tend to default to reasoning in high-resource languages at test time.<n>Cultural reasoning degrades performance on reasoning tasks but benefits cultural tasks, while safety evaluations exhibit language-specific behavior.
arXiv Detail & Related papers (2025-05-23T02:46:18Z)
Logic-of-Thought: Empowering Large Language Models with Logic Programs for Solving Puzzles in Natural Language [67.51318974970985]
Solving puzzles in natural language poses a long-standing challenge in AI.<n>We propose Logic-of-Thought, a framework that bridges large language models with logic programming.<n>We evaluate our method on various grid puzzles and dynamic puzzles involving actions, demonstrating near-perfect accuracy across all tasks.
arXiv Detail & Related papers (2025-05-22T01:37:40Z)
Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique [66.94905631175209]
We propose a novel inference-time scaling approach -- stepwise natural language self-critique (PANEL)<n>It employs self-generated natural language critiques as feedback to guide the step-level search process.<n>This approach bypasses the need for task-specific verifiers and the associated training overhead.
arXiv Detail & Related papers (2025-03-21T17:59:55Z)
Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization [49.362750475706235]
Reinforcement Learning (RL) plays a crucial role in aligning large language models with human preferences and improving their ability to perform complex tasks.<n>We introduce Direct Q-function Optimization (DQO), which formulates the response generation process as a Markov Decision Process (MDP) and utilizes the soft actor-critic (SAC) framework to optimize a Q-function directly parameterized by the language model.<n> Experimental results on two math problem-solving datasets, GSM8K and MATH, demonstrate that DQO outperforms previous methods, establishing it as a promising offline reinforcement learning approach for aligning language models.
arXiv Detail & Related papers (2024-10-11T23:29:20Z)
Quantifying over Optimum Answer Sets [6.390468088226495]
ASP(Q) lacks a method for encoding modeling in an elegant and compact way. We propose an extension of ASP(Q) in which component programs may contain weak constraints. We showcase the modeling capabilities of the new formalism through various application scenarios.
arXiv Detail & Related papers (2024-08-14T17:53:13Z)
Static Analysis of Logic Programs via Boolean Networks [0.0]
"What can be said about stable models of a logic program from its static information?" has been investigated and proved useful in many circumstances. The proposed connection can bring the existing results in the rich history on static analysis of Boolean networks to explore and prove more theoretical results on ASP. The newly obtained insights have the potential to benefit many problems in the field of ASP.
arXiv Detail & Related papers (2024-07-12T06:07:05Z)
Language Models can be Logical Solvers [99.40649402395725]
We introduce LoGiPT, a novel language model that directly emulates the reasoning processes of logical solvers. LoGiPT is fine-tuned on a newly constructed instruction-tuning dataset derived from revealing and refining the invisible reasoning process of deductive solvers.
arXiv Detail & Related papers (2023-11-10T16:23:50Z)
Large Language Models as Analogical Reasoners [155.9617224350088]
Chain-of-thought (CoT) prompting for language models demonstrates impressive performance across reasoning tasks. We introduce a new prompting approach, analogical prompting, designed to automatically guide the reasoning process of large language models.
arXiv Detail & Related papers (2023-10-03T00:57:26Z)
Leveraging Large Language Models to Generate Answer Set Programs [5.532477732693001]
Large language models (LLMs) have demonstrated exceptional performance in various natural language processing tasks. This paper proposes a neuro-symbolic method that combines the strengths of large language models and answer set programming.
arXiv Detail & Related papers (2023-07-15T03:40:55Z)
Coupling Large Language Models with Logic Programming for Robust and General Reasoning from Text [5.532477732693001]
We show that a large language model can serve as a highly effective few-shot semantically. It can convert natural language sentences into a logical form that serves as input for answer set programs. We demonstrate that this method achieves state-of-the-art performance on several benchmarks, including bAbI, StepGame, CLUTRR, and gSCAN.
arXiv Detail & Related papers (2023-07-15T03:29:59Z)

This list is automatically generated from the titles and abstracts of the papers in this site.