Related papers: STARQA: A Question Answering Dataset for Complex Analytical Reasoning over Structured Databases

STARQA: A Question Answering Dataset for Complex Analytical Reasoning over Structured Databases

URL: http://arxiv.org/abs/2509.19508v1
Date: Tue, 23 Sep 2025 19:26:16 GMT
Title: STARQA: A Question Answering Dataset for Complex Analytical Reasoning over Structured Databases
Authors: Mounica Maddela, Lingjue Xie, Daniel Preotiuc-Pietro, Mausam,
Abstract summary: We introduce STARQA, the first public human-created dataset of complex analytical reasoning questions and answers on three specialized relational-domain databases.<n>In this paper, we introduce STARQA, the first public human-created dataset of complex analytical reasoning questions and answers on three specialized relational-domain databases.
Score: 27.66819120859756
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Abstract: Semantic parsing methods for converting text to SQL queries enable question answering over structured data and can greatly benefit analysts who routinely perform complex analytics on vast data stored in specialized relational databases. Although several benchmarks measure the abilities of text to SQL, the complexity of their questions is inherently limited by the level of expressiveness in query languages and none focus explicitly on questions involving complex analytical reasoning which require operations such as calculations over aggregate analytics, time series analysis or scenario understanding. In this paper, we introduce STARQA, the first public human-created dataset of complex analytical reasoning questions and answers on three specialized-domain databases. In addition to generating SQL directly using LLMs, we evaluate a novel approach (Text2SQLCode) that decomposes the task into a combination of SQL and Python: SQL is responsible for data fetching, and Python more naturally performs reasoning. Our results demonstrate that identifying and combining the abilities of SQL and Python is beneficial compared to using SQL alone, yet the dataset still remains quite challenging for the existing state-of-the-art LLMs.

Related papers

From Queries to Insights: Agentic LLM Pipelines for Spatio-Temporal Text-to-SQL [8.496933324334167]
We present a naive text-to-Act baseline (Rellama-sqlcoder-8b) with orchestration by a Mistral-based Rellama-sqlcoder-8b.<n>We evaluate on 35 natural-language queries over the NYC and Tokyo check-in, covering spatial, temporal multi-dataset reasoning.<n>The agent achieves substantially higher accuracy than the dataset 91.4% vs. 28.6% and enhances usability through maps, and plots structured natural-language summaries.
arXiv Detail & Related papers (2025-10-29T22:18:57Z)
Text to Query Plans for Question Answering on Large Tables [4.917892629916144]
We propose a novel framework that transforms natural language queries into query plans.<n>We enable complex analytical functions, such as principal component analysis and anomaly detection.<n>We validate our framework through experiments on both standard databases and large scientific tables.
arXiv Detail & Related papers (2025-08-26T07:35:26Z)
Text2VectorSQL: Towards a Unified Interface for Vector Search and SQL Queries [36.92547259037192]
The proliferation of unstructured data poses a fundamental challenge to traditional database infrastructure.<n>While Text-to-BIRD has democratized access to structured data, it remains incapable of interpreting semantic or multi-modal queries.<n>We introduce and formalize Text2 Vector, a novel task to establish a unified natural language for seamlessly querying both structured and unstructured data.
arXiv Detail & Related papers (2025-06-29T03:17:42Z)
Weaver: Interweaving SQL and LLM for Table Reasoning [62.55797244714265]
Weaver generates a flexible, step-by-step plan that combinessql for structured data retrieval with LLMs for semantic processing.<n>Weaver consistently outperforms state-of-the-art methods across four TableQA datasets.
arXiv Detail & Related papers (2025-05-25T03:27:37Z)
LogicCat: A Chain-of-Thought Text-to-SQL Benchmark for Complex Reasoning [12.249447967086828]
LogicCat is the first Text-to-sense benchmark dataset specifically designed for complex reasoning and chain-of-thought parsing.<n>We show that LogicCat substantially increases the task difficulty for current state-of-the-art models to 33.20% execution accuracy.
arXiv Detail & Related papers (2025-05-24T15:23:43Z)
E-SQL: Direct Schema Linking via Question Enrichment in Text-to-SQL [1.187832944550453]
We introduce E-Seek, a novel pipeline specifically designed to address these challenges through direct schema linking and candidate predicate augmentation.<n>E-Seek enhances the natural language query by incorporating relevant database items (i.e., tables, columns, and values) and conditions directly into the question andsql construction plan, bridging the gap between the query and the database structure.<n> Comprehensive evaluations illustrate that E-Seek achieves competitive performance, particularly excelling in complex queries with a 66.29% execution accuracy on the test set.
arXiv Detail & Related papers (2024-09-25T09:02:48Z)
RoundTable: Leveraging Dynamic Schema and Contextual Autocomplete for Enhanced Query Precision in Tabular Question Answering [11.214912072391108]
Real-world datasets often feature a vast array of attributes and complex values. Traditional methods cannot fully relay the datasets size and complexity to the Large Language Models. We propose a novel framework that leverages Full-Text Search (FTS) on the input table.
arXiv Detail & Related papers (2024-08-22T13:13:06Z)
UQE: A Query Engine for Unstructured Databases [71.49289088592842]
We investigate the potential of Large Language Models to enable unstructured data analytics. We propose a new Universal Query Engine (UQE) that directly interrogates and draws insights from unstructured data collections.
arXiv Detail & Related papers (2024-06-23T06:58:55Z)
Semantic Decomposition of Question and SQL for Text-to-SQL Parsing [2.684900573255764]
We propose a new modular Query Plan Language (QPL) that systematically decomposessql queries into simple and regular sub-queries. Experimental results demonstrate that QPL is more effective than text-to-QPL for semantically equivalent queries.
arXiv Detail & Related papers (2023-10-20T15:13:34Z)
SQL-PaLM: Improved Large Language Model Adaptation for Text-to-SQL (extended) [53.95151604061761]
This paper introduces the framework for enhancing Text-to- filtering using large language models (LLMs) With few-shot prompting, we explore the effectiveness of consistency decoding with execution-based error analyses. With instruction fine-tuning, we delve deep in understanding the critical paradigms that influence the performance of tuned LLMs.
arXiv Detail & Related papers (2023-05-26T21:39:05Z)
Improving Text-to-SQL Semantic Parsing with Fine-grained Query Understanding [84.04706075621013]
We present a general-purpose, modular neural semantic parsing framework based on token-level fine-grained query understanding. Our framework consists of three modules: named entity recognizer (NER), neural entity linker (NEL) and neural entity linker (NSP)
arXiv Detail & Related papers (2022-09-28T21:00:30Z)
A Survey on Text-to-SQL Parsing: Concepts, Methods, and Future Directions [102.8606542189429]
The goal of text-to-corpora parsing is to convert a natural language (NL) question to its corresponding structured query language () based on the evidences provided by databases. Deep neural networks have significantly advanced this task by neural generation models, which automatically learn a mapping function from an input NL question to an output query.
arXiv Detail & Related papers (2022-08-29T14:24:13Z)
Weakly Supervised Text-to-SQL Parsing through Question Decomposition [53.22128541030441]
We take advantage of the recently proposed question meaning representation called QDMR. Given questions, their QDMR structures (annotated by non-experts or automatically predicted) and the answers, we are able to automatically synthesizesql queries. Our results show that the weakly supervised models perform competitively with those trained on NL- benchmark data.
arXiv Detail & Related papers (2021-12-12T20:02:42Z)
Dual Reader-Parser on Hybrid Textual and Tabular Evidence for Open Domain Question Answering [78.9863753810787]
A large amount of world's knowledge is stored in structured databases. query languages can answer questions that require complex reasoning, as well as offering full explainability.
arXiv Detail & Related papers (2021-08-05T22:04:13Z)

This list is automatically generated from the titles and abstracts of the papers in this site.