Related papers: Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples

Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples

URL: http://arxiv.org/abs/2307.14565v2
Date: Wed, 9 Aug 2023 04:53:52 GMT
Title: Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples
Authors: Peng Li, Yeye He, Cong Yan, Yue Wang, Surajit Chaudhuri
Abstract summary: Auto-Tables can automatically transform non-relational tables into standard relational forms for downstream analytics. Our evaluation suggests that Auto-Tables can successfully synthesize transformations for over 70% of test cases at interactive speeds.
Score: 24.208275772387683
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Relational tables, where each row corresponds to an entity and each column corresponds to an attribute, have been the standard for tables in relational databases. However, such a standard cannot be taken for granted when dealing with tables "in the wild". Our survey of real spreadsheet-tables and web-tables shows that over 30% of such tables do not conform to the relational standard, for which complex table-restructuring transformations are needed before these tables can be queried easily using SQL-based analytics tools. Unfortunately, the required transformations are non-trivial to program, which has become a substantial pain point for technical and non-technical users alike, as evidenced by large numbers of forum questions in places like StackOverflow and Excel/Power-BI/Tableau forums. We develop an Auto-Tables system that can automatically synthesize pipelines with multi-step transformations (in Python or other languages), to transform non-relational tables into standard relational forms for downstream analytics, obviating the need for users to manually program transformations. We compile an extensive benchmark for this new task, by collecting 244 real test cases from user spreadsheets and online forums. Our evaluation suggests that Auto-Tables can successfully synthesize transformations for over 70% of test cases at interactive speeds, without requiring any input from users, making this an effective tool for both technical and non-technical users to prepare data for analytics.

Related papers

HCT-QA: A Benchmark for Question Answering on Human-Centric Tables [2.8913928436305505]
Human-centric tables (HCTs) possess a unique combination of high business value, intricate layouts, limited operational power at scale, and sometimes serve as the only data source for critical insights.<n>This paper describes HCT-QA, an extensive benchmark of HCTs, natural language queries, and related answers on thousands of tables.
arXiv Detail & Related papers (2025-03-09T11:02:11Z)
Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval [52.592071689901196]
We introduce a method that uncovers useful join relations for any query and database during table retrieval. Our method outperforms the state-of-the-art approaches for table retrieval by up to 9.3% in F1 score and for end-to-end QA by up to 5.4% in accuracy.
arXiv Detail & Related papers (2024-04-15T15:55:01Z)
WikiTableEdit: A Benchmark for Table Editing by Natural Language Instruction [56.196512595940334]
This paper investigates the performance of Large Language Models (LLMs) in the context of table editing tasks. We leverage 26,531 tables from the Wiki dataset to generate natural language instructions for six distinct basic operations. We evaluate several representative large language models on the WikiTableEdit dataset to demonstrate the challenge of this task.
arXiv Detail & Related papers (2024-03-05T13:33:12Z)
Retrieval-Based Transformer for Table Augmentation [14.460363647772745]
We introduce a novel approach toward automatic data wrangling. We aim to address table augmentation tasks, including row/column population and data imputation. Our model consistently and substantially outperforms both supervised statistical methods and the current state-of-the-art transformer-based models.
arXiv Detail & Related papers (2023-06-20T18:51:21Z)
MultiTabQA: Generating Tabular Answers for Multi-Table Question Answering [61.48881995121938]
Real-world queries are complex in nature, often over multiple tables in a relational database or web page. Our model, MultiTabQA, not only answers questions over multiple tables, but also generalizes to generate tabular answers.
arXiv Detail & Related papers (2023-05-22T08:25:15Z)
Generate, Transform, Answer: Question Specific Tool Synthesis for Tabular Data [6.3455238301221675]
Tabular question answering (TQA) presents a challenging setting for neural systems. TQA process tables directly, resulting in information loss as table size increases. We propose ToolWriter to generate query specific programs and detect when to apply them to transform tables.
arXiv Detail & Related papers (2023-03-17T17:26:56Z)
TRUST: An Accurate and End-to-End Table structure Recognizer Using Splitting-based Transformers [56.56591337457137]
We propose an accurate and end-to-end transformer-based table structure recognition method, referred to as TRUST. Transformers are suitable for table structure recognition because of their global computations, perfect memory, and parallel computation. We conduct experiments on several popular benchmarks including PubTabNet and SynthTable, our method achieves new state-of-the-art results.
arXiv Detail & Related papers (2022-08-31T08:33:36Z)
OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering [106.73213656603453]
We develop a simple table-based QA model with minimal annotation effort. We propose an omnivorous pretraining approach that consumes both natural and synthetic data.
arXiv Detail & Related papers (2022-07-08T01:23:45Z)
Table Retrieval May Not Necessitate Table-specific Model Design [83.27735758203089]
We focus on the task of table retrieval, and ask: "is table-specific model design necessary for table retrieval?" Based on an analysis on a table-based portion of the Natural Questions dataset (NQ-table), we find that structure plays a negligible role in more than 70% of the cases. We then experiment with three modules to explicitly encode table structures, namely auxiliary row/column embeddings, hard attention masks, and soft relation-based attention biases. None of these yielded significant improvements, suggesting that table-specific model design may not be necessary for table retrieval.
arXiv Detail & Related papers (2022-05-19T20:35:23Z)
MATE: Multi-view Attention for Table Transformer Efficiency [21.547074431324024]
More than 20% of relational tables on the web have 20 or more rows. Current Transformer models are typically limited to 512 tokens. We propose MATE, a novel Transformer architecture designed to model the structure of web tables.
arXiv Detail & Related papers (2021-09-09T14:39:30Z)
Capturing Row and Column Semantics in Transformer Based Question Answering over Tables [9.347393642549806]
We show that one can achieve superior performance on table QA task without using any of these specialized pre-training techniques. Experiments on recent benchmarks prove that the proposed methods can effectively locate cell values on tables (up to 98% Hit@1 accuracy on Wiki lookup questions)
arXiv Detail & Related papers (2021-04-16T18:22:30Z)

This list is automatically generated from the titles and abstracts of the papers in this site.