ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “tabular data”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CRcs.DBRecentJul 24, 2026

trasgoDP: An Open Source Framework for Releasing Noised Tabular Microdata under Local Differential Privacy

Judith Sáinz-Pardo Díaz, Álvaro López García

trasgoDP is an open-source Python framework for releasing tabular and location data under local differential privacy and geo-indistinguishability guarantees.

View →
cs.IRcs.AIcs.CLRecentJun 1, 2026

ODTQA-FoRe: An Open-Domain Tabular Question Answering Dataset for Future Data Forecasting and Reasoning

Zhensheng Wang, Xiaole Liu, Wenmian Yang, Kun Zhou +2 more

The paper introduces Open-Domain Tabular Question Answering for Future Data Forecasting and Reasoning, a new dataset and framework that enables LLMs to perform time-series forecasting and reasoning on…

View →
cs.LGcs.AIRecentMay 30, 2026

TabChange: Precise Attribute Changes in Tabular Data

Arjun Dahal, Yu Lei, Raghu N. Kacker, Richard Kuhn

TabChange proposes a novel framework to generate natural and minimally altered counterfactual instances in tabular data by precisely controlling attribute modifications based on their relationship str…

View →
cs.LGstat.MLEmpiricalRecentJul 17, 2026

Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data

Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung +4 more

This paper proposes a taxonomy-guided evaluation protocol for temporal fidelity in synthetic sequential tabular data, measuring timestamp validity, cross-sectional structure, within-entity dynamics, a…

View →
cs.DBcs.AIcs.CLEmpiricalRecentJun 26, 2026

Single and Multi Truth Data Fusion using Large Language Models

Hira Beril Kucuk, Norman W Paton, Jiaoyan Chen, Zhenyu Wu

This paper explores the use of Large Language Models (LLMs) in data fusion tasks for tabular data and shows their superiority over traditional methods.

View →
cs.LGEmpiricalRecentJul 6, 2026

TabPack: Efficient Hyperparameter Ensembles for Tabular Deep Learning

Yury Gorishniy, Akim Kotelnikov, Ivan Rubachev, Artem Babenko

This paper introduces TabPack, an efficient MLP ensemble for tabular data that samples and trains MLPs with different hyperparameters in parallel and selects ensemble members on-the-fly during trainin…

View →
cs.LGRecentJun 1, 2026

TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks

Andrej Tschalzev, Nick Erickson, Yuyang Wang, Huzefa Rangwala +3 more

The paper introduces TabPrep, a feature engineering pipeline that systematically improves performance across various tabular machine learning models by addressing structural data patterns ignored by c…

View →
cs.CLcs.AIEmpiricalRecentJun 30, 2026

When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors

Yuqing Yang, Qi Zhu, Zhen Han, Boran Han +4 more

This paper systematically evaluates tabular data referencing errors in large language models and presents methods to improve answer accuracy and detect errors.

View →
cs.CLcs.AIEmpiricalRecentJul 7, 2026

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen

The paper introduces DataGovBench, a benchmark for evaluating Large Language Models in real-world data analysis scenarios, revealing significant performance gaps with state-of-the-art models.

View →
cs.CLRecentMay 29, 2026

Semantic Triplet Restoration: A Novel Protocol for Hierarchical Table Understanding in Large Language Models

Yibin Zhao, Fangxin Shang, Dingrui Yang, Yuqi Wang

The paper introduces Semantic Triplet Restoration (STR), a novel protocol that converts complex table structures into atomic semantic triplets, improving table question answering by providing explicit…

View →
cs.CLcs.AIcs.CVRecentJun 4, 2026

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

AJ Carl P. Dy, Aivin V. Solatorio

This paper introduces a new benchmark dataset and evaluation framework for 'data snapshot extraction,' focusing on identifying and localizing semantically meaningful analytical artifacts within operat…

View →
cs.LGcs.AIcs.CRRecentMay 23, 2026

Geometry-Aware Tabular Diffusion

David Turtora Zagardo

The paper introduces Geometry-Aware Tabular Diffusion (GATD), a method that enhances tabular data synthesis by explicitly incorporating pairwise geometric relationships (angles and lengths) into the d…

View →
cs.HCcs.AIEmpiricalRecentJun 29, 2026

Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework

Yuchen He, Peizhi Ying, Liqi Cheng, Kuilin Peng +3 more

The paper builds a benchmark to evaluate the ability of multimodal large language models to extract accurate data tables from chart images, and proposes a human-centered approach to improve numerical…

View →
cs.LGcs.AIstat.MERecentMay 28, 2026

The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction

Shu Wan, Abhinav Gorantla, Huan Liu, K. Selçuk Candan

While restricting a model to the theoretical Markov boundary can significantly improve prediction, the practical process of discovering and using this boundary is often computationally infeasible and…

View →
cs.DBcs.DCEmpiricalRecentJun 12, 2026

Vivace: Exact Temporal OLAP over Interval Histories via Independent Serverless Execution

Woohyeok Park, Taeyoon Kim, Hyunjoon Kim, Kungyong Lee

This paper presents Vivace, a serverless system for exact temporal OLAP over interval histories, which addresses the issues of incomplete data and incorrect answers in serverless functions.

View →
cs.AIcs.CLRecentMay 28, 2026

Demystifying Data Organization for Enhanced LLM Training

Yalun Dai, Yangyu Huang, Tongshen Yang, Yonghan Wang +7 more

This paper proposes four guidelines and two novel data ordering methods (STR and SAW) to systematically optimize data organization, significantly enhancing the stability and performance of LLM trainin…

View →
cs.AIcs.DBRecentMay 27, 2026

A Query Engine for the Agents

Kenny Daniel

The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.

View →
cs.CLRecentJun 1, 2026

Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification

Sunisth Kumar, Xanh Ho, Tim Schopf, Andre Greiner-Petter +2 more

The paper explains the 'table-chart gap' in scientific claim verification by showing that multimodal LLMs successfully encode information from charts but fail to route it to the final prediction layer…

View →
cs.CLcs.AIcs.DBEmpiricalRecentJul 12, 2026

The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their Interactions

Filip Klubicka, Vasudevan Nedumpozhimana, Sneha Rautmare, Bora Caglayan +2 more

This paper explores the integration of several extensions for Natural Language to SQL (NL2SQL) translation, including NatSQL representation, preprocessing, fine-tuning, and reranker model.

View →