ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “data pipelines”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.IREmpiricalRecentJul 1, 2026

Trie-based Experiment Plans for Efficient IR Pipeline Experiments

Irene Anu, Craig Macdonald

This paper describes the use of a trie data structure to enhance experiment efficiency in comparative pipeline experiments for cascading retrieval pipelines using PyTerrier, observing a 26% reduction…

View →
cs.CRRecentMar 30, 2026

Attesting LLM Pipelines: Enforcing Verifiable Training and Release Claims

Zhuoran Tan, Jeremy Singer, Christos Anagnostopoulos

The paper proposes an attestation-aware promotion gate to mitigate supply-chain risks in LLM pipelines by cryptographically verifying and enforcing claims about training and release artifacts before d…

View →
cs.CLcs.AIEmpiricalRecentJul 27, 2026

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

Zhen Huang, Yikun Wang, Shijie Xia, Pengfei Liu

The paper introduces DataOrchestra, a framework for example-specific processing of pretraining data for Large Language Models, achieving stable gains and reducing compute.

View →
cs.DBcs.AIRecentMay 29, 2026

Sophrosyne: Agentic Exploration of Relational Data Systems Needs Moderation

Madhav Jivrajani, Ramnatthan Alagappan, Aishwarya Ganesan

The paper introduces Sophrosyne, a system that moderates LLM agent exploration in relational data systems, significantly reducing over-exploration and boosting SQL generation accuracy by guiding the a…

View →
cs.AIRecentMay 31, 2026

"Skill issues'': data-centric optimization of lakehouse agents

Nicole Rose Schneider, Davide Ghilardi, Giacomo Piccinini, Jacopo Tagliabue

The paper introduces a data-centric optimization pipeline to improve coding agents' ability to interact with a branching lakehouse, showing significant accuracy gains by treating agent evaluation as a…

View →
cs.CLcs.AIcs.DBEmpiricalRecentJul 12, 2026

The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their Interactions

Filip Klubicka, Vasudevan Nedumpozhimana, Sneha Rautmare, Bora Caglayan +2 more

This paper explores the integration of several extensions for Natural Language to SQL (NL2SQL) translation, including NatSQL representation, preprocessing, fine-tuning, and reranker model.

View →
cs.AIcs.DBRecentMay 27, 2026

A Query Engine for the Agents

Kenny Daniel

The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.

View →
cs.CRRecentMay 20, 2026

An Evidence-driven Protocol for Trustworthy CI Pipelines

Fernando Castillo, Eduardo Brito, Pille Pullonen-Raudvere, Sebastian Werner +1 more

The paper proposes an evidence-driven protocol combining Deterministic Build Systems and Trusted Execution Environments to provide cryptographically verifiable guarantees of software artifact integrit…

View →
cs.DBcs.DCEmpiricalRecentJun 12, 2026

Vivace: Exact Temporal OLAP over Interval Histories via Independent Serverless Execution

Woohyeok Park, Taeyoon Kim, Hyunjoon Kim, Kungyong Lee

This paper presents Vivace, a serverless system for exact temporal OLAP over interval histories, which addresses the issues of incomplete data and incorrect answers in serverless functions.

View →
cs.DCcs.MATheoreticalRecentJul 3, 2026

A Workflow-Aware Serving Layer for Agentic Applications

Jiayi Qian, Zishen Wan, Hanchen Yang, Chun Tao +2 more

The paper presents Dyserve, a workflow-aware serving layer for agentic AI applications that compiles per-node model and verifier choices into an integer linear program, allowing for efficient model se…

View →
cs.LGcs.AIRecentMay 29, 2026

dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment

David Fernández-Narro, Pablo Ferri, Ángel Sánchez-García, Juan M. García-Gómez +1 more

The paper introduces 'dashi,' an open-source Python library that provides comprehensive tools for characterizing dataset shifts (covariate, prior, concept) to ensure robust and trustworthy AI developm…

View →
cs.DBcs.AIcs.MAEmpiricalRecentJun 29, 2026

Experience Graphs: The Data Foundation for Self-Improving Agents

Gang Liao, Yujia He, Abdullah Ozturk, Zhouyang Li +21 more

This paper proposes Trellis, a data foundation that treats experience graphs from long-horizon agentic tasks as first-class, governed, queryable database state.

View →
cs.CYcs.CLRecentMay 29, 2026

Traceable by Design: An LLM Pipeline and Dashboard for EU Regulatory Consultation Analysis

Thales Bertaglia, Haoyang Gui, Catalina Goanta, Gerasimos Spanakis

The paper presents an end-to-end LLM pipeline and interactive dashboard designed to automatically extract and structure topics from massive volumes of regulatory consultation submissions, ensuring ful…

View →
cs.CLRecentMay 31, 2026

Benchmarking Local LLMs for Natural-Language-to-SQL Querying in Biopharmaceutical Manufacturing: An Empirical Benchmark on Consumer-Grade Hardware

Sagar Bhetwal, Rajan Bastakoti, Nirajan Acharya, Gaurav Kumar Gupta

This study benchmarks four local LLMs for natural-language-to-SQL querying in biopharma manufacturing, finding that general-purpose code-tuned models like Llama 3.1 8B and Qwen 2.5 Coder 7B outperform…

View →
cs.AIRecentMay 31, 2026

GovAI-Pipe: A Layered AI Governance Pipeline for Citizen-Facing AI in Turkey's e-Government Gateway

Ahmet Kaplan

The paper proposes GovAI-Pipe, a novel four-layer governance pipeline that operationalizes high-level AI policies (like the EU AI Act) into auditable, technical checkpoints for deploying AI in large-s…

View →
cs.CLcs.AIcs.IRRecentMay 28, 2026

Exploring Autonomous Agentic Data Engineering for Model Specialization

Yujie Luo, Xiangyuan Ru, Jingsheng Zheng, Jingjing Wang +9 more

The paper introduces Autonomous Agentic Data Engineering, demonstrating that LLMs can autonomously plan and optimize end-to-end data curation pipelines, leading to substantial performance gains in spe…

View →
cs.DCEmpiricalRecentJun 23, 2026

BiJuTy: An Interactive HPC-Aware Big Data Cluster Lifecycle Manager and Performance Assessment Utility for JupyterHub

Apurv Deepak Kulkarni, Jan Frenzel, Siavash Ghiasvand

BiJuTy is a user-friendly solution for executing complex big data processing workflows on high-performance computing systems within the Jupyter ecosystem.

View →
cs.DBcs.AIRecentMay 29, 2026

SpecDB: LLM-Generated Customized Databases via Feature-Oriented Decomposition

Yunkai Lou, Longbin Lai, Shunyang Li, Zhengping Qian +1 more

SpecDB is a novel system that uses LLMs to synthesize highly customized, purpose-built relational databases, achieving performance comparable to commercial systems while significantly reducing code si…

View →
cs.CLcs.AIEmpiricalRecentJul 7, 2026

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen

The paper introduces DataGovBench, a benchmark for evaluating Large Language Models in real-world data analysis scenarios, revealing significant performance gaps with state-of-the-art models.

View →