20 results for “Basic knowledge of data analysis techniques”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper introduces 'dashi,' an open-source Python library that provides comprehensive tools for characterizing dataset shifts (covariate, prior, concept) to ensure robust and trustworthy AI developm…
Sjoerd Vink, Suyang Li, Brian Montambault, Michael Behrisch +2 more
This paper introduces ZipLine, a system for integrative analysis of multivariate graphs through a unified predicate language and learning algorithm.
This study investigates how industrial practitioners perceive and manage data leakage in automotive perception systems, finding that leakage control is a socio-technical coordination problem requiring…
KoAT is a tool that automatically infers complexity bounds and proves termination of integer programs using an alternating modular analysis approach and a portfolio of techniques.
The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.
This paper proposes statistical procedures to identify coordinates responsible for change-points in multivariate time series data.
This paper extends gradient boosting to functions of vector inputs using a simple algorithm with histogram-based decision trees.
This paper provides non-asymptotic and explicit estimates for the exponential deviation inequalities of the HyperLogLog estimator.
This paper comparatively analyzes two automatic label error detection methods, Confident Learning and Dataset Cartography, demonstrating that targeted data filtering significantly improves model perfo…
This paper introduces a new benchmark dataset and evaluation framework for 'data snapshot extraction,' focusing on identifying and localizing semantically meaningful analytical artifacts within operat…
The twoblock clustering tree is introduced as a new type of interpretable regression tree for multivariate responses, using local multivariate linear models as leaves and twoblock dimension reduction…
The paper introduces Synthesis Data Reversion (SDR), a method that infers the data laundering transformation used in LLM training and synthesizes queries to restore the detection signals lost when pro…
This paper compares different approaches for embedding source code in machine learning models for learning analytics in programming education, specifically for the visual block-based programming langu…
This paper introduces structured cut sets, a novel preprocessing technique for Maximum k-Cut, and extends existing techniques from Maximum Cut. The rules are optimality-preserving and yield significan…
This paper explores the use of Large Language Models (LLMs) in data fusion tasks for tabular data and shows their superiority over traditional methods.
This paper presents an improved parallel version of MaxFEM algorithm for maximal frequent episode mining, achieving up to 8x speedup in C++ implementation and 35x improvement overall.
The paper proposes a conservative extension of Lustre's clock calculus to facilitate the embedding of ML models in reactive applications.
This paper presents a pipeline for transforming historical sources into structured data using machine learning tools and the GRAM-framework, enabling automated, skeletal graphing of actions.