ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “large-scale linear statistical models”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

math.STcs.LGmath.PREmpiricalRecentJun 4, 2026

How abundant are good interpolators?

August Y. Chen, Ahmed El Alaoui

This paper establishes a large deviation principle for the generalization error of interpolating classifiers in the overparametrized regime.

View →
math.STstat.MEstat.MLTheoreticalRecentJun 9, 2026

Conformal Prediction for Dyadic Regression Under Complex Missingness

Robert Lunde, Minjie Yang, Elizaveta Levina, Ji Zhu

This paper develops a framework for conformal prediction in dyadic regression problems under complex missingness mechanisms.

View →
math.NAstat.MLNEWTheoreticalJul 28, 2026

Sequential Preconditioned Conjugate Gradient Method for Linear Statistical Models

Guan-Yu Chen, Dong-Yue Xie, Xi Yang, Zun-Hao Zheng

This paper proposes a randomized iterative method called Sequential Preconditioned Conjugate Gradient Method (SPCG) for large-scale linear statistical models, which significantly reduces computational…

View →
math.STcs.LGstat.MLTheoreticalRecentJul 27, 2026

The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression

Kevin Han Huang, Haoyu Ye, Somak Laha, Morgane Austern

This paper derives deterministic equivalents for the prediction risk of over-parameterized linear regression with degenerate covariance matrices and dependent covariates, and identifies the configurat…

View →
math.STmath.NAstat.MLTheoreticalRecentJun 15, 2026

Optimal Multiscale Learning of Linear Operators

Jiaheng Chen, Daniel Sanz-Alonso

This paper analyzes the statistical and computational limits of learning bounded linear operators between Sobolev spaces from noisy data, and constructs a finite-resolution blockwise least-squares est…

View →
cs.LGstat.MLEmpiricalRecentJul 22, 2026

Efficient Clustering with Provable Guardrails for LLM Inference at Scale

Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang +2 more

This paper proposes a two-stage clustering algorithm to ensure per-sample guardrails in LLM-based applications, reducing inference cost and latency.

View →
stat.MLcs.LGstat.MEEmpiricalRecentJul 26, 2026

Distributional Split Criteria for Random Forests: Extensions, Shrinkage, and the Robustness of Mean Splitting

Silas Koemen

This paper introduces Distributional Random Forests, which replace mean-based CART splitting with criteria that compare full conditional response distributions in candidate children. The authors syste…

View →
cs.LGcs.AIRecentMay 31, 2026

What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear Regression

Wendao Wu, Fangqing Zhang, Haihan Zhang, Cong Fang

This paper develops a unified spectral analysis framework to explain how knowledge transfer (KT) works across different machine learning regimes, such as Knowledge Distillation and Weak-to-Strong gene…

View →
math.STstat.MEstat.MLTheoreticalRecentJul 23, 2026

Optimal use of a black-box learner in semiparametric estimation

Yihong Gu

This paper proposes a novel estimator for the target linear coefficient in a partial linear model with black-box nuisance estimation and establishes its unimprovable error rate.

View →
cs.LGcs.AIstat.MLRecentMay 29, 2026

InfoAtlas: A Foundation Model for Zero-Shot Statistical Dependence Estimate

Zhengyang Hu, Yanzhi Chen, Hanxiang Ren, Qunsong Zeng +4 more

InfoAtlas is a foundation model that estimates statistical mutual information (MI) in a single forward pass, achieving state-of-the-art accuracy with a massive speedup compared to traditional iterativ…

View →
stat.MLcs.LGstat.MEEmpiricalRecentJul 2, 2026

Autorelevance function and other feature relevance measures for univariate time series

Julian Cardenas, Jamie Arjona, Pedro Delicado

The paper proposes methodologies to measure lag relevance in machine learning forecasting models using Ghost variables, Shapley values, and additive importance measures. It also introduces auto-releva…

View →
cs.LGstat.MLEmpiricalRecentJul 20, 2026

Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation

Siddharth Mishra-Sharma

The paper presents a framework for model selection and parameter estimation using large language models and neural simulation-based inference.

View →