ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

~ similar to 2607.25272· 19 results

cs.LGcs.AIRecentMay 27, 2026

Learning Theory of the SVRG: Generalization and Convergence Analysis

Yunwen Lei, Zimeng Wang, Xiaoming Yuan

This paper provides the first non-vacuous generalization analysis for the Stochastic Variance Reduced Gradient (SVRG) method by establishing sharp, data-dependent algorithmic stability bounds, thereby…

View →
cs.LGmath.OCstat.MLTheoreticalRecentJul 20, 2026

Optimizing the Preconditioner: A Black-box Online-to-Nonconvex Conversion with Static Regret Minimization Oracles

Haichen Hu, David Simchi-Levi

This paper shows that stochastic nonconvex optimization can be reduced to ordinary static regret minimization in online convex optimization, and establishes convergence rates for smooth and Lipschitz…

View →
cs.AImath.OCRecentJun 1, 2026

Stochastic convergence of parallel asynchronous adaptive first-order methods

Serge Gratton, Philippe L. Toint

The paper analyzes a new class of asynchronous adaptive first-order optimization methods and proves their stochastic convergence rate is O(1/sqrt{t}) for non-convex functions.

View →
cs.ITTheoreticalRecentJun 26, 2026

Deriving Approximate Message Passing from the Convex Gaussian Min-Max Theorem

Vikrant Malik, Babak Hassibi

This paper establishes a direct connection between Approximate Message Passing (AMP) and the Convex Gaussian Min-max Theorem (CGMT) for regularized linear regression and M-estimation.

View →
cs.LGcs.AIEmpiricalRecentJun 4, 2026

PC Layer: Polynomial Weight Preconditioning for Improving LLM Pre-Training

Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more

This paper proposes a preconditioning layer for stable weight conditioning in LLM training.

View →
cs.LGcs.AIEmpiricalRecentJun 4, 2026

PC Layer: Polynomial Weight Preconditioning for Improving LLM Pre-Training

Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more

This paper proposes a preconditioning layer for stable weight conditioning in LLM training.

View →
cs.LGTheoreticalRecentJul 24, 2026

Complexity Bounds and Approaches to Learning Projected Gradient Descent Solver Iterates

Anjian Li, Ryne Beeson

This paper proposes a data collection strategy using solver iterates to augment datasets for training generative models, improving the efficiency of the data-model-optimization loop in one-sided box-c…

View →
cs.LGcs.AIRecentMay 27, 2026

Stochastic Gradient Descent with Momentum is Algorithmically Stable

Yunwen Lei, Zimeng Wang, Xiaoming Yuan

This paper provides a comprehensive generalization analysis of Stochastic Gradient Descent with Momentum (SGDM) by establishing tight, on-average model stability bounds that show SGDM can generalize w…

View →
cs.LGmath.OCstat.MLTheoreticalRecentJun 29, 2026

Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules

Muhammad Hamza, Ayush Goel

This paper introduces Curvature-Weighted Gradient Diversity (CWGD), a geometry-aware measure for optimization noise that reduces the asymptotic optimization error floor by up to a factor of two compar…

View →
math.NAcs.CEmath-phRecentMay 28, 2026

Multifidelity Proper Orthogonal Decomposition

Nicole Aretz, Karen Willcox

The paper introduces Multifidelity Proper Orthogonal Decomposition (MFPOD), a method that significantly reduces the computational cost of dimension reduction by intelligently combining data from cheap…

View →
eess.SPcs.LGEmpiricalRecentJun 22, 2026

Low-rank Updates in Slowly Time-varying Graphs for Spatial-Temporal Signal Interpolation

Saghar Bagheri, Gene Cheung, Tim Eadie, Antonio Ortega

This paper models the changes in graph adjacency matrices over time as a low-rank matrix and develops a method for jointly interpolating signals and estimating graph adjacency matrices using this assu…

View →
math.STstat.MEstat.MLTheoreticalRecentJul 23, 2026

Optimal use of a black-box learner in semiparametric estimation

Yihong Gu

This paper proposes a novel estimator for the target linear coefficient in a partial linear model with black-box nuisance estimation and establishes its unimprovable error rate.

View →
cs.CVRecentJun 1, 2026

VISReg: Variance-Invariance-Sketching Regularization for JEPA training

Haiyu Wu, Randall Balestriero, Morgan Levine

VISReg introduces a novel regularization technique that combines variance control with a Sliced-Wasserstein-based sketching objective to stabilize self-supervised learning, achieving state-of-the-art…

View →
cs.LGcs.AIRecentJun 1, 2026

FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo

Kyunghun Nam, Sumyeong Ahn

The paper proposes FOAM, an adaptive damping method that stabilizes the Shampoo optimization algorithm by dynamically controlling damping and eigendecomposition frequency, thereby reducing staleness-i…

View →
econ.EMmath.STstat.METheoreticalRecentJul 27, 2026

Debiased Machine Learning: Identification, Estimation, and Shape Constraints

Qihui Chen, Ka Yan Cheng, Zheng Fang

This paper develops a framework for identifying and estimating parameters of interest in automatic debiased machine learning using a Riesz representer, which is identified when it uniquely optimizes a…

View →
cs.LGcs.CEmath.NARecentMay 27, 2026

History-aware adaptive reduced-order models via incremental singular value decomposition

Amirpasha Hedayat, Ali Mohaghegh, Laura Balzano, Cheng Huang +1 more

The paper introduces a history-aware adaptive Reduced-Order Model (ROM) framework using incremental Singular Value Decomposition (iSVD) that maintains accuracy for online dynamics far beyond the initi…

View →
cs.DSTheoreticalRecentJul 9, 2026

Locally Approximating the Top Eigenvector of Bounded Entry Matrices

Nicolas Menand, Erik Waingarten

This paper presents a local computation algorithm to approximate the top eigenvector of a symmetric matrix with entries between -1 and 1, building on Swartworth and Woodruff's work.

View →