~ similar to 2607.25272· 19 results
This paper provides the first non-vacuous generalization analysis for the Stochastic Variance Reduced Gradient (SVRG) method by establishing sharp, data-dependent algorithmic stability bounds, thereby…
This paper shows that stochastic nonconvex optimization can be reduced to ordinary static regret minimization in online convex optimization, and establishes convergence rates for smooth and Lipschitz…
The paper analyzes a new class of asynchronous adaptive first-order optimization methods and proves their stochastic convergence rate is O(1/sqrt{t}) for non-convex functions.
This paper establishes a direct connection between Approximate Message Passing (AMP) and the Convex Gaussian Min-max Theorem (CGMT) for regularized linear regression and M-estimation.
Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more
This paper proposes a preconditioning layer for stable weight conditioning in LLM training.
Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more
This paper proposes a preconditioning layer for stable weight conditioning in LLM training.
This paper proposes a data collection strategy using solver iterates to augment datasets for training generative models, improving the efficiency of the data-model-optimization loop in one-sided box-c…
This paper provides a comprehensive generalization analysis of Stochastic Gradient Descent with Momentum (SGDM) by establishing tight, on-average model stability bounds that show SGDM can generalize w…
This paper introduces Curvature-Weighted Gradient Diversity (CWGD), a geometry-aware measure for optimization noise that reduces the asymptotic optimization error floor by up to a factor of two compar…
The paper introduces Multifidelity Proper Orthogonal Decomposition (MFPOD), a method that significantly reduces the computational cost of dimension reduction by intelligently combining data from cheap…
This paper models the changes in graph adjacency matrices over time as a low-rank matrix and develops a method for jointly interpolating signals and estimating graph adjacency matrices using this assu…
This paper proposes a novel estimator for the target linear coefficient in a partial linear model with black-box nuisance estimation and establishes its unimprovable error rate.
VISReg introduces a novel regularization technique that combines variance control with a Sliced-Wasserstein-based sketching objective to stabilize self-supervised learning, achieving state-of-the-art…
The paper proposes FOAM, an adaptive damping method that stabilizes the Shampoo optimization algorithm by dynamically controlling damping and eigendecomposition frequency, thereby reducing staleness-i…
This paper develops a framework for identifying and estimating parameters of interest in automatic debiased machine learning using a Riesz representer, which is identified when it uniquely optimizes a…
Amirpasha Hedayat, Ali Mohaghegh, Laura Balzano, Cheng Huang +1 more
The paper introduces a history-aware adaptive Reduced-Order Model (ROM) framework using incremental Singular Value Decomposition (iSVD) that maintains accuracy for online dynamics far beyond the initi…
This paper presents a local computation algorithm to approximate the top eigenvector of a symmetric matrix with entries between -1 and 1, building on Swartworth and Woodruff's work.