20 results for “preconditioned conjugate gradient method”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more
This paper proposes a preconditioning layer for stable weight conditioning in LLM training.
This paper proposes a randomized iterative method called Sequential Preconditioned Conjugate Gradient Method (SPCG) for large-scale linear statistical models, which significantly reduces computational…
The paper proposes FOAM, an adaptive damping method that stabilizes the Shampoo optimization algorithm by dynamically controlling damping and eigendecomposition frequency, thereby reducing staleness-i…
This paper shows that stochastic nonconvex optimization can be reduced to ordinary static regret minimization in online convex optimization, and establishes convergence rates for smooth and Lipschitz…
This paper provides the first non-vacuous generalization analysis for the Stochastic Variance Reduced Gradient (SVRG) method by establishing sharp, data-dependent algorithmic stability bounds, thereby…
The paper analyzes a new class of asynchronous adaptive first-order optimization methods and proves their stochastic convergence rate is O(1/sqrt{t}) for non-convex functions.
This paper proposes methods to make computationally expensive Caputo-based optimization more viable, while preserving its memory structure.
This paper proposes a data collection strategy using solver iterates to augment datasets for training generative models, improving the efficiency of the data-model-optimization loop in one-sided box-c…
This paper provides a comprehensive generalization analysis of Stochastic Gradient Descent with Momentum (SGDM) by establishing tight, on-average model stability bounds that show SGDM can generalize w…
Introduce Deep Second-Order Stochastic Residual Method (D2SRM) for high-dimensional, Hessian-dependent fully nonlinear parabolic PDEs, establish well-posedness, and develop population-level convergenc…
This paper introduces DDC, a Dead-Direction Conditioner that keeps a deep network's optimization on the symmetry quotient by conditioning the optimizer's state in the orbit decomposition of a $G$-inva…
The paper proposes using a Physics-Informed Neural Network (PINN) residual as an efficient, physics-guided indicator to guide adaptive mesh refinement (AMR) for classical finite-difference PDE solvers…
This paper proposes a Multi-stage Constrained Optimization Framework (MCOF) for Variational Autoencoders (VAEs) to address challenges in sampling, identifying active decision variables, and enforcing…
Haina Jiang, Liam Wang, Peng-Chen Chen, Min Seop Kwak +3 more
This paper proposes Error-conditioned Neural Solvers (ENS) for neural surrogate models to iteratively correct predictions by using the PDE residual field as input, achieving higher accuracy than optim…
This paper constructs strictly convex quadratic problems and initial points for which the long Barzilai--Borwein method does not converge root-superlinearly.
This paper proves the conjecture that Local SGD outperforms Mini-batch SGD under bounded second-order heterogeneity for general convex objectives, improving the convergence guarantee and lower bounds.
This paper establishes a direct connection between Approximate Message Passing (AMP) and the Convex Gaussian Min-max Theorem (CGMT) for regularized linear regression and M-estimation.
This paper analyzes Bregman ADMM for nonconvex linearly constrained problems under two-sided relative smoothness, showing convergence to strict saddle points and almost-sure second-order stationarity.
This paper introduces NOTES, a method for efficient and transferable inverse design of physical systems using neural operators, dimensionality reduction, and evolutionary optimization.