ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “preconditioned conjugate gradient method”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.LGcs.AIEmpiricalRecentJun 4, 2026

PC Layer: Polynomial Weight Preconditioning for Improving LLM Pre-Training

Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more

This paper proposes a preconditioning layer for stable weight conditioning in LLM training.

View →
math.NAstat.MLTheoreticalRecentJul 28, 2026

Sequential Preconditioned Conjugate Gradient Method for Linear Statistical Models

Guan-Yu Chen, Dong-Yue Xie, Xi Yang, Zun-Hao Zheng

This paper proposes a randomized iterative method called Sequential Preconditioned Conjugate Gradient Method (SPCG) for large-scale linear statistical models, which significantly reduces computational…

View →
cs.LGcs.AIRecentJun 1, 2026

FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo

Kyunghun Nam, Sumyeong Ahn

The paper proposes FOAM, an adaptive damping method that stabilizes the Shampoo optimization algorithm by dynamically controlling damping and eigendecomposition frequency, thereby reducing staleness-i…

View →
cs.LGmath.OCstat.MLTheoreticalRecentJul 20, 2026

Optimizing the Preconditioner: A Black-box Online-to-Nonconvex Conversion with Static Regret Minimization Oracles

Haichen Hu, David Simchi-Levi

This paper shows that stochastic nonconvex optimization can be reduced to ordinary static regret minimization in online convex optimization, and establishes convergence rates for smooth and Lipschitz…

View →
cs.LGcs.AIRecentMay 27, 2026

Learning Theory of the SVRG: Generalization and Convergence Analysis

Yunwen Lei, Zimeng Wang, Xiaoming Yuan

This paper provides the first non-vacuous generalization analysis for the Stochastic Variance Reduced Gradient (SVRG) method by establishing sharp, data-dependent algorithmic stability bounds, thereby…

View →
cs.AImath.OCRecentJun 1, 2026

Stochastic convergence of parallel asynchronous adaptive first-order methods

Serge Gratton, Philippe L. Toint

The paper analyzes a new class of asynchronous adaptive first-order optimization methods and proves their stochastic convergence rate is O(1/sqrt{t}) for non-convex functions.

View →
math.OCcs.LGcs.NETheoreticalRecentJul 16, 2026

Fast and Scalable Caputo Fractional Gradient Descent via Perturbation-Preserving Memory Compression

Hwanseo Lee, Junseo Lee, Hyunju Kim

This paper proposes methods to make computationally expensive Caputo-based optimization more viable, while preserving its memory structure.

View →
cs.LGTheoreticalRecentJul 24, 2026

Complexity Bounds and Approaches to Learning Projected Gradient Descent Solver Iterates

Anjian Li, Ryne Beeson

This paper proposes a data collection strategy using solver iterates to augment datasets for training generative models, improving the efficiency of the data-model-optimization loop in one-sided box-c…

View →
cs.LGcs.AIRecentMay 27, 2026

Stochastic Gradient Descent with Momentum is Algorithmically Stable

Yunwen Lei, Zimeng Wang, Xiaoming Yuan

This paper provides a comprehensive generalization analysis of Stochastic Gradient Descent with Momentum (SGDM) by establishing tight, on-average model stability bounds that show SGDM can generalize w…

View →
math.NAmath.OCstat.MLTheoreticalRecentJul 18, 2026

A Deep Second-Order Stochastic Residual Method for Fully Nonlinear Parabolic PDEs

Zhenhua Zhao, Jihao Long

Introduce Deep Second-Order Stochastic Residual Method (D2SRM) for high-dimensional, Hessian-dependent fully nonlinear parabolic PDEs, establish well-posedness, and develop population-level convergenc…

View →
cs.LGmath.DGmath.OCEmpiricalRecentJun 28, 2026

Dead-Direction Conditioners: Gauge-Equivariant Preconditioning for Deep Networks

Tejas Pradeep Shirodkar

This paper introduces DDC, a Dead-Direction Conditioner that keeps a deep network's optimization on the symmetry quotient by conditioning the optimizer's state in the orbit decomposition of a $G$-inva…

View →
math.NAcs.CEcs.LGRecentJun 1, 2026

Physics-Informed Residuals for Adaptive Mesh Refinement in Finite-Difference PDE Solvers

Henry Kasumba, Ronald Katende

The paper proposes using a Physics-Informed Neural Network (PINN) residual as an efficient, physics-guided indicator to guide adaptive mesh refinement (AMR) for classical finite-difference PDE solvers…

View →
cs.LGmath.OCstat.MLEmpiricalRecentJul 26, 2026

A Multi-stage Constrained Optimization Framework for Data-driven Problems

Ye Shi

This paper proposes a Multi-stage Constrained Optimization Framework (MCOF) for Variational Autoencoders (VAEs) to address challenges in sampling, identifying active decision variables, and enforcing…

View →
cs.LGcs.AIcs.CVEmpiricalRecentJun 25, 2026

Error-Conditioned Neural Solvers

Haina Jiang, Liam Wang, Peng-Chen Chen, Min Seop Kwak +3 more

This paper proposes Error-conditioned Neural Solvers (ENS) for neural surrogate models to iteratively correct predictions by using the PDE residual field as input, achieving higher accuracy than optim…

View →
math.OCcs.AIcs.LGTheoreticalRecentJul 23, 2026

Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$

Dawei Li, Xiaotian Jiang, Mingyi Hong

This paper constructs strictly convex quadratic problems and initial points for which the long Barzilai--Borwein method does not converge root-superlinearly.

View →
cs.LGmath.OCstat.MLTheoreticalRecentJul 16, 2026

What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity

Kumar Kshitij Patel, Rustem Islamov, Sebastian U Stich, Aurelien Lucchi +2 more

This paper proves the conjecture that Local SGD outperforms Mini-batch SGD under bounded second-order heterogeneity for general convex objectives, improving the convergence guarantee and lower bounds.

View →
cs.ITTheoreticalRecentJun 26, 2026

Deriving Approximate Message Passing from the Convex Gaussian Min-Max Theorem

Vikrant Malik, Babak Hassibi

This paper establishes a direct connection between Approximate Message Passing (AMP) and the Convex Gaussian Min-max Theorem (CGMT) for regularized linear regression and M-estimation.

View →
math.OCcs.LGTheoreticalRecentJun 26, 2026

Second-Order KKT Guarantees for Bregman ADMM in Nonconvex and Non-Lipschitz Optimization

Shuang Li, Zhihui Zhu, Qiuwei Li

This paper analyzes Bregman ADMM for nonconvex linearly constrained problems under two-sided relative smoothness, showing convergence to strict saddle points and almost-sure second-order stationarity.

View →
cs.LGEmpiricalRecentJul 8, 2026

Neural Operator-enabled Topology-informed Evolutionary Strategy for PDE-Constrained Optimization

Xiangming Huang, Guannan Zhang, Lu Lu, Raphaël Pestourie

This paper introduces NOTES, a method for efficient and transferable inverse design of physical systems using neural operators, dimensionality reduction, and evolutionary optimization.

View →