ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Hessian”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.LGcs.AIcs.NETheoreticalRecentJun 27, 2026

Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks

Yuto Omae, Kazuki Sakai, Yohei Kakimoto, Makoto Sasaki +2 more

This paper derives the gradient of the Wolkowicz-Styan upper bound on the maximum eigenvalue of the cross-entropy loss Hessian in three-layer NNs to characterize directions leading to flat minima and…

View →
math.NAmath.OCstat.MLTheoreticalRecentJul 18, 2026

A Deep Second-Order Stochastic Residual Method for Fully Nonlinear Parabolic PDEs

Zhenhua Zhao, Jihao Long

Introduce Deep Second-Order Stochastic Residual Method (D2SRM) for high-dimensional, Hessian-dependent fully nonlinear parabolic PDEs, establish well-posedness, and develop population-level convergenc…

View →
cs.LGmath.OCstat.MLTheoreticalRecentJun 29, 2026

Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules

Muhammad Hamza, Ayush Goel

This paper introduces Curvature-Weighted Gradient Diversity (CWGD), a geometry-aware measure for optimization noise that reduces the asymptotic optimization error floor by up to a factor of two compar…

View →
cs.LGRecentJun 1, 2026

Riemannian Gradient Descent for Low-Rank Architectures

Nicholas Knight

The paper investigates applying Riemannian optimization techniques to low-rank matrix parameters for deep learning, but finds that the proposed methods do not conclusively outperform the AdamW baselin…

View →
cs.CLcs.AIEmpiricalRecentJul 23, 2026

Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it

Federico Boggia

This paper measures and analyzes the overuse of epanorthosis, a rhetorical figure, in large language models and proposes techniques to mitigate it.

View →
cs.DCcs.LGEmpiricalRecentJul 8, 2026

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining

Jieying Wang, Shuyuan Fan, Mingkai Zheng, Zhao Zhang

The paper presents GIFT, a method for reducing communication volume in large language model pretraining by transforming gradients into a near-isotropic space before quantization.

View →
cs.LGcs.AImath.OCRecentMay 28, 2026

Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization

Ruoran Xu, Borong She, Xiaobo Jin, Qiufeng Wang

The paper introduces Singularity-aware Adam (S-Adam), a novel optimizer that stabilizes deep learning training in non-smooth loss landscapes by dynamically damping updates based on local geometric ins…

View →
cs.LGcs.CEmath.NARecentMay 31, 2026

Cellular Sheaf Neural Operators for Structure-Preserving Surrogate Modeling of Constrained PDEs

Lennon J. Shikhman, Shane Gilbertie

The paper introduces Cellular Sheaf Neural Operators, a discretization-aware framework that models constrained PDEs by representing physical states on oriented cell complexes to enforce structure-pres…

View →
math.OCmath.NAstat.MLTheoreticalRecentJul 16, 2026

Tamed Stochastic Gradient Hamiltonian Monte Carlo

Zhuoran Wang, Ying Zhang

This paper introduces a new tamed stochastic gradient Hamiltonian Monte Carlo (tSGHMC) algorithm for sampling and optimization problems with large stochastic gradients, providing error bounds and perf…

View →
cs.DCEmpiricalRecentJul 2, 2026

Elasticity in Parallel Sparse Triangular Solve

Raphael S. Steiner, Christos K. Matzoros, Pál András Papp, Toni Böhnlein +1 more

This paper introduces Stale Synchronous Parallel mode of execution for parallel sparse triangular linear system solve and presents a scheduler that achieves geometric-mean speed-ups of 7-30% over Grow…

View →
math.OCcs.AIcs.LGTheoreticalRecentJul 23, 2026

Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$

Dawei Li, Xiaotian Jiang, Mingyi Hong

This paper constructs strictly convex quadratic problems and initial points for which the long Barzilai--Borwein method does not converge root-superlinearly.

View →
stat.MLcs.LGEmpiricalRecentJul 21, 2026

Algebraic Signatures for Structural Learning in Probability Tensors

Akihiro Maeda, Shohei Hidaka, Satoshi Aoki

This paper presents a method for identifying probabilistic structures from empirical probability tensors using algebraic statistics and Kronecker-stack class of configuration matrices.

View →
cs.LGRecentJun 1, 2026

Expressivity of congruence-based architectures for DNNs on positive-definite matrices

Antonin Oswald, Estelle Massart

The paper analyzes congruence-based neural architectures for classifying positive-definite matrices, demonstrating that common semi-orthogonality constraints severely limit the model's expressivity.

View →
stat.MLcs.LGEmpiricalRecentJun 28, 2026

Gradient boosting with vector-valued leafs

David Cortes

This paper extends gradient boosting to functions of vector inputs using a simple algorithm with histogram-based decision trees.

View →
stat.MLcs.LGmath.PRTheoreticalRecentJul 7, 2026

A Convex Approximation Framework for Neural Likelihood-Based Bayesian Inverse Problems

Fabian Schneider, Tapio Helin, Leila Taghizadeh

This paper improves the foundations of neural likelihood approximation for Bayesian inverse problems by making the learning problem strictly convex and showing convergence to the true likelihood.

View →
cs.NEEmpiricalRecentJul 3, 2026

Microcosmos: Reimagining Artificial Life for the GPU Era

Mark Tensen, Ciaran Regan, Bert Wang-Chak Chan, Mizuki Oka +2 more

The paper introduces Microcosmos, a GPU-accelerated, differentiable simulation engine for artificial lifeforms as elastic filament chains in a 2D viscous fluid world, validated through experiments.

View →
cs.NEEmpiricalRecentJun 12, 2026

A Programmer's Guide to Cascaded Adaptive Combiners: Online Learning by Biologically Accurate Models of Multilayer Neuron Networks

Martin Nilsson, Denis Kleyko

This paper introduces a mechanistic neuronal network model for multilayer learning, offering biological insights and an alternative to backpropagation.

View →