20 results for “Hessian”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Yuto Omae, Kazuki Sakai, Yohei Kakimoto, Makoto Sasaki +2 more
This paper derives the gradient of the Wolkowicz-Styan upper bound on the maximum eigenvalue of the cross-entropy loss Hessian in three-layer NNs to characterize directions leading to flat minima and…
Introduce Deep Second-Order Stochastic Residual Method (D2SRM) for high-dimensional, Hessian-dependent fully nonlinear parabolic PDEs, establish well-posedness, and develop population-level convergenc…
This paper introduces Curvature-Weighted Gradient Diversity (CWGD), a geometry-aware measure for optimization noise that reduces the asymptotic optimization error floor by up to a factor of two compar…
The paper investigates applying Riemannian optimization techniques to low-rank matrix parameters for deep learning, but finds that the proposed methods do not conclusively outperform the AdamW baselin…
This paper measures and analyzes the overuse of epanorthosis, a rhetorical figure, in large language models and proposes techniques to mitigate it.
The paper presents GIFT, a method for reducing communication volume in large language model pretraining by transforming gradients into a near-isotropic space before quantization.
The paper introduces Singularity-aware Adam (S-Adam), a novel optimizer that stabilizes deep learning training in non-smooth loss landscapes by dynamically damping updates based on local geometric ins…
The paper introduces Cellular Sheaf Neural Operators, a discretization-aware framework that models constrained PDEs by representing physical states on oriented cell complexes to enforce structure-pres…
This paper introduces a new tamed stochastic gradient Hamiltonian Monte Carlo (tSGHMC) algorithm for sampling and optimization problems with large stochastic gradients, providing error bounds and perf…
This paper introduces Stale Synchronous Parallel mode of execution for parallel sparse triangular linear system solve and presents a scheduler that achieves geometric-mean speed-ups of 7-30% over Grow…
This paper constructs strictly convex quadratic problems and initial points for which the long Barzilai--Borwein method does not converge root-superlinearly.
This paper presents a method for identifying probabilistic structures from empirical probability tensors using algebraic statistics and Kronecker-stack class of configuration matrices.
The paper analyzes congruence-based neural architectures for classifying positive-definite matrices, demonstrating that common semi-orthogonality constraints severely limit the model's expressivity.
This paper extends gradient boosting to functions of vector inputs using a simple algorithm with histogram-based decision trees.
This paper improves the foundations of neural likelihood approximation for Bayesian inverse problems by making the learning problem strictly convex and showing convergence to the true likelihood.
Mark Tensen, Ciaran Regan, Bert Wang-Chak Chan, Mizuki Oka +2 more
The paper introduces Microcosmos, a GPU-accelerated, differentiable simulation engine for artificial lifeforms as elastic filament chains in a 2D viscous fluid world, validated through experiments.
This paper introduces a mechanistic neuronal network model for multilayer learning, offering biological insights and an alternative to backpropagation.