ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “proper loss optimization”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.LGcs.AIEmpiricalRecentJun 4, 2026

PC Layer: Polynomial Weight Preconditioning for Improving LLM Pre-Training

Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more

This paper proposes a preconditioning layer for stable weight conditioning in LLM training.

View →
stat.MLcs.LGRecentJun 1, 2026

Doing well with less! On Sampling Techniques for Empirical Pairwise Loss Estimation/Minimization

Louise Davy, Stephan Clémençon, Charlotte Laclau

This paper introduces survey sampling techniques to estimate or minimize empirical pairwise loss functions, showing that targeting informative pairs significantly reduces computational cost while main…

View →
cs.LGcs.CRRecentJun 1, 2026

Near-Optimal Pure Machine Unlearning for Smooth Strongly Convex Losses

Matthew Regehr, Gautam Kamath, Andrew Lowy

The paper establishes tight upper and lower bounds on the statistical cost of approximate machine unlearning for smooth strongly convex losses, showing that the optimal unlearning rate depends critica…

View →
cs.LGmath.OCstat.MLTheoreticalRecentJul 20, 2026

Optimizing the Preconditioner: A Black-box Online-to-Nonconvex Conversion with Static Regret Minimization Oracles

Haichen Hu, David Simchi-Levi

This paper shows that stochastic nonconvex optimization can be reduced to ordinary static regret minimization in online convex optimization, and establishes convergence rates for smooth and Lipschitz…

View →
cs.LGcs.AImath.OCRecentMay 28, 2026

Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization

Ruoran Xu, Borong She, Xiaobo Jin, Qiufeng Wang

The paper introduces Singularity-aware Adam (S-Adam), a novel optimizer that stabilizes deep learning training in non-smooth loss landscapes by dynamically damping updates based on local geometric ins…

View →
cs.LGcs.AIcs.CVRecentMay 30, 2026

On the Difficulty of Learning a Meta-network for Training Data Selection

Zilin Du, Junqi Zhao, Boyang Albert Li

This paper analyzes the poor performance of Meta-learning for Training-data Selection (MTS) and proposes that increasing the batch size and incorporating informative features can significantly improve…

View →
cs.CLRecentMay 29, 2026

Towards Efficient LLMs Annealing with Principled Sample Selection

Yuanjian Xu, Jianing Hao, Wanbo Zhang, Zhong Li +1 more

The paper proposes DiReCT, a novel framework that treats data selection during LLM annealing as a constrained optimization problem based on the spectral geometry of the loss landscape, achieving state…

View →
cs.LGcs.AIcs.NETheoreticalRecentJun 27, 2026

Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks

Yuto Omae, Kazuki Sakai, Yohei Kakimoto, Makoto Sasaki +2 more

This paper derives the gradient of the Wolkowicz-Styan upper bound on the maximum eigenvalue of the cross-entropy loss Hessian in three-layer NNs to characterize directions leading to flat minima and…

View →
cs.LGstat.MLRecentJun 2, 2026

Online Learning with Gradient-Variation Interval Regret

Yan-Feng Xie, Shuche Wang, Peng Zhao, Zhi-Hua Zhou

The paper proposes a novel online learning algorithm that achieves an interval regret bound scaling with gradient variation, providing strong theoretical guarantees for non-stationary environments.

View →
econ.EMcs.LGstat.MLTheoreticalRecentJul 21, 2026

Optimizing Regret

Irene Aldridge

This paper derives the complete theory of covariance regret functional for decision making, providing insights on steepest-descent directions and boundary-optimal solutions.

View →
cs.LGmath.OCstat.MLTheoreticalRecentJul 3, 2026

On the Convergence of Adam, Revisited

Steven Heilman, Sampad Mohanty

This paper shows that projected Adam with arbitrary moment decay parameters can have non-zero average regret in online optimization.

View →
cs.DMTheoreticalRecentJun 27, 2026

Local Minima in Quadratic-Penalty Relaxations of Binary Linear Programs

Cheng-Han Huang, Yongliang Sun, Chaoyan Huang, Ismail Alkhouri +1 more

The paper establishes conditions for QUBO formulations of combinatorial optimization problems that guarantee valid binary and feasible local minimizers using gradient-based methods.

View →
cs.LGcs.NEEmpiricalRecentJul 23, 2026

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay

Xiaolong Li, Zhangchen Zhou, Zhi-Qin John Xu

This paper explains the concept of 'weight-norm criticality' in deep neural network training, which is a stability issue induced by the interaction between normalization and weight decay.

View →
cs.LGcs.AImath.OCNEWTheoreticalJul 29, 2026

Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise

Vaneet Aggarwal

This paper proposes HT-PAder, a parameter-free algorithm for online convex optimization in non-stationary environments with heavy-tailed noise, achieving an expected universal dynamic regret.

View →
cs.LGcs.AIRecentMay 28, 2026

Foundation-Preserving Adaptation via Generalized Rayleigh-Quotient Optimization

Dongjun Kim, Adrian de Wynter, Huancheng Chen, Heasung Kim +1 more

The paper introduces FoLoRA, a novel optimization framework that uses a generalized Rayleigh quotient to achieve a superior balance between adapting foundation models to specific tasks and preserving…

View →
cs.DSTheoreticalRecentJun 15, 2026

Approximation Preserving Coresets

Milind Prabhu, Chris Schwiegelshohn, Sudarshan Shyam

This paper introduces approximation-preserving coresets, which provide weaker guarantees than strong coresets but stronger guarantees than weak coresets for preserving the costs of good solutions in b…

View →
cs.LGEmpiricalRecentJul 1, 2026

Neural Certificate Pricing for Combinatorial Optimization Problems

Jingyi Chen, Xinyuan Zhang, Xinwu Qian

This paper introduces Neural Certificate Pricing (NCP), an unsupervised learning framework that exploits the asymmetry between certifiable discrete structures and structural feasibility in combinatori…

View →
cs.LGcs.AIcs.CERecentJun 1, 2026

On the Generalization in Topology Optimization via Sensitivity-Conditioned Bernoulli Flow Matching

Mohammad Rashed, Duarte F. Valoroso Madeira, Babak Gholami, Caglar Guerbuez +2 more

The paper proposes using pseudo-sensitivities, derived from adjoint sensitivity fields, as an optimal conditioning signal in a Bernoulli flow-matching framework to significantly improve the out-of-dis…

View →
cs.SEEmpiricalRecentJul 13, 2026

Which Optimizer, At What Budget? A Tournament of Optimizers for Search-Based SE

Kishan Kumar Ganguly, Tim Menzies

The authors compare and evaluate 20 software configuration optimizers based on six assumptions about the data, and find that no single optimizer outperforms others across all budgets. They propose a t…

View →
cs.LGcs.CRRecentMay 20, 2026

Choose Wisely and Privately: Proactive Client Selection for Fair and Efficient Federated Learning

Adda Akram Bendoukha, Heber Hwang Arcolezi, Nesrine Kaaniche, Aymen Boudguiga

The paper proposes a proactive client selection framework that optimizes the selection of client subsets to ensure high data utility and fairness before federated learning begins, leading to faster an…

View →