20 results for “proper loss optimization”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more
This paper proposes a preconditioning layer for stable weight conditioning in LLM training.
This paper introduces survey sampling techniques to estimate or minimize empirical pairwise loss functions, showing that targeting informative pairs significantly reduces computational cost while main…
The paper establishes tight upper and lower bounds on the statistical cost of approximate machine unlearning for smooth strongly convex losses, showing that the optimal unlearning rate depends critica…
This paper shows that stochastic nonconvex optimization can be reduced to ordinary static regret minimization in online convex optimization, and establishes convergence rates for smooth and Lipschitz…
The paper introduces Singularity-aware Adam (S-Adam), a novel optimizer that stabilizes deep learning training in non-smooth loss landscapes by dynamically damping updates based on local geometric ins…
This paper analyzes the poor performance of Meta-learning for Training-data Selection (MTS) and proposes that increasing the batch size and incorporating informative features can significantly improve…
Yuanjian Xu, Jianing Hao, Wanbo Zhang, Zhong Li +1 more
The paper proposes DiReCT, a novel framework that treats data selection during LLM annealing as a constrained optimization problem based on the spectral geometry of the loss landscape, achieving state…
Yuto Omae, Kazuki Sakai, Yohei Kakimoto, Makoto Sasaki +2 more
This paper derives the gradient of the Wolkowicz-Styan upper bound on the maximum eigenvalue of the cross-entropy loss Hessian in three-layer NNs to characterize directions leading to flat minima and…
The paper proposes a novel online learning algorithm that achieves an interval regret bound scaling with gradient variation, providing strong theoretical guarantees for non-stationary environments.
This paper derives the complete theory of covariance regret functional for decision making, providing insights on steepest-descent directions and boundary-optimal solutions.
This paper shows that projected Adam with arbitrary moment decay parameters can have non-zero average regret in online optimization.
Cheng-Han Huang, Yongliang Sun, Chaoyan Huang, Ismail Alkhouri +1 more
The paper establishes conditions for QUBO formulations of combinatorial optimization problems that guarantee valid binary and feasible local minimizers using gradient-based methods.
This paper explains the concept of 'weight-norm criticality' in deep neural network training, which is a stability issue induced by the interaction between normalization and weight decay.
This paper proposes HT-PAder, a parameter-free algorithm for online convex optimization in non-stationary environments with heavy-tailed noise, achieving an expected universal dynamic regret.
Dongjun Kim, Adrian de Wynter, Huancheng Chen, Heasung Kim +1 more
The paper introduces FoLoRA, a novel optimization framework that uses a generalized Rayleigh quotient to achieve a superior balance between adapting foundation models to specific tasks and preserving…
This paper introduces approximation-preserving coresets, which provide weaker guarantees than strong coresets but stronger guarantees than weak coresets for preserving the costs of good solutions in b…
This paper introduces Neural Certificate Pricing (NCP), an unsupervised learning framework that exploits the asymmetry between certifiable discrete structures and structural feasibility in combinatori…
The paper proposes using pseudo-sensitivities, derived from adjoint sensitivity fields, as an optimal conditioning signal in a Bernoulli flow-matching framework to significantly improve the out-of-dis…
The authors compare and evaluate 20 software configuration optimizers based on six assumptions about the data, and find that no single optimizer outperforms others across all budgets. They propose a t…
The paper proposes a proactive client selection framework that optimizes the selection of client subsets to ensure high data utility and fairness before federated learning begins, leading to faster an…