~ similar to 2607.23397· 19 results
This paper characterizes the gap between neural network performance and their neural tangent kernel limit on compositional tasks, attributing it to a mismatch between kernel smoothness bias and target…
This paper proposes Hierarchical Block-Local Learning (HBLL), a framework for training deep neural networks without full end-to-end backpropagation, achieving $\mathcal{O}(\log N)$ parallel time compl…
This paper characterizes continual classification in homogeneous models as sequential projections and identifies regularity properties for local linear convergence.
This paper investigates limitations of learning tanh neural networks under finite-precision computations and Lp accuracy guarantees.
This paper investigates limitations of learning tanh neural networks under finite-precision computations and Lp accuracy guarantees.
This paper explores how different components of the Transformer feedforward block architecture impact rank preservation across depth during initialization.
Yuto Omae, Kazuki Sakai, Yohei Kakimoto, Makoto Sasaki +2 more
This paper derives the gradient of the Wolkowicz-Styan upper bound on the maximum eigenvalue of the cross-entropy loss Hessian in three-layer NNs to characterize directions leading to flat minima and…
This paper provides the first algorithmic separation between constant-depth and logarithmic-depth networks, identifying a class of Boolean functions that logarithmic-depth networks can learn efficient…
Kilian Rueß, Gennadiy Averkov, Florestan Brunck, Moritz Grillo +6 more
This paper proves that the maximum of up to 10 real numbers can be exactly represented by a ReLU network with two hidden layers, and shows that the same depth bound holds for all continuous piecewise-…
Rania Zitouni, Nadine Bousdjira, Sarah Hasnaoui, Amel Sadoun +1 more
This paper compares and optimizes CUDA strategies for a shallow neural network, achieving a 1.41x speedup on a large dataset.
The scaling exponent in neural scaling laws is not fixed but systematically depends on the optimizer used, with preconditioned optimizers generally yielding steeper scaling.
Maksym Zubkov, Carol Wu, Shiwei Yang, Param Mody +1 more
This paper studies the expressivity of shallow polynomial neural networks with monomial activation functions over finite fields, quantifying it by the cardinality of the neuromanifold and deriving low…
The study finds that specific, interpretable neuron populations (Rosetta Neurons) exhibit predictable, scale-dependent changes in selectivity and specialization as neural models grow larger.
This paper introduces a mechanistic neuronal network model for multilayer learning, offering biological insights and an alternative to backpropagation.
This paper shows that in the large depth and width asymptotics, Dropout and Random Gradient Masking (RaM) converge to the same limiting dynamics for ResNets.