~ similar to 2607.18574· 20 results
A new error backpropagation method called supervised counterstream learning is proposed for deep associative networks, which only requires recognition of errors during training and backpropagates corr…
This paper introduces modulo error routing to extend Error Diffusion beyond binary classification and achieves high performance on MNIST and CIFAR-10 under Dale's principle.
The paper introduces Score Broadcast and Decorrelation (SBD), a general theoretical framework that unifies broadcast-based credit assignment across various differentiable loss functions by leveraging…
This paper introduces DDC, a Dead-Direction Conditioner that keeps a deep network's optimization on the symmetry quotient by conditioning the optimizer's state in the orbit decomposition of a $G$-inva…
The paper introduces MENTIS, a geometry-first framework that measures how preference alignment structurally changes the internal computations of language models, finding that these changes are selecti…
This paper systematically characterises the sensitivity of emergent misalignment (EM) in LLMs to various training choices, finding that the choice of optimiser has the largest effect on misalignment r…
Magnus Jørgenvåg, David Kaczér, Lasse Ruttert, Marvin Gülhan +2 more
This paper demonstrates that reinforcement learning (RL) can cause emergent misalignment (EM) in open-weight models, showing that even seemingly harmless or natural reward signals can induce significa…
The paper formally addresses the challenging question of cross-domain transferability of latent predictive models by proposing a structured framework that quantifies the relationship between source an…
Incorporating short-term synaptic plasticity (STP) into a PFC-inspired reservoir model significantly stabilizes goal-conditioned dynamics, particularly under state noise, suggesting STP dynamically mo…
Jiahe Guo, Xiangran Guo, Jiaxuan Chen, Weixiang Zhao +5 more
This paper introduces the concept of Safety Geometry Collapse, demonstrating that multimodal inputs degrade the safety separation of LLMs, and proposes ReGap, a training-free method that adaptively co…
The paper proposes CTRL-STEER, a closed-loop framework that adaptively adjusts intervention strength to stabilize concept regulation and improve task success in Vision-Language-Action models without r…
This paper proposes the stable signal principle to explain the convergence behavior of retraining in performative prediction systems, revealing regularization as a natural force to control performativ…
The paper demonstrates that LLM performance in zero-shot annotation is significantly limited by the alignment between the model's internal understanding and the task definition, showing that prompt-ba…
This paper introduces Neuro-Symbolic Manifold Alignment (NSMA), a method that embeds rule decisions as anchors inside the latent space of a neural policy to improve generalization and outperforms stat…
The paper proposes a local perturbation theory showing that cross-domain interference in multi-domain RL occurs via a low-dimensional shared conflict subspace, which can be selectively mitigated by sh…
The paper proposes evaluating certified training methods by comparing their Pareto fronts across the natural-certified accuracy trade-off, revealing superior performance and previously unappreciated c…