~ similar to 2607.19973· 19 results
The paper demonstrates that the location and nature of state encoding in sequence models are not fixed architectural traits but are highly dependent on the specific task, showing that the encoding pro…
The paper introduces HOLA (Hippocampal Linear Attention), a semiparametric test-time memory system that combines a compressive linear-attention state with a bounded exact cache for key-value associati…
The paper identifies five persistent, deep-seated behavioral patterns ('training strata') in LLMs, observed through long-term, intimate human-AI interaction, suggesting that training artifacts survive…
The paper tracks the developmental emergence of attention circuits in 1B-class language models, finding that the formation of induction and attention-sink circuits are distinct, temporally separated t…
This paper discusses the current understanding of Large Language Models (LLMs), their capabilities, and their relationship to human cognition, with a focus on emerging capabilities and mechanistic imp…
This paper introduces Accessibility Plasticity, a principle of adaptive computation that allows systems to adapt by reorganizing which existing computations can interact and participate, reducing the…
A modular, event-driven neuromorphic architecture for spiking neural network inference is presented, allowing for flexible configuration of neuron model, precision, and partitioning.
This paper challenges the claim that neural networks have met the challenge of systematicity in language and thought as proposed by Fodor and Pylyshyn, demonstrating limitations in a recent neural net…
This paper compares the cost-performance trade-off of Hebbian learning, Dense Difference Target Propagation (DDTP), and backpropagation (BP) using mutual-information-based measures.
The authors show that an explicit information bottleneck in a recurrent neural network is necessary for rotational and out-of-distribution generalization in a time-series prediction task, and that the…
A new error backpropagation method called supervised counterstream learning is proposed for deep associative networks, which only requires recognition of errors during training and backpropagates corr…
The paper introduces NEST, a graph-theoretic representational ontology for modeling cognition as structured state formation and transformation.
This paper explores how different components of the Transformer feedforward block architecture impact rank preservation across depth during initialization.
Przemyslaw Biecek, Luca Longo, Jianlong Zhou, Thomas Fel +2 more
The paper advocates for the establishment of Model Science, a systematic discipline that moves beyond simple benchmarking to deeply analyze AI models' internal workings and failure modes.
This paper investigates the use of contrastive objectives for brain decoding using functional MRI (fMRI) activity and shows that linear contrastive decoders outperform other methods.