~ similar to 2607.16761· 19 results
This paper explores how different components of the Transformer feedforward block architecture impact rank preservation across depth during initialization.
This paper derives a method to predict and mitigate intruder dimensions caused by LoRA fine-tuning in deep learning models, improving performance and reducing forgetting.
Tim Nielen, Sameer Ambekar, Johannes Kiechle, Daniel M. Lang +1 more
This paper identifies prediction bias, a failure mode of entropy minimization in test-time adaptation, and proposes Distribution Shift Bias Reduction (DSBR) to stabilize adaptation and prevent model c…
Qiao Xiao, Boqian Wu, Patrik Okanovic, Tomasz Sternal +5 more
The paper introduces Sparse Memory-Efficient Training (SMET), a method that stabilizes and optimizes Dynamic Sparse Training (DST) for large language models, enabling stable and memory-efficient spars…
This paper explores the possibility of using intrinsic device noise in analog neuromorphic hardware as a consolidation mechanism instead of an accuracy tax.
VideoMLA introduces a novel Multi-Head Latent Attention (MLA) mechanism that replaces per-head KV caches with a shared low-rank content latent, significantly reducing memory and improving throughput f…
Qi Liu, Mingdi Sun, Yongyi He, Zhi Zheng +4 more
The paper proposes EKSFT, a selective fine-tuning method that masks high-entropy or high-KL divergence tokens during Supervised Fine-Tuning (SFT) to prevent distribution shift and improve subsequent R…
This paper compares two theoretical frameworks for hierarchical neural networks with a finite but large number of hidden units and shows that training input-to-hidden weights reduces generalization er…
This paper presents a Noise-modulated Neural Network (NNN) that learns and infers with noise, reconstructing backpropagation from forward-pass statistics alone.
This paper evaluates the consistency and effectiveness of decoding methods for diffusion large language models (dLLMs) across diverse evaluation settings and reveals their sensitivity to prompt templa…
Ruixuan Huang, Yipei Wang, Wenyi Fang, Hantao Huang +6 more
The paper proposes methods for detecting training instability in large language models using internal monitors based on the functional role of critical modules and earliest computational sites.
The paper systematically characterizes column-level activation sparsity across various diffusion model architectures, demonstrating that element-level sparsity metrics significantly overestimate the a…
Longxuan Yu, Yunshu Wu, Yu Fu, Siheng Xiong +4 more
The paper introduces DSL-LLaDA, a method that lightly adapts a pre-trained masked diffusion language model to perform continuous denoising in embedding space, significantly improving text generation q…
This paper evaluates rank-order N-of-M encoding as an alternative to threshold-binary encoder in Sparse Distributed Memory systems and shows its outperformance in capacity experiments and robustness e…
肖代替了视觉令牌的永久删除,通过可恢复的路由来改进视觉语言模型的性能
The paper introduces TRACER, a novel regularization framework that uses Weighted Moving Average (WMA) distillation to robustly finetune multimodal models, mitigating catastrophic forgetting and improv…