ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

~ similar to 2607.16761· 19 results

cs.LGcs.AITheoreticalRecentJul 15, 2026

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

Katie Everett

This paper explores how different components of the Transformer feedforward block architecture impact rank preservation across depth during initialization.

View →
cs.LGstat.MLEmpiricalRecentJul 26, 2026

The Intruder Threshold: A Spectral Law for LoRA Fine-Tuning

Peng Xie

This paper derives a method to predict and mitigate intruder dimensions caused by LoRA fine-tuning in deep learning models, improving performance and reducing forgetting.

View →
cs.LGcs.CVRecentJun 1, 2026

Entropy Minimization without Model Collapse: Mitigating Prediction Bias in Medical Imaging

Tim Nielen, Sameer Ambekar, Johannes Kiechle, Daniel M. Lang +1 more

This paper identifies prediction bias, a failure mode of entropy minimization in test-time adaptation, and proposes Distribution Shift Bias Reduction (DSBR) to stabilize adaptation and prevent model c…

View →
cs.LGcs.AIRecentMay 30, 2026

Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling

Qiao Xiao, Boqian Wu, Patrik Okanovic, Tomasz Sternal +5 more

The paper introduces Sparse Memory-Efficient Training (SMET), a method that stabilizes and optimizes Dynamic Sparse Training (DST) for large language models, enabling stable and memory-efficient spars…

View →
cs.LGcs.NEEmpiricalRecentJul 8, 2026

Intrinsic-Noise Consolidation: A Doob-Barrier-Conditioned Diffusion Turns Analog Device Noise into a Continual-Learning Resource

Gunner Levi Howe

This paper explores the possibility of using intrinsic device noise in analog neuromorphic hardware as a consolidation mechanism instead of an accuracy tax.

View →
cs.CVcs.AIRecentMay 28, 2026

VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion

Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral, Adil Kaan Akan +3 more

VideoMLA introduces a novel Multi-Head Latent Attention (MLA) mechanism that replaces per-head KV caches with a shared low-rank content latent, significantly reducing memory and improving throughput f…

View →
cs.AIRecentMay 28, 2026

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models

Qi Liu, Mingdi Sun, Yongyi He, Zhi Zheng +4 more

The paper proposes EKSFT, a selective fine-tuning method that masks high-entropy or high-KL divergence tokens during Supervised Fine-Tuning (SFT) to prevent distribution shift and improve subsequent R…

View →
cs.LGmath.STstat.MLTheoreticalRecentJul 26, 2026

A Statistical Difference between Single-Layer Learning and Hierarchical Learning in Wide Neural Networks

Sumio Watanabe

This paper compares two theoretical frameworks for hierarchical neural networks with a finite but large number of hidden units and shows that training input-to-hidden weights reduces generalization er…

View →
cs.NEEmpiricalRecentJul 29, 2026

Reconstructing Backpropagation from Forward Fluctuations in Noise-modulated Neural Networks

Shuhei Ikemoto

This paper presents a Noise-modulated Neural Network (NNN) that learns and infers with noise, reconstructing backpropagation from forward-pass statistics alone.

View →
cs.CLcs.LGEmpiricalRecentJun 28, 2026

Understanding Evaluation Illusion in Diffusion Large Language Models

Hengxiang Zhang, Jiaxi Ren, Hongxin Wei

This paper evaluates the consistency and effectiveness of decoding methods for diffusion large language models (dLLMs) across diverse evaluation settings and reveals their sensitivity to prompt templa…

View →
cs.CLEmpiricalRecentJun 26, 2026

Mechanism-Driven Monitors for Preemptive Detection of LLM Training Instability

Ruixuan Huang, Yipei Wang, Wenyi Fang, Hantao Huang +6 more

The paper proposes methods for detecting training instability in large language models using internal monitors based on the functional role of critical modules and earliest computational sites.

View →
cs.ARcs.PFRecentMay 30, 2026

Regular-Activation Concentration: Characterizing Column-Level Output Sparsity Across Diffusion Model Architectures

Dazhi Yang, Shafayat Mowla Anik, Byeong Kil Lee, Jeeho Ryoo

The paper systematically characterizes column-level activation sparsity across various diffusion model architectures, demonstrating that element-level sparsity metrics significantly overestimate the a…

View →
cs.CLcs.AIRecentMay 31, 2026

DSL-LLaDA: Scaling Continuous Denoising to 8B Masked Diffusion LMs

Longxuan Yu, Yunshu Wu, Yu Fu, Siheng Xiong +4 more

The paper introduces DSL-LLaDA, a method that lightly adapts a pre-trained masked diffusion language model to perform continuous denoising in embedding space, significantly improving text generation q…

View →
cs.LGcs.NEEmpiricalRecentJul 3, 2026

Rank-Order N-of-M Codes for Sparse Distributed Memory: Disentangling Representation and Learning Effects in Noise Robustness Against Contemporary Neuromorphic Architectures

Joy Bose

This paper evaluates rank-order N-of-M encoding as an alternative to threshold-binary encoder in Sparse Distributed Memory systems and shows its outperformance in capacity experiments and robustness e…

View →
cs.CVcs.AIEmpiricalRecentJun 10, 2026

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models

Cheng-Yu Yang, Shao-Yuan Lo, Yu-Lun Liu

肖代替了视觉令牌的永久删除,通过可恢复的路由来改进视觉语言模型的性能

View →
cs.LGcs.AIcs.CVRecentMay 28, 2026

TRACER: Persistent Regularization for Robust Multimodal Finetuning

Hesam Asadollahzadeh, Feng Liu, Christopher Leckie, Sarah M. Erfani

The paper introduces TRACER, a novel regularization framework that uses Weighted Moving Average (WMA) distillation to robustly finetune multimodal models, mitigating catastrophic forgetting and improv…

View →