ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “weight sparsity”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.LGcs.AIEmpiricalRecentJun 4, 2026

PC Layer: Polynomial Weight Preconditioning for Improving LLM Pre-Training

Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more

This paper proposes a preconditioning layer for stable weight conditioning in LLM training.

View →
cs.CLcs.LGEmpiricalRecentJul 8, 2026

PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning

Yazdan Jamshidi, Alexey Shvets

The paper proposes PALS, a method for adjusting per-layer sparsity based on activation magnitudes in transformer models, achieving better performance than uniform one-shot pruning methods.

View →
cs.AREmpiricalRecentJul 24, 2026

The Sparsity Tax: Weight Sparsity Trade-offs in Event-Driven SIMD and SIMT Neuromorphic Cores

Mattias Westerink, Sameed Sohail, Berend-Jan van der Zwaag, Sabih Gerez +1 more

This paper compares three neuromorphic core designs for handling weight sparsity in event-driven neural networks and quantifies the performance and energy costs.

View →
cs.LGcs.AIEmpiricalRecentJul 9, 2026

SLORR: Simple and Efficient In-Training Low-Rank Regularization

David González-Martínez, Shiwei Liu

The paper introduces SLORR, a simple and stateless framework for in-training low-rank regularization of neural networks using two variants based on the Hoyer sparsity metric and the nuclear norm.

View →
cs.DSmath.COTheoreticalRecentJul 9, 2026

Optimal Sparsifiers for Abelian Cayley Graphs

Arpon Basu, Pravesh K. Kothari, Raghu Meka, Stefan Tudose

This paper proves that for any finite abelian group, there exists a spectral sparsifier for its Cayley graph with log(|G|) generators. This result improves upon previous work for constructing code spa…

View →
cs.DSTheoreticalRecentJul 17, 2026

A Unified Theory of Sparsification

Sanjeev Khanna, Aaron Putterman, Madhu Sudan

This paper introduces a structural theorem for the sparsifiability of real-valued codes, which generalizes both combinatorial and continuous notions of sparsification.

View →
cs.DScs.DCTheoreticalRecentJul 27, 2026

Parallel Spectral Graph Sparsification via Low Diameter Decompositions

Yves Baumann, Gernot Zöcklein

A new solver-free parallel spectral sparsification algorithm for weighted graphs is presented, relying on low-diameter decompositions and independent sampling, eliminating dependence on target approxi…

View →
cs.LGcs.AIRecentMay 31, 2026

When Data Is Scarce: Scaling Sparse Language Models with Repeated Training

Boqian Wu, Qiao Xiao, Patrik Okanovic, Tomasz Sternal +5 more

This paper introduces a new scaling law for sparse language models trained with limited data, demonstrating that sparsity can significantly improve performance and delay data saturation during multi-e…

View →
cs.DScs.CCmath.CORecentMay 29, 2026

High-Dimensional Expanders, the Sparsest Cut Problem, and Steurer's Conjecture

Farzam Ebrahimnejad, Shayan Oveis Gharan

The paper refutes Steurer's conjecture regarding the existence of large constant-separated sets within families of unit-norm vectors with low average correlation, using high-dimensional expanders to s…

View →
q-bio.NCcs.LGRecentJun 1, 2026

How Optimality Structures Sparse Dictionaries: A Theory for Understanding SAE Representations

William Dorrell

The paper theoretically analyzes the properties that optimal sparse autoencoder (SAE) dictionaries must satisfy, deriving constraints that explain observed SAE behaviors like hierarchical splitting an…

View →
cs.IRcs.AIcs.LGRecentMay 28, 2026

No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval

Lixuan Guo, Yifei Wang, Tiansheng Wen, Aosong Feng +2 more

The paper introduces Single-stage Sparse Retrieval (SSR), a method that replaces computationally expensive vector clustering with sparse autoencoding to achieve highly efficient multi-vector retrieval…

View →
cs.LGcs.AIRecentJun 1, 2026

FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo

Kyunghun Nam, Sumyeong Ahn

The paper proposes FOAM, an adaptive damping method that stabilizes the Shampoo optimization algorithm by dynamically controlling damping and eigendecomposition frequency, thereby reducing staleness-i…

View →
cs.ARcs.PFRecentMay 30, 2026

Regular-Activation Concentration: Characterizing Column-Level Output Sparsity Across Diffusion Model Architectures

Dazhi Yang, Shafayat Mowla Anik, Byeong Kil Lee, Jeeho Ryoo

The paper systematically characterizes column-level activation sparsity across various diffusion model architectures, demonstrating that element-level sparsity metrics significantly overestimate the a…

View →