~ similar to 2606.17712· 20 results
This paper proposes FSZ, a GPU error-bounded lossy compressor with three innovations for higher compression ratios and throughput within a single CUDA kernel.
This paper measures the lower bound for the shortest program generating a sequence, proving a conservation law and providing a deterministic engine to recover generating programs for certain sequences…
Proposed a novel algorithm-hardware co-design, HiKV, to tackle the memory bottleneck in long-context large language models by exploiting KV cache redundancy through hierarchical importance awareness,…
Clark Hash is a stateless, deterministic quantization method that significantly reduces the storage size of neural embeddings while maintaining high accuracy for cosine similarity search.
This paper introduces HbHAI techniques using key-dependent hash functions for AI analysis with unprecedented performance and compression rate reduction.
This paper proposes a compression scheme for transmitting sparse local updates in distributed learning systems, and provides a converse based on f-divergence to characterize the communication-accuracy…
Guoxin Ma, Yibing Liu, Chengzhengxu Li, Yu Liang +6 more
The paper introduces Thinking as Compression (TaC), a novel paradigm showing that the inherent reasoning process of a large language model can naturally compress long context inputs, outperforming ded…
This paper systematically analyzes combining dimensionality reduction and quantization to compress text embeddings, showing that this combined approach achieves substantial compression (e.g., 0.1% siz…
This paper evaluates rank-order N-of-M encoding as an alternative to threshold-binary encoder in Sparse Distributed Memory systems and shows its outperformance in capacity experiments and robustness e…
The paper shows that deterministic cache eviction cannot ensure consistent serving-time error estimation and proposes a randomized approach to restore identifiability and provide error certificates.
This paper introduces an open-source framework for evaluating the efficacy of AI agents powered by open-weight large language models on data preparation tasks in research using locally deployable mode…
The paper proposes SubFit, a novel compression technique that achieves superior LLM compression by replacing non-contiguous, submodule-level components (Attention and FeedForward) with lightweight res…
ACRONYM is a novel algorithm-hardware co-designed platform that enables high-recall, continuous approximate nearest neighbor search in memory for dynamic vector databases, achieving massive throughput…
This paper proposes a novel Simultaneous Data Compression and Encryption (SDCE) system that combines chaotic map-based encryption with Huffman encoding to securely and efficiently transmit large video…