20 results for “INT8 quantization”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Songchen Xu, Ting Song, Shaohan Huang, Zhiliang Peng +6 more
The paper introduces VibeVoice-ASR-BitNet, a compressed real-time speech recognition model optimized for edge CPUs using heterogeneous quantization and custom SIMD kernels.
The paper introduces Fast-TurboQuant, a multiplier-free projection architecture for large language models that uses a structured fast Johnson-Lindenstrauss transform instead of dense matrices, resulti…
This paper proposes KroQuant, a post-training quantization method for diffusion transformers using learned Kronecker-structured invertible transforms, which reduces parameters, improves speed, and mai…
This paper proposes a method for quantizing large language models with cross-layer error compensation and finite-sample feature-statistics matching.
HARP introduces a novel, adaptive, learnable orthogonal processor that significantly improves the robustness and accuracy of extreme low-bit LLM quantization compared to fixed methods.
The paper introduces HOPE, a mathematical framework for network compression that shifts representation deconstruction from the discrete domain to a Hilbert space, enabling unbiased architectural decis…
The paper analyzes the failure modes of aggressive 2-bit quantization in large reasoning models, proposing lightweight controls like FP16 planning and loop rescue to restore accuracy and achieve pract…
The paper proposes a scale-invariant scaling formula for Ozaki scheme II to emulate high-precision matrix multiplication using low-precision integer matrix operations, ensuring CRT uniqueness conditio…
Xiangyu Gao, Winston Li, Jiakang Li, Zirui Li +3 more
The paper introduces Accordion, an end-to-end framework that significantly improves the efficiency of compiling fermionic Hamiltonians into quantum circuits for simulation on constrained quantum hardw…
The paper investigates which quantum encodings can be applied directly to classical data point clouds while preserving the topological invariants necessary for topological data analysis (TDA).
Clark Hash is a stateless, deterministic quantization method that significantly reduces the storage size of neural embeddings while maintaining high accuracy for cosine similarity search.
This paper investigates the necessity of interaction for order-optimal 1-bit mean estimation in nonparametric finite-moment classes.
Xinjue Wang, Xiuheng Wang, Yejun Zhang, Sergiy A. Vorobyov +2 more
The paper investigates whether using fine-grained, tensorized adapters (CP components) instead of standard LoRA ranks improves the accuracy-budget trade-off in PEFT, finding that while they fill budge…