ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “INT8 quantization”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDcs.CLeess.ASEmpiricalRecentJul 23, 2026

VibeVoice-ASR-BitNet Technical Report

Songchen Xu, Ting Song, Shaohan Huang, Zhiliang Peng +6 more

The paper introduces VibeVoice-ASR-BitNet, a compressed real-time speech recognition model optimized for edge CPUs using heterogeneous quantization and custom SIMD kernels.

View →
cs.LGcs.ITeess.SPEmpiricalRecentJun 19, 2026

Fast-TurboQuant: A Multiplier-Free Online Vector Quantization Approach

Pedro M. R. Pereira, Felipe A. P. de Figueiredo, Rausley A. A. de Souza

The paper introduces Fast-TurboQuant, a multiplier-free projection architecture for large language models that uses a structured fast Johnson-Lindenstrauss transform instead of dense matrices, resulti…

View →
cs.LGcs.CVEmpiricalRecentJul 23, 2026

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann

This paper proposes KroQuant, a post-training quantization method for diffusion transformers using learned Kronecker-structured invertible transforms, which reduces parameters, improves speed, and mai…

View →
cs.NEEmpiricalRecentJul 16, 2026

Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models

Ryona Noda

This paper proposes a method for quantizing large language models with cross-layer error compensation and finite-sample feature-statistics matching.

View →
cs.LGcs.AIRecentMay 28, 2026

HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization

Artur Zagitov, Gleb Molodtsov, Aleksandr Beznosikov

HARP introduces a novel, adaptive, learnable orthogonal processor that significantly improves the robustness and accuracy of extreme low-bit LLM quantization compared to fixed methods.

View →
cs.LGcs.AIstat.MLTheoreticalRecentJul 23, 2026

Hilbert Operator for Progressive Encoding (HOPE): A Mathematical Framework for Deconstructing Learned Representations in Deep Networks

Hossein Mobahi, Peter L. Bartlett

The paper introduces HOPE, a mathematical framework for network compression that shifts representation deconstruction from the discrete domain to a Hilbert space, enabling unbiased architectural decis…

View →
cs.AIcs.LGRecentJun 1, 2026

Extreme Low-Bit Inference in Reasoning Models: Failure Modes and Targeted Recovery

Ekaterina Alimaskina, Darya Rudas, Denis Shveykin, Gleb Molodtsov +2 more

The paper analyzes the failure modes of aggressive 2-bit quantization in large reasoning models, proposing lightweight controls like FP16 planning and loop rescue to restore accuracy and achieve pract…

View →
cs.MScs.DCmath.NATheoreticalRecentJun 28, 2026

Improved Scaling for Fast Mode of Ozaki Scheme II

Shota Kawakami, Daisuke Takahashi

The paper proposes a scale-invariant scaling formula for Ozaki scheme II to emulate high-precision matrix multiplication using low-precision integer matrix operations, ensuring CRT uniqueness conditio…

View →
cs.ARRecentMay 31, 2026

Linear Complexity Fermionic Simulation on Quantum Devices with Hardware Connectivity Constraints

Xiangyu Gao, Winston Li, Jiakang Li, Zirui Li +3 more

The paper introduces Accordion, an end-to-end framework that significantly improves the efficiency of compiling fermionic Hamiltonians into quantum circuits for simulation on constrained quantum hardw…

View →
quant-phcs.CGmath.ATRecentMay 27, 2026

Quantum encodings that preserve persistent homology

Arthur J. Parzygnat, Andrew Vlasic

The paper investigates which quantum encodings can be applied directly to classical data point clouds while preserving the topological invariants necessary for topological data analysis (TDA).

View →
cs.AIRecentMay 27, 2026

Clark Hash: Stateless Sparse Johnson-Lindenstrauss Quantization for Neural Embeddings

Stanislav Kirdey, Clark Labs Inc

Clark Hash is a stateless, deterministic quantization method that significantly reduces the storage size of neural embeddings while maintaining high accuracy for cosine similarity search.

View →
cs.ITcs.LGmath.STTheoreticalRecentJul 3, 2026

Open Problem: Is Interaction Necessary for Order-Optimal 1-bit Mean Estimation?

Ivan Lau, Jonathan Scarlett

This paper investigates the necessity of interaction for order-optimal 1-bit mean estimation in nonparametric finite-moment classes.

View →
cs.LGcs.AIcs.CLRecentMay 29, 2026

Finer Parameter Steps for Low-Rank PEFT: A Controlled Study with CP Tensor Adapters

Xinjue Wang, Xiuheng Wang, Yejun Zhang, Sergiy A. Vorobyov +2 more

The paper investigates whether using fine-grained, tensorized adapters (CP components) instead of standard LoRA ranks improves the accuracy-budget trade-off in PEFT, finding that while they fill budge…

View →