ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “fast matrix multiplication”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.MScs.DCmath.NATheoreticalRecentJun 28, 2026

Improved Scaling for Fast Mode of Ozaki Scheme II

Shota Kawakami, Daisuke Takahashi

The paper proposes a scale-invariant scaling formula for Ozaki scheme II to emulate high-precision matrix multiplication using low-precision integer matrix operations, ensuring CRT uniqueness conditio…

View →
cs.DCEmpiricalRecentJul 24, 2026

$g$MAGNUS: Fast SpGEMM on GPUs for Irregular Matrices via Hierarchical Multisplit

Jordi Wolfson-Pou, Ahmed Helal, Fabrizio Petrini

The paper introduces $g$MAGNUS, a new algorithm for sparse matrix-matrix multiplication on GPUs that addresses heavy rows by reordering intermediate products and achieves significant speedups.

View →
cs.CRcs.DSRecentApr 6, 2026

Packing Entries to Diagonals for Homomorphic Sparse-Matrix Vector Multiplication

Kemal Mutluergil, Deniz Elbek, Kamer Kaya, Erkay Savaş

This paper proposes methods to optimally permute the rows and columns of a sparse matrix to minimize the number of cyclic diagonals required for homomorphic sparse-matrix vector multiplication, signif…

View →
cs.CRcs.ARcs.LGRecentMar 20, 2026

Hawkeye: Reproducing GPU-Level Non-Determinism

Erez Badash, Dan Boneh, Ilan Komargodski, Megha Srivastava

Hawkeye is a system that allows perfect, precision-preserving reproduction of GPU-level matrix multiplication operations on a CPU, enabling efficient and trustworthy third-party auditing of machine le…

View →
cs.DSTheoreticalRecentJul 9, 2026

Locally Approximating the Top Eigenvector of Bounded Entry Matrices

Nicolas Menand, Erik Waingarten

This paper presents a local computation algorithm to approximate the top eigenvector of a symmetric matrix with entries between -1 and 1, building on Swartworth and Woodruff's work.

View →
cs.DScs.DCcs.MSEmpiricalRecentJul 27, 2026

Right Multiplication on Grammar-Compressed Matrices: A Streaming, Memory-Bounded GPU Engine

Francesco Tosoni, Gabriele Mencagli

This paper presents a method for compressing matrices using a RePair straight-line program (SLP), allowing matrix-vector products with time and space proportional to the compressed size, and demonstra…

View →
cs.LGcs.ITeess.SPEmpiricalRecentJun 19, 2026

Fast-TurboQuant: A Multiplier-Free Online Vector Quantization Approach

Pedro M. R. Pereira, Felipe A. P. de Figueiredo, Rausley A. A. de Souza

The paper introduces Fast-TurboQuant, a multiplier-free projection architecture for large language models that uses a structured fast Johnson-Lindenstrauss transform instead of dense matrices, resulti…

View →
eess.SYcs.CRRecentMar 24, 2026

Secure Two-Party Matrix Multiplication from Lattices and Its Application to Encrypted Control

Kaoru Teranishi

The paper proposes a provably secure, single-round two-party computation protocol for approximate matrix multiplication using lattice-based cryptography, demonstrated for secure control law implementa…

View →
cs.CRcs.DCcs.DSRecentApr 13, 2026

GPU Acceleration of Sparse Fully Homomorphic Encrypted DNNs

Lara D'Agata, Carlos Agulló-Domingo, Óscar Vera-López, Kaustubh Shivdikar +6 more

The paper proposes a novel, optimized sparse matrix multiplication method for fully homomorphic encrypted deep neural networks, achieving up to a 3.0x speedup on AMD GPUs compared to CPU implementatio…

View →
eess.SPcs.DCEmpiricalRecentJul 26, 2026

Parallel Cascaded Recursive Filtering on Multi-Core CPUs and GPUs

Haotian Zhai, Bernd-Peter Paris

This paper reformulates cascaded second-order filtering as a block-tridiagonal linear system and develops parallel solution algorithms, achieving high performance on SIMD cores, multi-core CPUs, and G…

View →
cs.DCEmpiricalRecentJul 2, 2026

Elasticity in Parallel Sparse Triangular Solve

Raphael S. Steiner, Christos K. Matzoros, Pál András Papp, Toni Böhnlein +1 more

This paper introduces Stale Synchronous Parallel mode of execution for parallel sparse triangular linear system solve and presents a scheduler that achieves geometric-mean speed-ups of 7-30% over Grow…

View →
cs.DCcs.AIcs.CRRecentMay 21, 2026

Secure and Parallel Determinant Computation for Large-Scale Matrices in Edge Environments

Prajwal Panth

The paper proposes a Secure Parallel Determinant Computation (SPDC) framework that enables efficient, privacy-preserving, and scalable matrix determinant calculation across multiple untrusted edge ser…

View →
cs.DCEmpiricalRecentJul 15, 2026

DRIFT: Direct Reduced Fourier Transforms for Distributed Spectral Neural Operators

Sana Taghipour Anvari, David Kaeli

This paper introduces the Distributed Truncated Spectral Transform (DTST) for Fourier Neural Operators (FNOs), achieving significant speedups in distributed computing.

View →
cs.CCcs.DSTheoreticalRecentJul 2, 2026

Partition Rank and Algebraic Circuit Lower Bounds

Cornelius Brand, Petteri Kaski, Jiaheng Wang

This paper connects a generalized notion of tensor rank with multiplicative complexity, enabling control of arithmetic complexity in any constant degree of multilinearity and applications to fine-grai…

View →
cs.PLcs.MScs.SEEmpiricalRecentJul 28, 2026

Progress in Benchmarking Generics for Mathematical Computation

Daniel Pang, Stephen M. Watt

This paper reports on SciGMark 1.5, a benchmark study of specialized and generic implementations in modern languages, examining the consequences of various generic-realization strategies and extending…

View →
math.NAcs.LGmath.OCEmpiricalRecentJul 24, 2026

Singular value soft-thresholding via the polar decomposition

Stephen Becker

The paper describes how to compute singular value soft-thresholding using matrix polar decomposition for faster GPU processing.

View →
cs.ARRecentJun 1, 2026

O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead

Danilo Cammarata, Angelo Garofalo, Luca Benini

O-POPE is a novel outer-product engine that accelerates floating-point GEMM by repurposing FPU pipeline registers as buffers, achieving high utilization and improved energy efficiency.

View →