ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Understanding of thermal simulation, GPU-accelerated computing, sparse solvers”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.AREmpiricalRecentJun 16, 2026

CUTh-Solver: GPU-Accelerated Sparse Matrix Solver for High-Resolution Thermal Simulation of 3D ICs

Chenghan Wang, Zhen Zhuang, Shui Jiang, Siyuan Liang +8 more

This paper proposes CUTh-Solver, a GPU-accelerated Preconditioned Conjugate Gradient (PCG)-based sparse solver framework for high-resolution 3D IC thermal simulation, achieving significant speedup ove…

View →
cs.GRcs.CVcs.DCEmpiricalRecentJul 17, 2026

Rendering 3D Gaussians on a Graph Processor

Nicholas Fry, Ignacio Alzugaray, Mark Pupilli, Paul H. J. Kelly +1 more

First implementation of a 3D Gaussian renderer on an Intelligence Processing Unit (IPU) with 1,472 independent tiles using only on-chip SRAM.

View →
cs.DCEmpiricalRecentJul 2, 2026

Elasticity in Parallel Sparse Triangular Solve

Raphael S. Steiner, Christos K. Matzoros, Pál András Papp, Toni Böhnlein +1 more

This paper introduces Stale Synchronous Parallel mode of execution for parallel sparse triangular linear system solve and presents a scheduler that achieves geometric-mean speed-ups of 7-30% over Grow…

View →
cs.DCEmpiricalRecentJul 24, 2026

$g$MAGNUS: Fast SpGEMM on GPUs for Irregular Matrices via Hierarchical Multisplit

Jordi Wolfson-Pou, Ahmed Helal, Fabrizio Petrini

The paper introduces $g$MAGNUS, a new algorithm for sparse matrix-matrix multiplication on GPUs that addresses heavy rows by reordering intermediate products and achieves significant speedups.

View →
physics.comp-phcs.DCphysics.flu-dynEmpiricalRecentJul 8, 2026

Scaling WaterLily.jl with MPI and an improved geometric multigrid solver

Bernat Font, Marin Lauber, Tzu-Yao Huang, Gabriel D. Weymouth

The paper presents improvements to the performance and scalability of WaterLily.jl, a scale-resolving incompressible flow solver, through the addition of MPI-based parallelism and optimizations to the…

View →
cs.ARcs.ETRecentJun 4, 2026

Space-CIM: Enabling Compute-In-Memory Accelerators for Thermally-Constrained Space Platforms

Sohan Salahuddin Mugdho, Md. Shahedul Hasan, Cheng Wang

This paper investigates the thermal constraints of deploying AI compute infrastructure in space, comparing GPUs and compute-in-memory (CIM) accelerators using a co-design methodology.

View →
cs.AREmpiricalRecentJul 22, 2026

DGNA: Dissecting GPU NUMA Architecture through Microbenchmarking and Data Analysis

Changxi Liu, Yun Chen, Trevor E. Carlson

This paper introduces DGNA, a methodology to unveil the Non-Uniform Memory Access (NUMA) architecture of GPU memory hierarchy through microbenchmarking and data analysis.

View →
cs.DCEmpiricalRecentJul 15, 2026

DRIFT: Direct Reduced Fourier Transforms for Distributed Spectral Neural Operators

Sana Taghipour Anvari, David Kaeli

This paper introduces the Distributed Truncated Spectral Transform (DTST) for Fourier Neural Operators (FNOs), achieving significant speedups in distributed computing.

View →
cs.PFcs.ARcs.DCRecentMay 27, 2026

Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory

Myeong Jun Jo

The paper introduces Rotary GPU, an exploratory execution approach demonstrating that large Mixture-of-Experts models can be run locally on consumer GPUs with limited VRAM, achieving usable decode thr…

View →
cs.CRcs.ARcs.LGRecentMar 20, 2026

Hawkeye: Reproducing GPU-Level Non-Determinism

Erez Badash, Dan Boneh, Ilan Komargodski, Megha Srivastava

Hawkeye is a system that allows perfect, precision-preserving reproduction of GPU-level matrix multiplication operations on a CPU, enabling efficient and trustworthy third-party auditing of machine le…

View →
cs.LGcs.AIRecentJun 1, 2026

FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo

Kyunghun Nam, Sumyeong Ahn

The paper proposes FOAM, an adaptive damping method that stabilizes the Shampoo optimization algorithm by dynamically controlling damping and eigendecomposition frequency, thereby reducing staleness-i…

View →
cs.CEphysics.comp-phRecentMay 27, 2026

Unified sparse framework for large-scale material point method simulations

Yidong Zhao, Lars Blatny, Xiang Feng, Mikkel M. Juel +2 more

This paper proposes a unified sparse background-grid framework for the Material Point Method (MPM), significantly reducing computational time and memory usage in large-scale simulations where the mate…

View →
cs.DScs.DCcs.MSEmpiricalRecentJul 27, 2026

Right Multiplication on Grammar-Compressed Matrices: A Streaming, Memory-Bounded GPU Engine

Francesco Tosoni, Gabriele Mencagli

This paper presents a method for compressing matrices using a RePair straight-line program (SLP), allowing matrix-vector products with time and space proportional to the compressed size, and demonstra…

View →
cs.AREmpiricalRecentJul 7, 2026

GPU-Accelerated Effective Resistance Analysis for 3D IC Power Delivery Network

Jingchao Hu, Cheng Zhuo, Zhou Jin

This paper proposes a GPU-accelerated framework for analyzing effective resistance in 3D IC power delivery networks, achieving significant speedup with negligible error.

View →
cs.CRcs.DSRecentApr 6, 2026

Packing Entries to Diagonals for Homomorphic Sparse-Matrix Vector Multiplication

Kemal Mutluergil, Deniz Elbek, Kamer Kaya, Erkay Savaş

This paper proposes methods to optimally permute the rows and columns of a sparse matrix to minimize the number of cyclic diagonals required for homomorphic sparse-matrix vector multiplication, signif…

View →