ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “NUMA balancing”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.LGcs.AIEmpiricalRecentJun 4, 2026

PC Layer: Polynomial Weight Preconditioning for Improving LLM Pre-Training

Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more

This paper proposes a preconditioning layer for stable weight conditioning in LLM training.

View →
cs.DCq-bio.NCEmpiricalRecentJul 24, 2026

NUMA balancing hampering performance of spiking network simulations

Melissa Lober, Alp Inangu, Gorka Peraza Coppola, Dennis Terhorst +8 more

Turning off automatic NUMA balancing in simulation of large-scale spiking networks reduces energy consumption by 30%.

View →
cs.AREmpiricalRecentJul 22, 2026

DGNA: Dissecting GPU NUMA Architecture through Microbenchmarking and Data Analysis

Changxi Liu, Yun Chen, Trevor E. Carlson

This paper introduces DGNA, a methodology to unveil the Non-Uniform Memory Access (NUMA) architecture of GPU memory hierarchy through microbenchmarking and data analysis.

View →
cs.LGcs.AIcs.CLRecentMay 29, 2026

Finer Parameter Steps for Low-Rank PEFT: A Controlled Study with CP Tensor Adapters

Xinjue Wang, Xiuheng Wang, Yejun Zhang, Sergiy A. Vorobyov +2 more

The paper investigates whether using fine-grained, tensorized adapters (CP components) instead of standard LoRA ranks improves the accuracy-budget trade-off in PEFT, finding that while they fill budge…

View →
cs.LGcs.AIRecentMay 31, 2026

Hybrid Imbalanced Regression Through Unified Data-Level and Algorithm-Level Balancing

Shermin Shahbazi, Hossein Mohammadi, Mohsen Afsharchi

The paper proposes a unified hybrid framework that combines data-level and algorithm-level balancing to effectively address the challenge of imbalanced regression, significantly improving predictive p…

View →
cs.DCEmpiricalRecentJun 29, 2026

Spandana: Reconciling Strict SLOs with Low Cost under Fine-Grained Load Fluctuations

Dilina Dehigama, Shyam Jesalpura, Zeyu Xu, Marton Nemeth +3 more

The paper introduces Spandana, an architecture that decouples SLO enforcement from cost optimization in cloud-based online services, achieving high utilization, strict SLO adherence, and cost savings.

View →
cs.MScs.DCmath.NAEmpiricalRecentJul 8, 2026

Multiple Double Arithmetic on NVIDIA Tensor Cores

Howard Chen, Jan Verschelde

The paper presents a solution to enable multiple double arithmetic on NVIDIA A100 tensor cores, which are unsuited for branching operations caused by renormalization.

View →
cs.DSTheoreticalRecentJun 19, 2026

Online Stacking with a Few Load/Unload Points

Martin Olsen

A simple online algorithm is presented for the stacking problem to avoid shifts with a sufficient condition involving stacking area dimension, load/unload points, and maximum items.

View →
cs.DCEmpiricalRecentJul 2, 2026

Elasticity in Parallel Sparse Triangular Solve

Raphael S. Steiner, Christos K. Matzoros, Pál András Papp, Toni Böhnlein +1 more

This paper introduces Stale Synchronous Parallel mode of execution for parallel sparse triangular linear system solve and presents a scheduler that achieves geometric-mean speed-ups of 7-30% over Grow…

View →
cs.CRcs.ARcs.LGRecentMar 20, 2026

Hawkeye: Reproducing GPU-Level Non-Determinism

Erez Badash, Dan Boneh, Ilan Komargodski, Megha Srivastava

Hawkeye is a system that allows perfect, precision-preserving reproduction of GPU-level matrix multiplication operations on a CPU, enabling efficient and trustworthy third-party auditing of machine le…

View →
cs.CRRecentMay 6, 2026

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation

Marco Arazzi, Vignesh Kumar Kembu, Antonino Nocera, Stjepan Picek +1 more

The paper introduces NeWTral, a framework that restores safety alignment to specialized LLM adapters without sacrificing their domain-specific knowledge, achieving a significant reduction in attack su…

View →
cs.DCcs.LGEmpiricalRecentJun 29, 2026

GPU Parallelization Strategies for Forward and Backward Propagation in Shallow Neural Networks: A CUDA-Based Comparative Study

Rania Zitouni, Nadine Bousdjira, Sarah Hasnaoui, Amel Sadoun +1 more

This paper compares and optimizes CUDA strategies for a shallow neural network, achieving a 1.41x speedup on a large dataset.

View →
cs.LGmath.DGmath.OCEmpiricalRecentJun 28, 2026

Dead-Direction Conditioners: Gauge-Equivariant Preconditioning for Deep Networks

Tejas Pradeep Shirodkar

This paper introduces DDC, a Dead-Direction Conditioner that keeps a deep network's optimization on the symmetry quotient by conditioning the optimizer's state in the orbit decomposition of a $G$-inva…

View →
cs.DCEmpiricalRecentJun 29, 2026

Energy-Aware Scheduling for Serverless LLM Serving on Shared GPUs

Tianyu Wang, Gourav Rattihalli, Aditya Dhakal, Longfei Shangguan +1 more

This paper presents Festina, a profiling-guided, power-aware control plane for minimizing energy consumption in serverless large language model (LLM) serving.

View →
cs.ROEmpiricalRecentJul 17, 2026

A New Implementation of NeoSLAM and a Comparative Evaluation with RatSLAM

Joao Victor T. Borges, Fabio Coelho, Paulo Padrao, Jose Fuentes +3 more

This paper presents a new modular architecture for NeoSLAM using modern frameworks, achieving real-time execution and minimal data discarding. It also compares NeoSLAM and RatSLAM across three dataset…

View →
cs.ROEmpiricalRecentJul 23, 2026

FORGE-plus: Force-Budgeted Recovery for Contact-Rich Assembly with a Frozen LLM Supervisor

Kyupaeck Jeff Rah, Midum Oh

A two-layer framework using a large language model for force-conditioned reinforce learning with recovery maneuvers and force signatures.

View →