20 results for “NUMA balancing”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang +3 more
This paper proposes a preconditioning layer for stable weight conditioning in LLM training.
Turning off automatic NUMA balancing in simulation of large-scale spiking networks reduces energy consumption by 30%.
This paper introduces DGNA, a methodology to unveil the Non-Uniform Memory Access (NUMA) architecture of GPU memory hierarchy through microbenchmarking and data analysis.
Xinjue Wang, Xiuheng Wang, Yejun Zhang, Sergiy A. Vorobyov +2 more
The paper investigates whether using fine-grained, tensorized adapters (CP components) instead of standard LoRA ranks improves the accuracy-budget trade-off in PEFT, finding that while they fill budge…
The paper proposes a unified hybrid framework that combines data-level and algorithm-level balancing to effectively address the challenge of imbalanced regression, significantly improving predictive p…
Dilina Dehigama, Shyam Jesalpura, Zeyu Xu, Marton Nemeth +3 more
The paper introduces Spandana, an architecture that decouples SLO enforcement from cost optimization in cloud-based online services, achieving high utilization, strict SLO adherence, and cost savings.
The paper presents a solution to enable multiple double arithmetic on NVIDIA A100 tensor cores, which are unsuited for branching operations caused by renormalization.
A simple online algorithm is presented for the stacking problem to avoid shifts with a sufficient condition involving stacking area dimension, load/unload points, and maximum items.
This paper introduces Stale Synchronous Parallel mode of execution for parallel sparse triangular linear system solve and presents a scheduler that achieves geometric-mean speed-ups of 7-30% over Grow…
Hawkeye is a system that allows perfect, precision-preserving reproduction of GPU-level matrix multiplication operations on a CPU, enabling efficient and trustworthy third-party auditing of machine le…
The paper introduces NeWTral, a framework that restores safety alignment to specialized LLM adapters without sacrificing their domain-specific knowledge, achieving a significant reduction in attack su…
Rania Zitouni, Nadine Bousdjira, Sarah Hasnaoui, Amel Sadoun +1 more
This paper compares and optimizes CUDA strategies for a shallow neural network, achieving a 1.41x speedup on a large dataset.
This paper introduces DDC, a Dead-Direction Conditioner that keeps a deep network's optimization on the symmetry quotient by conditioning the optimizer's state in the orbit decomposition of a $G$-inva…
This paper presents Festina, a profiling-guided, power-aware control plane for minimizing energy consumption in serverless large language model (LLM) serving.
This paper presents a new modular architecture for NeoSLAM using modern frameworks, achieving real-time execution and minimal data discarding. It also compares NeoSLAM and RatSLAM across three dataset…
A two-layer framework using a large language model for force-conditioned reinforce learning with recovery maneuvers and force signatures.