ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Parallel programming, MPI, HPC environments”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.DCEmpiricalRecentJun 30, 2026

Performance Analysis in Parallel Programming Education: A Comparative Usability Study

Anna-Lena Roth, David James, Jonas Posner, Michael Kuhn

The paper introduces EduMPI, a learning support tool for simplifying cluster usage and performance analysis of MPI parallel programs for students.

View →
cs.DCEmpiricalRecentJun 29, 2026

Towards Transparent Checkpointing with AI-driven Code Generation

Hai Duc Nguyen, Tekin Bicer, Kyle Chard, Ian Foster +1 more

Researchers used a large language model to generate checkpoint/restart code for MPI applications, achieving comparable efficiency to human-engineered solutions.

View →
physics.comp-phcs.DCphysics.flu-dynEmpiricalRecentJul 8, 2026

Scaling WaterLily.jl with MPI and an improved geometric multigrid solver

Bernat Font, Marin Lauber, Tzu-Yao Huang, Gabriel D. Weymouth

The paper presents improvements to the performance and scalability of WaterLily.jl, a scale-resolving incompressible flow solver, through the addition of MPI-based parallelism and optimizations to the…

View →
cs.DCEmpiricalRecentJul 21, 2026

A User-oriented Portable, Reproducible, and Scalable Software Ecosystem

Alfio Lazzaro, Utz-Uwe Haus, Sandrine Charousset, Nina Mujkanovic

This paper presents a software ecosystem enabling consistent development environments for running workflows across diverse hardware platforms.

View →
cs.DCEmpiricalRecentJun 30, 2026

An Empirical Analysis of High-Performance Computing Education in Germany

Anna-Lena Roth, Jonas Posner

This paper assesses HPC education at 102 academic institutions in Germany, identifying 178 HPC-related courses and evaluating their competency coverage and curricular placement, as well as examining l…

View →
cs.DCEmpiricalRecentJun 26, 2026

How far does a random forest generalize from a 54-run LAMMPS+SPICA benchmark?

Dennis Alves Pedersen, Paulo Henrique Leme Ramalho, Fábio Andrijauskas

This paper investigates the use of a Random Forest surrogate model to predict molecular dynamics workload performance and recommend optimal hybrid MPI+OpenMP configurations without exhaustive benchmar…

View →
cs.DCEmpiricalRecentJun 23, 2026

BiJuTy: An Interactive HPC-Aware Big Data Cluster Lifecycle Manager and Performance Assessment Utility for JupyterHub

Apurv Deepak Kulkarni, Jan Frenzel, Siavash Ghiasvand

BiJuTy is a user-friendly solution for executing complex big data processing workflows on high-performance computing systems within the Jupyter ecosystem.

View →
cs.DCEmpiricalRecentJul 2, 2026

Elasticity in Parallel Sparse Triangular Solve

Raphael S. Steiner, Christos K. Matzoros, Pál András Papp, Toni Böhnlein +1 more

This paper introduces Stale Synchronous Parallel mode of execution for parallel sparse triangular linear system solve and presents a scheduler that achieves geometric-mean speed-ups of 7-30% over Grow…

View →
cs.ETcs.DCEmpiricalRecentJul 21, 2026

Examining QRMI as a Unified Interface for Quantum-HPC Integration

Thomas Badts, Tim Boyle, Claudio Carvalho, Antonio Córcoles +24 more

The paper presents the Quantum Resource Management Interface (QRMI) as a standardized, vendor-agnostic middleware layer for integrating quantum resources into high-performance computing environments,…

View →
cs.PFcs.ARcs.DCRecentMay 27, 2026

Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory

Myeong Jun Jo

The paper introduces Rotary GPU, an exploratory execution approach demonstrating that large Mixture-of-Experts models can be run locally on consumer GPUs with limited VRAM, achieving usable decode thr…

View →
cs.DCEmpiricalRecentJun 19, 2026

rush: Scalable Asynchronous Distributed Computing via Shared State in R

Marc Becker, Bernd Bischl

An R package named rush is introduced, which provides a shared-state coordination layer for asynchronously parallelized iterative algorithms using a Redis database.

View →
cs.GRcs.CVcs.DCEmpiricalRecentJul 17, 2026

Rendering 3D Gaussians on a Graph Processor

Nicholas Fry, Ignacio Alzugaray, Mark Pupilli, Paul H. J. Kelly +1 more

First implementation of a 3D Gaussian renderer on an Intelligence Processing Unit (IPU) with 1,472 independent tiles using only on-chip SRAM.

View →
cs.PLcs.CCcs.FLRecentMay 30, 2026

Grid Programs: A Two-Dimensional, Variable-Free Model of Computation

Ezequiel López-Rubio

The paper introduces Grid Programs, a novel, Turing-complete model of computation where programs are two-dimensional arrangements of instructions, fundamentally departing from linear code structures.

View →
cs.DCEmpiricalRecentJun 29, 2026

StreamGuard: Low-Overhead Resilience for Real-time HPC Data Streams

Hai Duc Nguyen, Bogdan Nicolae, Tekin Bicer, Amal Gueroudji +3 more

This paper presents two techniques, dynamic checkpointing and progress-aware load redistribution, to maintain forward progress and balanced execution in real-time scientific workflows using the produc…

View →