ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “accelerator”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.AREmpiricalRecentJul 27, 2026

A Heterogeneous Neural Network Accelerator for End-to-End Multitask RF Signal Recognition

Zhifan Song, Haralampos-G. Stratigopoulos, Hassan Aboushady

This paper proposes a heterogeneous neural network accelerator for multi-task RF signal recognition, achieving high accuracy and low latency for automatic modulation recognition, hardware-Trojan cover…

View →
cs.PFcs.ARcs.DCRecentMay 27, 2026

Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory

Myeong Jun Jo

The paper introduces Rotary GPU, an exploratory execution approach demonstrating that large Mixture-of-Experts models can be run locally on consumer GPUs with limited VRAM, achieving usable decode thr…

View →
cs.ARRecentJun 1, 2026

CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees

Lorenzo Leone, Philip Wiese, Gamze İslamoğlu, Michael Rogenmoser +3 more

The paper introduces Chimera, a highly efficient and scalable MCU designed for ultra-low-power edge AI inference, achieving 3.1 TOPS/W by integrating a dedicated transformer accelerator and a QoS-guar…

View →
cs.ARcs.PFRecentMay 30, 2026

Regular-Activation Concentration: Characterizing Column-Level Output Sparsity Across Diffusion Model Architectures

Dazhi Yang, Shafayat Mowla Anik, Byeong Kil Lee, Jeeho Ryoo

The paper systematically characterizes column-level activation sparsity across various diffusion model architectures, demonstrating that element-level sparsity metrics significantly overestimate the a…

View →
cs.ARcs.ETRecentJun 4, 2026

Space-CIM: Enabling Compute-In-Memory Accelerators for Thermally-Constrained Space Platforms

Sohan Salahuddin Mugdho, Md. Shahedul Hasan, Cheng Wang

This paper investigates the thermal constraints of deploying AI compute infrastructure in space, comparing GPUs and compute-in-memory (CIM) accelerators using a co-design methodology.

View →
eess.AScs.SDNEWEmpiricalJul 28, 2026

faster-enhancer.c: A Dependency-Free int8 Runtime for Streaming Speech Enhancement on Commodity CPUs

Gyeongmin Kim

This paper implements and measures the cost of running a streaming speech enhancer on a CPU, achieving a 3.3x speedup compared to the fp32 ONNX Runtime graph.

View →
cs.ROcs.AIcs.DCRecentMay 27, 2026

CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive Control

Antoonio Buo, Vittorio Cammarota, Michele Avagnale, Pierluigi Arpenti +2 more

The paper introduces CA-AC-MPC, a CUDA-accelerated variant of Actor-Critic Model Predictive Control, which significantly reduces the training and inference latency of AC-MPC while maintaining state-of…

View →
cs.ROcs.AIcs.CVEmpiricalRecentJun 11, 2026

Mana: Dexterous Manipulation of Articulated Tools

Zhao-Heng Yin, Guanya Shi, Pieter Abbeel, C. Karen Liu

This paper presents Mana, a sim-to-real framework for dexterous articulated tool manipulation.

View →
cs.ARcs.ETcs.LGEmpiricalRecentJul 22, 2026

Leveraging ECRAM for Edge Continual Learning

Nabila Tasnim, Haoran Liu, Qing Cao, Saugata Ghose

The paper proposes CLASP, an end-to-end system with IMC acceleration for continual learning on edge platforms, addressing challenges of noisy computation and poor support for resource-efficient traini…

View →
cs.CRcs.ARcs.LGRecentMar 20, 2026

Hawkeye: Reproducing GPU-Level Non-Determinism

Erez Badash, Dan Boneh, Ilan Komargodski, Megha Srivastava

Hawkeye is a system that allows perfect, precision-preserving reproduction of GPU-level matrix multiplication operations on a CPU, enabling efficient and trustworthy third-party auditing of machine le…

View →
cs.DCEmpiricalRecentJul 17, 2026

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

Siyuan Shen, Anton Korzh, John Bachan, Tiancheng Chen +9 more

This paper explores methods to reduce latency in GPU collective communications for large language model inference, achieving near-optimal designs with barrier-free synchronization and efficient use of…

View →
cs.DCEmpiricalRecentJul 9, 2026

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend

Zheng Yu

The paper presents a field study on deploying large inference workloads on a non-GPU AI accelerator and identifies eight categories of limitations.

View →