~ similar to 2607.14908· 19 results
VideoMLA introduces a novel Multi-Head Latent Attention (MLA) mechanism that replaces per-head KV caches with a shared low-rank content latent, significantly reducing memory and improving throughput f…
The paper systematically characterizes column-level activation sparsity across various diffusion model architectures, demonstrating that element-level sparsity metrics significantly overestimate the a…
Jiayi Wu, Haoming Cai, Cornelia Fermuller, Christopher Metzler +1 more
Real2SAM2Real introduces a framework that uses explicit 3D caches, derived from 3D lifting models, to provide robust geometric guidance to Video Diffusion Models, significantly improving spatiotempora…
Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu +5 more
SANA-Streaming introduces a novel, efficient framework that enables real-time, high-resolution streaming video-to-video editing by combining a hybrid diffusion transformer with specialized training an…
Zhuoren Ye, Tianyu Wo, Dinghao Xue, Mingming Zhang +3 more
This paper proposes CrossPool, a serving engine for cold Machine Learning Models (LLMs) that separates weights and KV-cache into two GPU memory pools to improve GPU memory utilization and long-context…
Hawkeye is a system that allows perfect, precision-preserving reproduction of GPU-level matrix multiplication operations on a CPU, enabling efficient and trustworthy third-party auditing of machine le…
This paper characterizes the performance of vision-language model inference on the Qualcomm SM8750 using FastVLM-0.5B as a case study, showing significant speedups and energy savings for different pha…
Chunan Shi, Yilei Chen, Yilin Chen, Xupeng Miao +1 more
The paper proposes AsymCache, a computation-latency-aware KV cache management system that optimizes LLM inference by aligning cache eviction decisions with GPU attention kernel performance, significan…
This paper presents Roomie, an architecture for orchestrating model serving in multi-model scenarios on GPUs, reducing SLO violations by predicting and avoiding kernel-level interference.
This paper reformulates cascaded second-order filtering as a block-tridiagonal linear system and develops parallel solution algorithms, achieving high performance on SIMD cores, multi-core CPUs, and G…
Shuo Huai, Hao Kong, Xiangzhong Luo, Shiqing Li +4 more
This paper addresses the obstacles of using Crossbar-based In-Memory Processing (IMP) accelerators for deep neural networks (DNNs) by reusing bit-shift units for multiplication, applying pruning metho…
First implementation of a 3D Gaussian renderer on an Intelligence Processing Unit (IPU) with 1,472 independent tiles using only on-chip SRAM.
Tessera introduces a novel hardware architecture that achieves secure, near-line-rate weight streaming for DNNs on UMA edge accelerators by performing cache-line granularity decryption during DRAM fet…
The paper presents a field study on deploying large inference workloads on a non-GPU AI accelerator and identifies eight categories of limitations.
Hubert Dymarkowski, Xingjian Fu, Rappy Saha, Jude Haris +1 more
This paper presents FlexViT, a reconfigurable FPGA accelerator for efficient Vision Transformer (ViT) inference on edge devices, achieving up to 2.74x speedup on accelerator-executed layers.
This paper investigates the potential of real-world Processing-in-Memory (PIM) architectures, specifically using UPMEM, to accelerate cryptographic algorithms, demonstrating that distributing computat…