20 results for “diffusion transformer”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper investigates the application of transformer-based diffusion models for simulating and reconstructing hydrological time series using data from six sites in North-East France.
Liting Gao, Yonggang Zhu, Yaru Chen, Dongyu Wang +4 more
This paper proposes a hybrid two-stage diffusion transformer architecture for instruction-guided audio editing, balancing performance and efficiency.
The paper systematically characterizes column-level activation sparsity across various diffusion model architectures, demonstrating that element-level sparsity metrics significantly overestimate the a…
This paper explores the effects of geometric perturbations on the stability of diffusion architectures, specifically Diffusion Transformers and Convolutional UNets, using a unified framework and contr…
Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu +5 more
SANA-Streaming introduces a novel, efficient framework that enables real-time, high-resolution streaming video-to-video editing by combining a hybrid diffusion transformer with specialized training an…
This paper introduces Spectral Attention and Graph Convolutional Attention (GCA) for denoising graphs, which outperforms linear attention and provably utilizes the input graph spectrum.
Ruotong Liao, Guowen Huang, Qing Cheng, Guangyao Zhai +5 more
TunerDiT introduces a training-free progressive steering method to enhance multi-event video generation using Diffusion Transformers, achieving state-of-the-art performance by explicitly managing even…
This paper proposes KroQuant, a post-training quantization method for diffusion transformers using learned Kronecker-structured invertible transforms, which reduces parameters, improves speed, and mai…
The paper proposes VRPO, a reinforcement learning-based optimization strategy that replaces static alignment losses in diffusion models, significantly improving both convergence and image fidelity.
Calvin Yeung, Prathyush Poduval, Ali Zakeri, Zhuowen Zou +1 more
The paper introduces residualized temporal Sparse Autoencoders (SAEs) to analyze the full spatiotemporal structure of activations generated during the iterative denoising process of diffusion models,…
The paper repurposes a pre-trained speech classifier as the backbone for diffusion generation, reducing the need for two separately trained models.
This paper conditions a diffusion-based speech enhancement model on wav2vec 2.0 features using Feature-wise Linear Modulation (FiLM), achieving competitive performance on VoiceBank-DEMAND and LibriMix…
Longxuan Yu, Yunshu Wu, Yu Fu, Siheng Xiong +4 more
The paper introduces DSL-LLaDA, a method that lightly adapts a pre-trained masked diffusion language model to perform continuous denoising in embedding space, significantly improving text generation q…
The paper proposes SafeDIG, a robust safety steering framework that adapts Diffusion Transformers for text-to-image generation by treating safety control as position-aware sparse feature transfer, ens…
The paper introduces DLM-SWAI, a training-free method that effectively steers diffusion language models (DLMs) toward desired textual styles or properties by biasing the token distribution at each den…
Paul Jünger, Justin Lovelace, Linxi Zhao, Dongyoung Go +1 more
The paper introduces SARDI, a novel, training-free framework that uses low-confidence 'lookahead' tokens generated during the denoising process of discrete diffusion language models to dynamically gui…
The paper proposes a novel global sketch-based watermarking technique for diffusion language models that controls the entire sequence's statistics, offering an order-agnostic and context-independent a…
Yijie Jin, Jiajun Xu, Yuxuan Liu, Chenkai Xu +7 more
This paper proposes Multi-Block Diffusion Language Models (MBD-LMs) for text generation, which are obtained by post-training Block Diffusion Language Models (BD-LMs) with Multi-block Teacher Forcing (…