20 results for “diffusion generation”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper introduces WanSong, a diffusion-based model for long-form, commercial-grade song generation that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems,…
The paper repurposes a pre-trained speech classifier as the backbone for diffusion generation, reducing the need for two separately trained models.
Sergio Rozada, Yiming Qin, Manuel Madeira, Pascal Frossard +1 more
This paper introduces DiPhon, a diffusion framework for size-scalable graph generation, using a continuous diffusion process on the graphon space and a discretized graph-level process.
Yijie Jin, Jiajun Xu, Yuxuan Liu, Chenkai Xu +7 more
This paper proposes Multi-Block Diffusion Language Models (MBD-LMs) for text generation, which are obtained by post-training Block Diffusion Language Models (BD-LMs) with Multi-block Teacher Forcing (…
The paper introduces DLM-SWAI, a training-free method that effectively steers diffusion language models (DLMs) toward desired textual styles or properties by biasing the token distribution at each den…
Ruotong Liao, Guowen Huang, Qing Cheng, Guangyao Zhai +5 more
TunerDiT introduces a training-free progressive steering method to enhance multi-event video generation using Diffusion Transformers, achieving state-of-the-art performance by explicitly managing even…
The paper introduces FTDiff, a reinforcement learning fine-tuning framework that efficiently generates high-quality, drug-like molecules constrained by a target protein structure, outperforming existi…
This paper introduces Align4D, a framework for generating coherent video-3D pairs using any-modal input, achieving state-of-the-art quality and consistency in X-to-4D generation.
Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat +2 more
This paper introduces Flex-Forcing, a framework for video generation that enables a model to operate under both bidirectional and autoregressive generation regimes, achieving better video quality and…
The paper introduces MDM-VGB, a reward-guided sampler for Masked Diffusion Models, which extends the Jerrum-Sinclair backtracking Markov chain to a masked-state graph for effective high-reward generat…
Sijie Wang, Zhengyu Qing, Zhiqiang Tan, Yiming Yin +5 more
DigenRL is a disaggregated RL framework for diffusion-based generative LLMs that achieves 1.56-2.10x throughput improvements over state-of-the-art diffusion RL systems.
The paper proposes a novel global sketch-based watermarking technique for diffusion language models that controls the entire sequence's statistics, offering an order-agnostic and context-independent a…
Fei Deng, Yanwu Xu, Zhipeng Bao, Zhixing Zhang +3 more
BlazeEdit is a highly efficient, generalist image-to-image diffusion model designed for on-device deployment, consolidating multiple editing tasks into a compact 195M parameter model that runs quickly…
Haoxuan Wu, Lai Man Po, Mengyang Liu, Kun Li +2 more
The paper introduces PRISM, a method for decoding preference signals from noisy latents using a lightweight Query-based Aggregation head and a frozen video diffusion backbone, achieving state-of-the-a…
The paper introduces SynCity 3000, a framework for generating large, coherent 3D scenes using a convolutional generator, addressing the scarcity of 3D scene data for training.
The paper introduces Strong Stochastic Flow Maps (SSFMs), a novel framework that directly learns the strong solution map of additive-noise Stochastic Differential Equations (SDEs), enabling few-step s…
Introduces Expanding Generative Flows (EFlows) and Expanding Flow Maps (EFMs) for generating outputs of varying sizes in continuous and discrete state spaces.
Yue Li, Linying Xue, Kaiqing Lin, Hanyu Quan +4 more
The paper proposes AEGIS, a novel diffusion-guided method for injecting adversarial perturbations into the latent space to create generalizable and robust defenses against advanced facial deepfake man…
Yangtian Zhang, Zhe Wang, Arthur Gretton, Rex Ying +3 more
The paper introduces the Insertion Process (IP), a novel stochastic generative model that learns variable-length, non-monotonic sequence generation by explicitly modeling the insertion order of tokens…