20 results for “Feature-wise Linear Modulation (FiLM)”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper conditions a diffusion-based speech enhancement model on wav2vec 2.0 features using Feature-wise Linear Modulation (FiLM), achieving competitive performance on VoiceBank-DEMAND and LibriMix…
Zhida Zhang, Jie Ma, Zhan Peng, Haoxue Wu +4 more
SmartDirector is a novel framework that significantly improves cinematic video generation by using multiple keyframes to provide precise control over narrative structure and temporal pacing.
Haitian Li, Yanghao Zhou, Heyan Huang, Liangji Chen +14 more
The paper introduces MTAVG-Bench 2.0, a new benchmark designed to diagnose high-level failure modes of cinematic expressiveness in multi-talker audio-video generation, showing that even advanced model…
Low-Pass Flow Matching introduces a spectral bias into the flow matching process, allowing it to better model natural data by transitioning from a standard source spectrum to a frequency-decaying bias…
JenBridge is a novel, adaptive framework that generates high-fidelity, long-form video soundtracks, significantly improving narrative coherence and naturalness across scene transitions.
The paper proposes FocalSE, a method for enhancing speech in neural speech codecs by performing feature denoising, separation, and recognition in the continuous embedding space.
Haowen Hou, Zhen Huang, Zheming Liang, Qingyi Si +7 more
AdaCodec introduces a predictive visual coding scheme for video MLLMs, significantly improving efficiency and performance by transmitting only inter-frame changes and full reference frames when necess…
This paper proposes a method for handling overparameterized linear regression using early-stopped negative-shifted gradient descent, which allows for smooth filters and mixed-sign capabilities.
This paper proposes a lightweight underwater image enhancement framework with two components: MBRConv-DCT for injecting frequency priors during training and FGDPA for fusing spatial and spectral cues.
Pengzhen Chen, Yanwei Liu, Xiaoyan Gu, Antonios Argyriou +2 more
The paper introduces a novel third-order, rotation-invariant spherical bispectrum for watermarking panoramic images, enabling reliable watermark embedding and extraction under arbitrary 3D rotations.
Haoxuan Wu, Lai Man Po, Mengyang Liu, Kun Li +2 more
The paper introduces PRISM, a method for decoding preference signals from noisy latents using a lightweight Query-based Aggregation head and a frozen video diffusion backbone, achieving state-of-the-a…
The paper proposes a graph attention-based virtual metrology framework that accurately predicts film thickness in semiconductor deposition by modeling structured, directional dependencies among hetero…
This paper proposes a data-driven method for automatic blind audio equalization using a deep neural network and semantic embeddings.
Kui Jiang, Zefan Feng, Laibin Chang, Yan Luo +2 more
A new framework, CRWKV, is proposed for underwater image enhancement using a Clustering-aware Semantic Dynamic Reordering and Dark-response Modulated Local Propagation methods.
EigeNet introduces a geometry-informed multi-modal Transformer framework to achieve state-of-the-art few-shot novel view Room Impulse Response (RIR) prediction by effectively integrating spatial geome…