ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Feature-wise Linear Modulation (FiLM)”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

eess.AScs.SDEmpiricalRecentJun 21, 2026

Bridging Self-Supervised Learning and Speech Enhancement: A Wav2Vec2-Conditioned Framework

Shuubham Ojha, Carol Espy-Wilson

This paper conditions a diffusion-based speech enhancement model on wav2vec 2.0 features using Feature-wise Linear Modulation (FiLM), achieving competitive performance on VoiceBank-DEMAND and LibriMix…

View →
cs.CVcs.AIRecentMay 27, 2026

SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control

Zhida Zhang, Jie Ma, Zhan Peng, Haoxue Wu +4 more

SmartDirector is a novel framework that significantly improves cinematic video generation by using multiple keyframes to provide precise control over narrative structure and temporal pacing.

View →
cs.AIcs.MMcs.SDRecentMay 27, 2026

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation

Haitian Li, Yanghao Zhou, Heyan Huang, Liangji Chen +14 more

The paper introduces MTAVG-Bench 2.0, a new benchmark designed to diagnose high-level failure modes of cinematic expressiveness in multi-talker audio-video generation, showing that even advanced model…

View →
cs.LGRecentJun 1, 2026

Low-Pass Flow Matching

Francesco M. Ruscio, T. Konstantin Rusch

Low-Pass Flow Matching introduces a spectral bias into the flow matching process, allowing it to better model natural data by transitioning from a standard source spectrum to a frequency-decaying bias…

View →
cs.SDcs.AIcs.CVRecentJun 1, 2026

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

Jiashuo Yu, Yao Yao, Boyu Chen, Alex Wang

JenBridge is a novel, adaptive framework that generates high-fidelity, long-form video soundtracks, significantly improving narrative coherence and naturalness across scene transitions.

View →
eess.ASEmpiricalRecentJul 5, 2026

Noisy Environment Adaptation of Neural Speech Codec via Focal Mask and Noise Feature Separation

Shaokai Li, Weiping Tu, Yuhong Yang

The paper proposes FocalSE, a method for enhancing speech in neural speech codecs by performing feature denoising, separation, and recognition in the continuous embedding space.

View →
cs.CVcs.AIcs.CLRecentJun 1, 2026

AdaCodec: A Predictive Visual Code for Video MLLMs

Haowen Hou, Zhen Huang, Zheming Liang, Qingyi Si +7 more

AdaCodec introduces a predictive visual coding scheme for video MLLMs, significantly improving efficiency and performance by transmitting only inter-frame changes and full reference frames when necess…

View →
cs.LGmath.STstat.MLTheoreticalRecentJul 24, 2026

Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent

Peng Zhao

This paper proposes a method for handling overparameterized linear regression using early-stopped negative-shifted gradient descent, which allows for smooth filters and mixed-sign capabilities.

View →
cs.CVEmpiricalRecentJun 29, 2026

Real-Time Underwater Image Enhancement via Frequency-Guided Dual-Path Attention

Leshen Zhang, Ao Li, Ce Zhu

This paper proposes a lightweight underwater image enhancement framework with two components: MBRConv-DCT for injecting frequency priors during training and FGDPA for fusing spatial and spectral cues.

View →
cs.CVcs.AIcs.CRRecentMay 26, 2026

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

Pengzhen Chen, Yanwei Liu, Xiaoyan Gu, Antonios Argyriou +2 more

The paper introduces a novel third-order, rotation-invariant spherical bispectrum for watermarking panoramic images, enabling reliable watermark embedding and extraction under arbitrary 3D rotations.

View →
cs.CVEmpiricalRecentJun 18, 2026

Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models

Haoxuan Wu, Lai Man Po, Mengyang Liu, Kun Li +2 more

The paper introduces PRISM, a method for decoding preference signals from noisy latents using a lightweight Query-based Aggregation head and a frozen video diffusion backbone, achieving state-of-the-a…

View →
cs.CERecentMay 30, 2026

Graph Attention-Based Virtual Metrology for Film Deposition Processes in Semiconductor Manufacturing

Tao Han, Suk Ki Lee, Hyunwoong Ko

The paper proposes a graph attention-based virtual metrology framework that accurately predicts film thickness in semiconductor deposition by modeling structured, directional dependencies among hetero…

View →
cs.SDeess.ASEmpiricalRecentJul 26, 2026

Automatic Audio Equalization with Semantic Embeddings

Eloi Moliner, Vesa Välimäki, Konstantinos Drossos, Matti S. Hämäläinen

This paper proposes a data-driven method for automatic blind audio equalization using a deep neural network and semantic embeddings.

View →
cs.CVEmpiricalRecentJul 23, 2026

CLUIE: Clustering-Aware Recurrent Propagation with Local Structural Compensation for Underwater Image Enhancement

Kui Jiang, Zefan Feng, Laibin Chang, Yan Luo +2 more

A new framework, CRWKV, is proposed for underwater image enhancement using a Clustering-aware Semantic Dynamic Reordering and Dark-response Modulated Local Propagation methods.

View →
cs.SDcs.AIcs.MMRecentMay 27, 2026

EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction

Chong Jing, Zitong Lan, Junan Zhang, Zhizheng Wu

EigeNet introduces a geometry-informed multi-modal Transformer framework to achieve state-of-the-art few-shot novel view Room Impulse Response (RIR) prediction by effectively integrating spatial geome…

View →