20 results for “WavLM”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper introduces WavLM for vocal effort classification and improves performance through data augmentation and Gaussian-neighbor soft labels, achieving a new state-of-the-art on AVID.
Liwen Jing, Yisha Lu, Tingting Yang, Li Sun +4 more
The paper introduces SpikeWFM, a novel hybrid architecture combining spiking neural networks (SNNs) and transformers, which significantly improves the robustness and accuracy of wireless foundation mo…
Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu +2 more
The paper proposes VLM3, a simple, scalable method that demonstrates standard Vision Language Models (VLMs) can natively learn 3D understanding by focusing on architectural simplicity and specific dat…
The paper proposes a novel multimodal framework for session-based music recommendation that jointly models audio, lyric, and semantic content signals within a unified LLM-based sequential reasoning sy…
The paper introduces MLLM-Microscope, a system that analyzes the internal structure of multimodal large language models (MLLMs), finding that modality fusion significantly impacts the linearity and di…
Chanwoo Cho, Wooseok Kim, Yonglak Son, Young Seo Lee +1 more
The paper proposes Voltron, a framework for executing large language model inferences on multiple user-end devices at the edge to achieve higher accuracy and satisfy QoS requirements.
The paper proposes DLLM-VSR, a novel Diffusion Large Language Model framework for Visual Speech Recognition, achieving state-of-the-art performance by treating transcription as iterative masked denois…
Hao Zhang, Yiwen Zhao, Yixuan Zhang, Yiwen Shao +1 more
The paper introduces an agentic soundscape construction framework for controllable compositional audio generation, which makes explicit the scene planning, source selection, temporal layout, and rende…
Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang +3 more
The paper proposes ESRT, an edge-cloud framework that achieves state-of-the-art, bandwidth-efficient, and privacy-preserving many-to-many speech translation across 45 languages by splitting the model…
This paper conditions a diffusion-based speech enhancement model on wav2vec 2.0 features using Feature-wise Linear Modulation (FiLM), achieving competitive performance on VoiceBank-DEMAND and LibriMix…
This paper introduces BayLing-Duplex, a native full-duplex Speech Language Model that decides when to listen, speak, and stop without relying on an external Voice Activity Detection module.
Chong Li, Zhengdao Yu, Nelson Lossing, Thibaut Tachon +5 more
The paper introduces Mpipe, a method for multimodal-aware heterogeneous parallel scheduling in large-scale multimodal language model training, achieving significant speedups on Ascend 910C NPU cluster…
Ke-Han Lu, Keqi Deng, Ruchao Fan, Rui Zhao +1 more
The paper proposes SpeechKV, a method to compress speech sequences inside large language models using a learned pooling, maintaining performance and delivering decoding speedup.
This paper presents the topology-independent distributed multichannel Wiener filter (TI-dMWF) algorithm for distributed node-specific signal estimation in wireless acoustic sensor networks, enabling o…
This survey provides a comprehensive taxonomy and vulnerability-centric analysis of adversarial attacks targeting Multimodal Large Language Models (MLLMs), offering an explanatory framework for enhanc…
This paper evaluates biases in multimodal speech recognition by testing how pairing different faces with the same audio affects transcription accuracy, finding significant quality-of-service drops acr…
This paper introduces WanSong, a diffusion-based model for long-form, commercial-grade song generation that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems,…