ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “WavLM”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDcs.LGEmpiricalRecentJun 25, 2026

Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings

Zahra Omidi, John H. L. Hansen

This paper introduces WavLM for vocal effort classification and improves performance through data augmentation and Gaussian-neighbor soft labels, achieving a new state-of-the-art on AVID.

View →
eess.SPcs.AIcs.LGRecentMay 28, 2026

SpikeWFM: Spiking-Aided Wireless Foundation Model for Robust Channel Prediction

Liwen Jing, Yisha Lu, Tingting Yang, Li Sun +4 more

The paper introduces SpikeWFM, a novel hybrid architecture combining spiking neural networks (SNNs) and transformers, which significantly improves the robustness and accuracy of wireless foundation mo…

View →
cs.CVcs.AIRecentMay 28, 2026

VLM3: Vision Language Models Are Native 3D Learners

Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu +2 more

The paper proposes VLM3, a simple, scalable method that demonstrates standard Vision Language Models (VLMs) can natively learn 3D understanding by focusing on architectural simplicity and specific dat…

View →
cs.IRcs.AIcs.LGRecentMay 28, 2026

Multimodal Music Recommendation System using LLMs

Srikar Prabhas Kandagatla, Sreehitha R. Narayana, Chandana Magapu, Swetha Mohan +5 more

The paper proposes a novel multimodal framework for session-based music recommendation that jointly models audio, lyric, and semantic content signals within a unified LLM-based sequential reasoning sy…

View →
cs.CLcs.AIRecentMay 30, 2026

MLLM-Microscope: Unlocking Hidden Structure Within Multimodal Large Language Models

Ravil Mussabayev, Rustam Mussabayev

The paper introduces MLLM-Microscope, a system that analyzes the internal structure of multimodal large language models (MLLMs), finding that modality fusion significantly impacts the linearity and di…

View →
cs.DCEmpiricalRecentJul 8, 2026

Voltron: Enabling Elastic Multi-Device Execution of LLM Inference for Empowered Edge Intelligence

Chanwoo Cho, Wooseok Kim, Yonglak Son, Young Seo Lee +1 more

The paper proposes Voltron, a framework for executing large language model inferences on multiple user-end devices at the edge to achieve higher accuracy and satisfy QoS requirements.

View →
cs.AIcs.CVeess.ASRecentMay 27, 2026

Diffusion Large Language Models for Visual Speech Recognition

Jeong Hun Yeo, Chae Won Kim, Hyeongseop Rha, Yong Man Ro

The paper proposes DLLM-VSR, a novel Diffusion Large Language Model framework for Visual Speech Recognition, achieving state-of-the-art performance by treating transcription as iterative masked denois…

View →
cs.SDEmpiricalRecentJul 23, 2026

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision

Hao Zhang, Yiwen Zhao, Yixuan Zhang, Yiwen Shao +1 more

The paper introduces an agentic soundscape construction framework for controllable compositional audio generation, which makes explicit the scene planning, source selection, temporal layout, and rende…

View →
cs.AIRecentMay 27, 2026

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang +3 more

The paper proposes ESRT, an edge-cloud framework that achieves state-of-the-art, bandwidth-efficient, and privacy-preserving many-to-many speech translation across 45 languages by splitting the model…

View →
eess.AScs.SDEmpiricalRecentJun 21, 2026

Bridging Self-Supervised Learning and Speech Enhancement: A Wav2Vec2-Conditioned Framework

Shuubham Ojha, Carol Espy-Wilson

This paper conditions a diffusion-based speech enhancement model on wav2vec 2.0 features using Feature-wise Linear Modulation (FiLM), achieving competitive performance on VoiceBank-DEMAND and LibriMix…

View →
cs.CLeess.ASEmpiricalRecentJun 12, 2026

BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

Qingkai Fang, Shoutao Guo, Yang Feng

This paper introduces BayLing-Duplex, a native full-duplex Speech Language Model that decides when to listen, speak, and stop without relying on an external Voice Activity Detection module.

View →
cs.DCEmpiricalRecentJul 3, 2026

HyperParallel-Mpipe: A Composable Algebra System for Optimizing MLLM Training over Supernode Clusters

Chong Li, Zhengdao Yu, Nelson Lossing, Thibaut Tachon +5 more

The paper introduces Mpipe, a method for multimodal-aware heterogeneous parallel scheduling in large-scale multimodal language model training, achieving significant speedups on Ascend 910C NPU cluster…

View →
eess.AScs.SDEmpiricalRecentJul 7, 2026

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

Ke-Han Lu, Keqi Deng, Ruchao Fan, Rui Zhao +1 more

The paper proposes SpeechKV, a method to compress speech sequences inside large language models using a learned pooling, maintaining performance and delivering decoding speedup.

View →
eess.AScs.ITeess.SPTheoreticalRecentJul 6, 2026

Distributed Multichannel Wiener Filtering for Topology-Unconstrained Wireless Acoustic Sensor Networks

Paul Didier, Pourya Behmandpoor, Henri Gode, Toon van Waterschoot +3 more

This paper presents the topology-independent distributed multichannel Wiener filter (TI-dMWF) algorithm for distributed node-specific signal estimation in wireless acoustic sensor networks, enabling o…

View →
cs.CRcs.AIRecentMar 30, 2026

Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey

Bhavuk Jain, Sercan Ö. Arık, Hardeo K. Thakur

This survey provides a comprehensive taxonomy and vulnerability-centric analysis of adversarial attacks targeting Multimodal Large Language Models (MLLMs), offering an explanatory framework for enhanc…

View →
cs.CLRecentMay 28, 2026

Your Multimodal Speech Model Says I Have a Face for Radio

Maya K. Nachesa, Vlad Niculae, Vagrant Gautam

This paper evaluates biases in multimodal speech recognition by testing how pairing different faces with the same audio affects transcription accuracy, finding significant quality-of-service drops acr…

View →
eess.AScs.CVEmpiricalRecentJul 16, 2026

WanSong v1.0 Technical Report

Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou

This paper introduces WanSong, a diffusion-based model for long-form, commercial-grade song generation that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems,…

View →