ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “audio-visual speech recognition”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

eess.AScs.CVEmpiricalRecentJul 17, 2026

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze +18 more

The paper introduces Audio-Visual Flamingo (AV-Flamingo), an open-source audio-visual large language model designed for understanding and reasoning over long and complex real-world audio-visual videos…

View →
cs.SDcs.AIRecentMay 30, 2026

Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty

Zhou Yang, Yueyi Yang

This paper investigates if upper-face affective cues enhance audiovisual sentence recognition, especially when audio is degraded, finding that while mouth cues are crucial for robustness, upper-face c…

View →
cs.CLRecentMay 28, 2026

Your Multimodal Speech Model Says I Have a Face for Radio

Maya K. Nachesa, Vlad Niculae, Vagrant Gautam

This paper evaluates biases in multimodal speech recognition by testing how pairing different faces with the same audio affects transcription accuracy, finding significant quality-of-service drops acr…

View →
eess.ASEmpiricalRecentJul 3, 2026

CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds

Harshit Rajgarhia, Shuubham Ojha, Akhil Pothanapalli, Rachuri Lokesh +3 more

The paper introduces CaReCoS, a benchmark for multimodal reasoning over medical acoustic spectrograms, and evaluates the performance of vision and omni models, finding a maximum accuracy of 51.2%.

View →
cs.CVcs.AIcs.CRRecentApr 17, 2026

NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition

Junguang Yao, Wenye Liu, Stjepan Picek, Yue Zheng

NeuroLip proposes an event-based spatiotemporal framework for visual speaker recognition that achieves robust cross-scene generalization by capturing fine-grained lip dynamics, outperforming existing…

View →
cs.AIcs.CVeess.ASRecentMay 27, 2026

Diffusion Large Language Models for Visual Speech Recognition

Jeong Hun Yeo, Chae Won Kim, Hyeongseop Rha, Yong Man Ro

The paper proposes DLLM-VSR, a novel Diffusion Large Language Model framework for Visual Speech Recognition, achieving state-of-the-art performance by treating transcription as iterative masked denois…

View →
cs.CVcs.AIRecentMay 30, 2026

V-LynX: Token Interface Alignment for Video+X LLMs

Jungin Park, Jiyoung Lee, Kwanghoon Sohn

V-LynX is a framework that enhances Video LLMs by integrating new modalities into their existing token interface, achieving state-of-the-art performance across diverse video understanding tasks.

View →
eess.ASEmpiricalRecentJun 19, 2026

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi

Sujith Pulikodan, Agneedh Basu, Saurabh Kumar, Pranav Bhat +4 more

The paper introduces a new inclusive, multimodal Hindi ASR benchmark with real-world recordings and diverse demographic groups, enabling more robust and realistic evaluation.

View →
cs.SDcs.AIEmpiricalRecentJul 21, 2026

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

Cheng Siong Chin, Jianhua Zhang, Mohan Venkateshkumar

Caption Studio is a transparency-first speech and audio intelligence platform that provides automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle ge…

View →
eess.SPcs.SDEmpiricalRecentJul 18, 2026

Efficient Audio-Visual Event Recognition via Knowledge Distillation and Dynamic INT8 Quantization of a Hybrid Cross-Attention Network

Parinaz Binandeh Dehaghani, Danilo Pena, A. Pedro Aguiar

This paper proposes a compression framework for efficient audio-visual event recognition using transformer-based models, combining architectural compression, knowledge distillation, and dynamic INT8 q…

View →
cs.SDeess.ASEmpiricalRecentJul 26, 2026

Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

Ryo Magoshi, Jaeyoung Lee, Shinsuke Sakai, Tatsuya Kawahara

The study investigates the limitations of Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription using Grapheme-to-Phoneme (G2P) labels and proposes a new approach based on continuous Artic…

View →
eess.ASEmpiricalRecentJul 8, 2026

UBG-Net: An Uncertainty-aware Bayesian Gating Network for Robust Audio-Visual Speech Recognition

Jinjie Fu, Hang Chen, Wu Guo, Zhijun Zhang +2 more

This paper proposes a framework, UBG-Net, for robust audio-visual speech recognition using a Modality Uncertainty-aware Bayesian Fusion mechanism and Distribution Uncertainty-aware Hierarchical Voting…

View →
cs.CVcs.SDEmpiricalRecentJul 3, 2026

$C^3$ASD: Multi-Level Consistency-Driven Representation Learning

Jin Hong, Jisoo Park, Junseok Kwon

This paper proposes a multi-level consistency-driven framework, $C^3$ASD, for robust active speaker detection in video, addressing the limitations of recent audio-visual fusion methods.

View →
eess.AScs.SDEmpiricalRecentJul 7, 2026

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

Ke-Han Lu, Keqi Deng, Ruchao Fan, Rui Zhao +1 more

The paper proposes SpeechKV, a method to compress speech sequences inside large language models using a learned pooling, maintaining performance and delivering decoding speedup.

View →
eess.ASEmpiricalRecentJul 20, 2026

The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

Xuanji He, Gaoyang Dong, Xiaoxiao Li, Minchuan Chen +1 more

The paper presents the tttAI system for time-stamped speaker-attributed speech recognition in smart-glasses recordings, achieving a tcpCER of 7.10% on Track 1 and 34.04% on Track 2.

View →
cs.CLcs.AIeess.ASRecentMay 31, 2026

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more

PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…

View →
cs.SDcs.AIRecentJun 1, 2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu +21 more

MOSS-Audio is a unified audio-language model designed for comprehensive understanding of speech, environmental sounds, and music, achieving strong performance across various audio-grounded tasks.

View →
eess.ASEmpiricalRecentJul 19, 2026

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

Xiaoyu Yang, Xuenan Xu, Wenyi Yu, Siyin Wang +9 more

The paper proposes SALMONN-2, an ALLM built on a unified SSL encoder, and presents a multi-layer feature fusion adapter to better exploit hierarchical SSL encoder representations. It also explores mul…

View →