ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

14 results for “Understanding of spatial audio and immersive educational environments.”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDEmpiricalRecentJul 8, 2026

EscFOA: Enhancing Spatial Learning for Visually Impaired Learners via Generative Spatial Audio in 360-Degree Educational Environments

Ziyu Luo, Xiaowei Dai, Siying Zhu, Xiaoming Chen

This paper proposes EscFOA, a framework that uses geometry-aware spatial audio to enhance immersive educational environments for visually impaired learners, outperforming conventional audio methods.

View →
eess.AScs.SDEmpiricalRecentJun 23, 2026

Evaluation of Headrest-Integrated Loudspeakers for Enhanced Spatial Audio Immersion in Automotive Cabins

Martin Wolters, Jacobo Giralt, Harald Mundt, Arijit Biswas

This paper conducts subjective assessments to evaluate the preference and spatial audio attributes of headrest-integrated speakers for immersive audio scenarios in automobiles.

View →
eess.AScs.AIcs.CLRecentMay 29, 2026

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

Jun-Hak Yun, Seung-Bin Kim, Seong-Whan Lee

ImmersiveTTS is an environment-aware text-to-speech model that generates natural speech seamlessly integrated within environmental contexts by explicitly modeling cross-modal interactions, achieving s…

View →
cs.SDEmpiricalRecentJul 23, 2026

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision

Hao Zhang, Yiwen Zhao, Yixuan Zhang, Yiwen Shao +1 more

The paper introduces an agentic soundscape construction framework for controllable compositional audio generation, which makes explicit the scene planning, source selection, temporal layout, and rende…

View →
cs.SDcs.AIcs.MMRecentMay 27, 2026

EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction

Chong Jing, Zitong Lan, Junan Zhang, Zhizheng Wu

EigeNet introduces a geometry-informed multi-modal Transformer framework to achieve state-of-the-art few-shot novel view Room Impulse Response (RIR) prediction by effectively integrating spatial geome…

View →
cs.MMcs.HCEmpiricalRecentJul 20, 2026

Toward Site-Aware MR Art Exhibitions: A SLAM-Based Deployment Pipeline for Spatial Coherence and Exhibition Experience

Yawei Zhao, Yuming Zhu, Hao Li, Yuqi Liang +3 more

This paper presents a practical pipeline for designing and deploying large-scale Mixed Reality art exhibitions using SLAM-based alignment, and evaluates its impact on technical stability and user expe…

View →
cs.SDstat.APEmpiricalRecentJun 23, 2026

Statistical validation and full-sphere extension of a Bayesian model for human static sound localisation

Roberto Barumerli, Fabian Brinkmann, Emanuele Zanoni, Anton Hoyer +2 more

This paper validates a Bayesian sound localisation model using statistical methods and compares four HRTF template interpolation methods.

View →
cs.SDcs.AIEmpiricalRecentJun 26, 2026

From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection

Stefano Giacomelli, Stefano Damiano, Claudia Rinaldi, Fabio Graziosi +1 more

This paper proposes AT2SELD framework to extend pretrained GP-AT models for spatially grounded Sound Event Localization and Detection using spectral FOA descriptors, NAS, and calibration.

View →
eess.ASEmpiricalRecentJul 18, 2026

NABEATs: Noise-Aware Audio Representation Learning

Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern, Christoph Boeddeker +2 more

The paper introduces Noise-Aware BEATs (NABEATs), a noise-aware audio self-supervised learning framework that estimates clean BEATs representations from noisy audio signals using an auxiliary referenc…

View →
cs.SDcs.AIcs.CLRecentMay 28, 2026

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings

Yonggang Zhu, Liting Gao, Aidong Men, Wenwu Wang

The paper introduces COMET, a novel PLS-SVD framework, to analyze the audio-text modality gap in CLAP models, showing that shared concepts are captured by a small subset of axes, and proposes a spectr…

View →
cs.SDcs.AIRecentJun 1, 2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu +21 more

MOSS-Audio is a unified audio-language model designed for comprehensive understanding of speech, environmental sounds, and music, achieving strong performance across various audio-grounded tasks.

View →
cs.HCcs.SDEmpiricalRecentJul 25, 2026

Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models

Shuoyang Jasper Zheng, Anna Xambó Sedó, Nick Bryan-Kinns

This paper proposes material explainability as a way to make AI models accessible and inclusive design materials for artists, designers, and makers, using a case study of building a repository for neu…

View →