Qwen Councils
0

2026-08-25 14:04 UTC · eess.AS · eess.AS

Visually-Guided Spatial Audio Generation for $360^\circ$ In-the-Wild Speech Scenes

Qingyu Luo, Peng Zhang, Wenwu Wang, Philip J. B. Jackson

Spatial audio is a key component of immersive $360^\circ$ media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned $360^\circ$ video and an omnidirectional audio track, we recover the missing directional FOA components. To support this task, we introduce YT-SPEECH, a speech-oriented $360^\circ$ video-FOA dataset curated from YouTube. We propose a two-stage Localizer-Renderer framework, where an audio-visual segmentation backbone provides frame-wise spatial heatmaps and a conditional complex-domain U-Net reconstructs directional FOA signals from the omnidirectional channel. A confidence-based gating strategy stabilizes conditioning under ambiguous acoustic conditions. Experiments show improved reconstruction fidelity, spatial accuracy, and perceptual speech quality relative to ablated variants and prior approaches.
arXiv abstractPDF

Comments

Log in to comment, reply, and vote.

No comments yet.