Aivora
arXivAudio AIAdvanced

WorldSonus: Bringing Real-Time Spatial Audio to World Models

WorldSonus:為虛擬世界模型注入即時且具空間感的立體聲

2 min read
WorldSonus: Bringing Real-Time Spatial Audio to World Models
The 30-second version

While current world models excel at visual synthesis, they usually remain silent. WorldSonus addresses this gap by tackling real-time streaming, dynamic control, and spatial alignment. It features a streaming causal autoregressive diffusion architecture with a low real-time factor (RTF) of 0.41. Using chunk-indexed prompt scheduling, users can control sound events mid-stream. Additionally, training with stereo and ambisonic datasets ensures synthesized audio aligns perfectly with camera and scene dynamics.

Key points

01

Low-Latency Streaming

Employs a causal autoregressive diffusion architecture to generate audio in chunks, achieving a low real-time factor of 0.41.

02

Mid-Stream Dynamic Control

Features an audio-centric captioning pipeline and chunk-indexed prompt scheduling for interactive sound adjustments during generation.

03

Spatially Aligned Audio

Incorporates high-quality stereo and ambisonic supervision to align audio effects with scene geometry and camera motion.

How it works

WorldSonus Interactive Spatial Audio Synthesis Architecture
Scene geometryMid-stream instructionStreaming synthesisStereo outputVideo & Motion InputChunk PromptsCausal AR DiffusionReal-time ChunksAligned Spatial Stereo

Why it matters

Traditional video-to-audio models are often offline and non-interactive. WorldSonus breaks these boundaries by enabling real-time, text-controlled, and spatially-aware stereo. This is crucial for enhancing immersion in simulators, game engines, and virtual world models, unlocking true multimodal interactive environments.

Who it affects

  • AI Developer
  • AI Researcher
  • Content Creator

How to use it

  1. 1Real-time stereo sound synthesis for interactive video games and 3D world models.
  2. 2Spatially aware audio generation for visual simulators with camera motion tracking.
  3. 3Interactive video post-production with mid-stream prompt adjustments for sound effects.

Limitations & caveats

  • Highly dependent on the availability and quality of stereo and ambisonic training data.
  • As an autoregressive diffusion model, there remains a potential risk of error accumulation over very long sequences.

Related

Fine-Tuning NVIDIA Nemotron ASR for Regional Dialects: A Guide to Dialect Adaptation
NVIDIA DeveloperAudio AI

Fine-Tuning NVIDIA Nemotron ASR for Regional Dialects: A Guide to Dialect Adaptation

如何微調 NVIDIA Nemotron 語音辨識模型?以沙烏地阿拉伯方言為例的跨語言適應指南

This guide details how to fine-tune NVIDIA Nemotron 3.5 ASR using the NeMo framework, leveraging minimal curation, replay mixing, and bucketed batches to drastically improve regional dialect recognition while preserving baseline language accuracy.

2 min read