VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published 9 days ago • 175
SPEAR Collection A Unified SSL Framework for Learning Speech and Audio Representations • 2 items • Updated 8 days ago
GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark Paper • 2606.28884 • Published Jun 27 • 1
GigaSpeech Series Collection Evolving, Large-Scale, and Multi-domain ASR Corpus • 6 items • Updated Jun 26
UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning Paper • 2606.04939 • Published Jun 3
Evaluating the Expressive Appropriateness of Speech in Rich Contexts Paper • 2605.09413 • Published May 10 • 5
Evaluating the Expressive Appropriateness of Speech in Rich Contexts Paper • 2605.09413 • Published May 10 • 5
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Paper • 2605.06407 • Published May 7