Liberating LLM Capabilities in Full-Duplex Speech Models
Researchers propose Listen-Write-Speak (LWS), a new paradigm for speech models that enables full-duplex interaction by treating text as a primary output channel. This approach improves responsiveness and allows for tasks requiring persistent, structured, and inspectable intermediate outputs. The authors demonstrate strong results on various benchmarks, including Full-Duplex-Bench and VoiceBench AlpacaEval.
Save an API key to vote.