Make AI phone and voice agents sound and react like a person, answering in under a tenth of a second, so talking to software stops feeling like waiting on a machine.
ScalingStage 5 of 5
Sonic-3.6 out Aug 2026; #1 on Artificial Analysis voice boards; 61 locales.
Updated 27 Aug 2026·Checked 11 Oct·0 updates this week
Milestones
No announced next step
Sonic 2.0: twice as large, still about 90 ms11 Mar 2025Complete.
Sonic 3 with emotion and laughterOct 2025Complete.
Real-time voice models (Sonic TTS, Ink STT), state space models
Cartesia builds very fast speech models for live voice agents: Sonic turns text into speech and Ink turns speech into text. Founded by Stanford researchers behind state space models, an alternative to transformers.
Sonic-3.6 (Aug 2026) tops Artificial Analysis's text-to-speech leaderboards, starting to speak in under 90 ms.
Co-founder Albert Gu co-invented Mamba; the team builds on state space models, which keep a fixed-size memory.
Raised $91M by its 2025 Series A (Kleiner Perkins), and a reported $100M Series B in Oct 2025.