Cartesia

Real-time voice models (Sonic TTS, Ink STT), state space models

USprivateest. 2023cartesia.ai (opens cartesia.ai)Checked 11 Oct

The brief

Analyst view
  • Sonic-3.6 (Aug 2026) tops Artificial Analysis's text-to-speech leaderboards, starting to speak in under 90 ms.
  • Co-founder Albert Gu co-invented Mamba; the team builds on state space models, which keep a fixed-size memory.
  • Raised $91M by its 2025 Series A (Kleiner Perkins), and a reported $100M Series B in Oct 2025.

Technical approach

As reported
Architecture

State space models

Instead of re-reading everything said so far like a transformer, the model updates a fixed-size summary, so cost grows linearly with length.

Latency

Speak within a blink

Streaming the first audio in under 90 ms keeps a phone-agent conversation from feeling like it lags.

Full stack

Ears and voice

Ink transcribes the caller and detects when they finished speaking; Sonic answers. Both are tuned for live agents.

Primer