# Sonic by Cartesia

Canonical page: https://technooptimist.io/ai/sonic
Frontier: [AI](https://technooptimist.io/ai.md) · Field: Voice & music · Company: [Cartesia](https://technooptimist.io/companies/cartesia.md)
Stage: Scaling (5 of 5: Research → Proto → Pilot → Deploy → Scaling)
Updated: last update 2026-08-27, facts checked 2026-10-11
Goal: Make AI phone and voice agents sound and react like a person, answering in under a tenth of a second, so talking to software stops feeling like waiting on a machine.
Status: Sonic-3.6 out Aug 2026; #1 on Artificial Analysis voice boards; 61 locales.

## Milestones

| Status | Milestone | Date | Slip |
| --- | --- | --- | --- |
| Done | Sonic 2.0: twice as large, still about 90 ms | 11 Mar 2025 |   |
| Done | Sonic 3 with emotion and laughter | Oct 2025 |   |
| Done | Ink-2 speech-to-text tops streaming accuracy board | 9 Jul 2026 |   |
| Done | Sonic-3.6: rebuilt data, model and training | 27 Aug 2026 |   |

## Most important updates

- **27 Aug 2026**: Cartesia ships Sonic-3.6, topping Artificial Analysis text-to-speech rankings. Two months after Sonic-3.5, Cartesia rebuilt data, architecture, training and evaluation for Sonic-3.6. ([our story](https://technooptimist.io/stories/cartesia-ships-sonic-3-6-topping-artificial-analysis-text-to-speech-dbbf267d.md) · [original source](https://cartesia.ai/blog/sonic-3.6))
- **9 Jul 2026**: Cartesia releases Ink-2, a speech-to-text model for live voice agents. Ink-2 transcribes speech as it streams and decides when the speaker has finished from meaning rather than silence. ([our story](https://technooptimist.io/stories/cartesia-releases-ink-2-a-speech-to-text-model-for-live-voice-agents-347c61d1.md) · [original source](https://cartesia.ai/blog/ink-2))
- **11 Mar 2025**: Cartesia raises $64M Series A led by Kleiner Perkins and launches Sonic 2.0. Cartesia raised $64 million led by Kleiner Perkins, for $91 million raised in total, and released Sonic 2.0 on its state space architecture. ([our story](https://technooptimist.io/stories/cartesia-raises-64m-series-a-led-by-kleiner-perkins-and-launches-sonic-c493a296.md) · [original source](https://cartesia.ai/blog/series-a))

## Current obstacles

- **Crowded market**: ElevenLabs, OpenAI and Google all sell real-time voices; leads on leaderboards shift within months.
- **Voice cloning misuse**: Instant cloning of any voice from seconds of audio enables fraud and impersonation unless consent is enforced.

## Physics limits

- **Speed vs quality**: Starting to talk before the full sentence is planned limits how well the model can set intonation for what comes later.
- **Network delay adds up**: Even a 90 ms model sits inside a chain of transcription, reasoning and network hops that sets how fast an agent can really reply.

## How it works

### Stream: Speak while still thinking

Audio is generated piece by piece and streamed, so the first syllable plays in under 90 ms while the rest is still being made.

### Memory: A fixed-size running summary

State space layers carry a compact summary forward instead of re-reading all prior audio, keeping each step cheap.

### Voice: Clone from a short sample

Given a brief recording, the model reproduces that voice and accent; emotion and pace can be steered per line.

### Agent loop: Paired with Ink

Ink transcribes the caller and spots the end of a turn from meaning, not silence; Sonic then speaks the reply.

## Spec sheet

| Spec | Sonic | Sonic-3.5 | Provenance |
| --- | --- | --- | --- |
| Time to first audio | Under 90 ms |   | reported |
| Generation speed | About 132 characters/s | 68 (v3 Conversational) | reported |
| Locales | 61 (11 Indic) |   | reported |
| Preset voices | 500+ |   | reported |
| Artificial Analysis Elo (controlled) | 1120, ranked #1 | 1095 (Sonic-3.5) | reported |
| Blind test vs Eleven v3 (US English) | 92% preferred Sonic-3.6 |   | reported |

Provenance: reported = stated by the company; estimated = our estimate; sample = a sample figure.

## About Cartesia

Cartesia builds very fast speech models for live voice agents: Sonic turns text into speech and Ink turns speech into text. Founded by Stanford researchers behind state space models, an alternative to transformers.

- Sonic-3.6 (Aug 2026) tops Artificial Analysis's text-to-speech leaderboards, starting to speak in under 90 ms.
- Co-founder Albert Gu co-invented Mamba; the team builds on state space models, which keep a fixed-size memory.
- Raised $91M by its 2025 Series A (Kleiner Perkins), and a reported $100M Series B in Oct 2025.

- Founded: 2023
- Headquarters: US
- Status: private
- Raised: $186M
- Last round: Series B $100M, Oct 2025, led by Kleiner Perkins
- Website: https://cartesia.ai
- People: Karan Goel (Co-founder & CEO); Albert Gu (Co-founder & Chief Scientist)
- Partners, customers, investors: Kleiner Perkins (investor); Index Ventures (investor); NVIDIA (investor)

Company page: https://technooptimist.io/companies/cartesia.md

## Related programs

Everything in AI: https://technooptimist.io/ai.md

---

The Techno Optimist tracks every frontier of technology, program by program: https://technooptimist.io/ (a Markdown version of any page: add .md to its URL; site map for agents: https://technooptimist.io/llms.txt).
Data API (JSON, realtime stream): https://technooptimist.io/developers. People read 5 pages free, then Reader $10/month; these Markdown pages are open to agents.
