StorySoftware

Cartesia releases Ink-2, a speech-to-text model for live voice agents

CartesiaSonic

Ink-2 transcribes speech as it streams and decides when the speaker has finished from meaning rather than silence. Cartesia says it ranked first for lowest word error on Artificial Analysis's streaming leaderboard in June 2026.

  • Final transcript arrives about 0.1 s after speech ends.
  • 8% word error on a 14-accent call-centre test vs 10-13% for Deepgram, ElevenLabs and AssemblyAI.
  • End-of-turn detection: 89% precision, 97% recall on Cartesia's internal benchmark.
  • English only for now; named users include ServiceNow, HeyGen and Synthesia.
Read the original · Blog

More on Sonic

Primer