StorySoftware
Cartesia releases Ink-2, a speech-to-text model for live voice agents
Ink-2 transcribes speech as it streams and decides when the speaker has finished from meaning rather than silence. Cartesia says it ranked first for lowest word error on Artificial Analysis's streaming leaderboard in June 2026.
- Final transcript arrives about 0.1 s after speech ends.
- 8% word error on a 14-accent call-centre test vs 10-13% for Deepgram, ElevenLabs and AssemblyAI.
- End-of-turn detection: 89% precision, 97% recall on Cartesia's internal benchmark.
- English only for now; named users include ServiceNow, HeyGen and Synthesia.