What is Cartesia?
Cartesia builds real-time voice models on a state space model (SSM) architecture designed around latency rather than retrofitted for it. Its Sonic-3 model reaches time-to-first-audio of roughly 40 milliseconds, and its instant voice cloning builds a usable clone from seconds of sample audio — then synthesises that cloned voice across 42 languages, with regional accent variants for several of them.