
Speech-to-speech model benchmarks on live phone calls
Cekura Bench publishes voice AI benchmarks you can verify. Our new speech-to-speech benchmark tests 9 realtime voice models, including GPT Realtime 2.1, Gemini Live, Grok and Phonic, as complete phone agents on live calls: 82 scenarios, three runs each. Models are ranked on reliability, data accuracy, stalled calls, response time and cost, and every call transcript is public. Cekura Bench also covers voice agent benchmarks and STT benchmarks, with TTS benchmarks coming soon.
Cekura Bench provides verified benchmarks for speech-to-speech AI models tested on live phone calls, evaluating nine models across 82 scenarios. The platform ranks models based on reliability, data accuracy, response time, and cost, with all call transcripts publicly accessible.
Scored deterministically. Only candidates that fire a story trigger are sent to a model, so this one has no written angle.
Gaps in our data, not findings about the product. Their weight is redistributed across the 5 we did measure.
A source that found nothing is a measurement. A source that has not run is a gap. Neither means the launch lacks the thing.