Response Time
How fast does the model respond after the caller stops speaking?
Measures response latency with P50 for the typical caller experience and P90 for slower tail responses.
Compare leading AI models across response quality, speed, and reliability. Find the right balance for your voice agent and choose with confidence.
What we evaluate
Retell tests the behaviors that make a voice agent feel fast, reliable, and ready for production.
Response Time
Measures response latency with P50 for the typical caller experience and P90 for slower tail responses.
Grounding
Grounded
Answers stay within supported information.
Unsupported
Answers contain unsupported or unverifiable claims.
Averaged across every model scored on Retell's grounding suite: answers supported by the prompt, tools, or retrieved context.
Tool Calling
Tests whether the model selects the right action, passes valid arguments, and uses tool results correctly.
Instruction Following
Checks whether the model follows role constraints, business rules, formatting requirements, and refusal rules.
Task Completion
Evaluates whether the model resolves the task, asks for missing information, and avoids dead ends.
Model Benchmarks
Compare quality, latency, and reliability across top providers, with filters that make the tradeoffs easier to see.
Whether answers stay within the prompt, tools, and retrieved context.