Better conversations start with the right model.

Compare leading AI models across response quality, speed, and reliability. Find the right balance for your voice agent and choose with confidence.

What we evaluate

Model performance, measured where voice agents fail.

Retell tests the behaviors that make a voice agent feel fast, reliable, and ready for production.

Response Time

How fast does the model respond after the caller stops speaking?

Caller speech ends00:12.49
Model starts00:13.07
P50
P90

Measures response latency with P50 for the typical caller experience and P90 for slower tail responses.

Grounding

Does the model stay within the provided context?

Grounded

Answers stay within supported information.

Unsupported

Answers contain unsupported or unverifiable claims.

0%50%100%

Averaged across every model scored on Retell's grounding suite: answers supported by the prompt, tools, or retrieved context.

Tool Calling

Does the model choose and execute the correct tool?

Tests whether the model selects the right action, passes valid arguments, and uses tool results correctly.

Instruction Following

Does the model follow the system and task instructions?

Checks whether the model follows role constraints, business rules, formatting requirements, and refusal rules.

Task Completion

Does the model complete the user's goal?

Evaluates whether the model resolves the task, asks for missing information, and avoids dead ends.

Model Benchmarks

We benchmarked the leading models for you.

Compare quality, latency, and reliability across top providers, with filters that make the tradeoffs easier to see.

All models

Whether answers stay within the prompt, tools, and retrieved context.