Guava Voice Index · v1.0 · September 2026
How close is voice AI
to a human on the call?
Voice agent quality gets reported as a demo or a single latency number, never as how close an agent gets to a human handling the same call. The Guava Voice Index measures exactly that: a composite score across five pillars, responsiveness, conversational flow, fidelity, resolution and TTS quality, combining automated metrics from open source benchmarks with blind human evaluations.
- 5 pillars
- 12 metrics, automated and human
- 4 systems scored
- 60/40 content to experience
Composite scores
Out of 100. Higher is better. September 2026.
Highest composite reached
Nothing scored above 60. The gap between the best voice agent tested and the top of the scale is far larger than the gap between first and fourth, and the margin between first and second is 1.77 points.
Disclosure
Guava publishes this index and Guava competes in it. That is worth saying plainly, and it is why the method is public rather than summarised: the formula, the pillar weights, all 12 submetrics, the external sources, and the 2 pillars below where another system scored higher. The margin at the top is 1.77 points out of 100, close enough that a different weighting would change the order.
Five pillars, every system, published in full
Each bar is the points a system earned out of the points that pillar is worth, ranked best first. Content pillars carry 60 of the 100 points, experience pillars 40.
How long the caller waits between finishing a sentence and hearing a reply.
Whether the agent starts, stops and yields at the points a person would.
Whether what the agent says is supported by the source material it was given.
Whether the caller's actual reason for calling was handled end to end.
Whether the delivery carries the prosody, pacing and repair a listener expects.
Daytona’s two open fronts are conversational flow and resolution. Resolution is the heavier at 30 points and measures whether the caller’s reason for calling was handled end to end, so it is the one we are working on first. TTS quality is the number worth staring at for the whole field: Daytona scores 1.7 of the 15 points available, the best any other system managed was 1.1, and nobody is close to a human. That is headroom, not a verdict.
Against a live human
The human bar is still ahead.
Blind pairwise listening tests, minimum 10 evaluators per pair, comparing voice agents against recordings of human agents handling the same calls. Each figure is how often evaluators preferred the human, across all four systems.
Response timing felt more natural
Listeners picked the human nine times out of ten.
Recovered better from interruption
The widest gap of the three, and the hardest to close.
Most fitting dialogue for the conversation
Close to a coin flip. The only dimension near parity.
Timing and interruption recovery are not close. The one dimension near parity is choosing what to say, which is also the pillar carrying the most weight.
The method
All 12 metrics, their sources and weights, the validation gates, the duplication rules and the human protocol are published in full. Automated metrics come from EVA-Bench (ServiceNow) and word error rate from Coval.
Have a system scored
The next round takes submissions from anyone running a production voice agent, whether it is a single model or a stack you assembled yourself. Tell us what to evaluate and we will come back with the call set, the consent requirements and what we need from you to run it.
Scoring runs in batches, not on demand. The automated half can be reproduced from the published methodology; the human half runs on a commissioned evaluator panel, which is why it cannot be self-served.
Your expertise finds its voice.
Guava Voice Index v1.0 · published 2026-09-23 · last updated 2026-09-23. Scores are versioned; when the methodology changes, the version changes with it.