SOC 2 Type II · HITRUST i1 · PCI DSS L1 · Reports available

Guava Voice Index · v1.0 · September 2026

How close is voice AI
to a human on the call?

Voice agent quality gets reported as a demo or a single latency number, never as how close an agent gets to a human handling the same call. The Guava Voice Index measures exactly that: a composite score across five pillars, responsiveness, conversational flow, fidelity, resolution and TTS quality, combining automated metrics from open source benchmarks with blind human evaluations.

  • 5 pillars
  • 12 metrics, automated and human
  • 4 systems scored
  • 60/40 content to experience

Composite scores

Out of 100. Higher is better. September 2026.

Guava Daytona
58.88
ElevenLabs + GPT 5.4
57.11
Grok Think Fast 1.0 / Grok STT
53.60
ElevenLabs + Claude Haiku 4.5
51.06
050100
58.88

Highest composite reached

Nothing scored above 60. The gap between the best voice agent tested and the top of the scale is far larger than the gap between first and fourth, and the margin between first and second is 1.77 points.

Disclosure

Guava publishes this index and Guava competes in it. That is worth saying plainly, and it is why the method is public rather than summarised: the formula, the pillar weights, all 12 submetrics, the external sources, and the 2 pillars below where another system scored higher. The margin at the top is 1.77 points out of 100, close enough that a different weighting would change the order.

Five pillars, every system, published in full

Each bar is the points a system earned out of the points that pillar is worth, ranked best first. Content pillars carry 60 of the 100 points, experience pillars 40.

Guava DaytonaElevenLabs + GPT 5.4Grok Think Fast 1.0 / Grok STTElevenLabs + Claude Haiku 4.5
Responsiveness 10 pts Daytona leads

How long the caller waits between finishing a sentence and hearing a reply.

Guava Daytona
4.2
Grok Think Fast 1.0 / Grok STT
3.3
ElevenLabs + Claude Haiku 4.5
3.0
ElevenLabs + GPT 5.4
2.5
Conversational flow 15 pts Room to improve

Whether the agent starts, stops and yields at the points a person would.

Grok Think Fast 1.0 / Grok STT
8.3
Guava Daytona
7.7
ElevenLabs + Claude Haiku 4.5
7.2
ElevenLabs + GPT 5.4
6.3
Fidelity 30 pts Daytona leads

Whether what the agent says is supported by the source material it was given.

Guava Daytona
25.2
ElevenLabs + GPT 5.4
24.6
Grok Think Fast 1.0 / Grok STT
23.4
ElevenLabs + Claude Haiku 4.5
21.3
Resolution 30 pts Room to improve

Whether the caller's actual reason for calling was handled end to end.

ElevenLabs + GPT 5.4
23.7
Guava Daytona
20.1
ElevenLabs + Claude Haiku 4.5
19.2
Grok Think Fast 1.0 / Grok STT
17.7
TTS quality 15 pts Daytona leads

Whether the delivery carries the prosody, pacing and repair a listener expects.

Guava Daytona
1.7
Grok Think Fast 1.0 / Grok STT
1.1
ElevenLabs + GPT 5.4
0.3
ElevenLabs + Claude Haiku 4.5
0.3

Daytona’s two open fronts are conversational flow and resolution. Resolution is the heavier at 30 points and measures whether the caller’s reason for calling was handled end to end, so it is the one we are working on first. TTS quality is the number worth staring at for the whole field: Daytona scores 1.7 of the 15 points available, the best any other system managed was 1.1, and nobody is close to a human. That is headroom, not a verdict.

Against a live human

The human bar is still ahead.

Blind pairwise listening tests, minimum 10 evaluators per pair, comparing voice agents against recordings of human agents handling the same calls. Each figure is how often evaluators preferred the human, across all four systems.

91%

Response timing felt more natural

Listeners picked the human nine times out of ten.

93%

Recovered better from interruption

The widest gap of the three, and the hardest to close.

55%

Most fitting dialogue for the conversation

Close to a coin flip. The only dimension near parity.

Timing and interruption recovery are not close. The one dimension near parity is choosing what to say, which is also the pillar carrying the most weight.

The method

All 12 metrics, their sources and weights, the validation gates, the duplication rules and the human protocol are published in full. Automated metrics come from EVA-Bench (ServiceNow) and word error rate from Coval.

Read the full methodology →

Have a system scored

The next round takes submissions from anyone running a production voice agent, whether it is a single model or a stack you assembled yourself. Tell us what to evaluate and we will come back with the call set, the consent requirements and what we need from you to run it.

Scoring runs in batches, not on demand. The automated half can be reproduced from the published methodology; the human half runs on a commissioned evaluator panel, which is why it cannot be self-served.

Your expertise finds its voice.

Guava Voice Index v1.0 · published 2026-09-23 · last updated 2026-09-23. Scores are versioned; when the methodology changes, the version changes with it.