Dolphy Docs
Trust and quality

Measuring quality

An agent looking good in a few sample conversations is not enough to establish reliability. Evaluation combines fixed cases representing real user intent, expected evidence and observable answer behaviour.

Golden sets

A golden set is a versioned collection of questions and expected outcomes. Where possible, each case includes:

  • The user's natural message and necessary conversation context
  • The source or acceptable evidence that should be retrieved
  • Expected language, brevity and component behaviour
  • An observable result such as a request, handoff or action

The set should not contain only easy questions with available answers. Include unknown, conflicting and multi-step cases, plus prompts likely to trigger an irrelevant card.

Retrieval and answer metrics

Recall@k reports whether correct evidence appears in the first k results; MRR reflects how high the first correct result ranks. At the answer layer, grounding, factual correctness, language, brevity and UI or action choice are evaluated separately. One aggregate score can hide which layer failed.

How evaluation runs

Every case goes through a real answer turn: the same routing, the same retrieval, the same system prompt and the same component decisions. There are two layers.

  • Deterministic checks are free: expected and forbidden text, card count, list length, word count, handoff and request expectations.
  • The judge layer evaluates the answer with a different model family. A different family is chosen deliberately so the answering model does not judge its own output. Dimensions stay binary rather than graded: grounding, correct answer, brevity, language and card suitability.

The judge run costs money and is bounded by a budget gate; the run stops when the budget is exceeded. Deterministic checks always run; the judge layer runs on request.

The voice channel is not measured

Evaluation currently covers written answers only. Voice conversation behaviour is exercised separately in real calls; none of the figures on this page apply to the voice channel.

Dolphy laboratory result

In a 30-case ecommerce laboratory evaluation dated 9 September 2026, the deterministic transition increased total passing cases from 23/30 to 29/30 and reduced p50 response time from 4,865 ms to 3,836 ms in the same run. Card suitability reached 29/30, while language and brevity checks reached 30/30.

That result is historical and can no longer be reproduced: the case set was lost together with its test business during the database outage of 10 September 2026. The same figures cannot be regenerated today, so this paragraph is kept as a record rather than a current measurement.

The current measurement is weaker and less flattering. On 25 September 2026 a 10-case journey set was run: pilot-dentomega stayed at 7/10, while ferran-test fell from 7/10 to 5/10. The drop is attributed to judge variance and to dates in the cases having gone stale, not to a prompt change; repeating the same case produced a different result. This shows that a single run on a small set is not enough to decide anything.

Both results are comparative measurements on a limited laboratory set. They are not guarantees for every tenant or for live workload.

For industry context, McKinsey estimates the productivity value of generative AI in customer operations at roughly 30–45% of current function costs. That figure comes from industry research; it is not a measured Dolphy customer outcome.

The documentation separates tasks, explanations and reference material following Diátaxis, and avoids unmeasured superiority claims following Google's developer documentation guidance.

On this page