Shared knowledge, designed for speech
The voice agent transcribes customer speech, uses the same knowledge and answer pipeline, then synthesises the response. There is no separate “voice knowledge base”; updates made for web and messaging also reach the voice channel.
The voice channel runs on two provider lines. The priority line runs on our own servers, using live streaming infrastructure for call transport and speech synthesis we host ourselves. The second line is a fallback kept for rollback; new work is designed for the priority line, and that line's specific constraints do not bind the fallback.
One conversation turn
- Microphone audio is transcribed in real time.
- End of turn is detected without treating every short pause as completion.
- The question is searched through the shared retrieval pipeline.
- The answer is shaped for spoken length and structure.
- Text is synthesised while interruptions are managed if the user speaks.
Language
The conversation starts in the visitor's browser language. If the visitor switches to another language the business has enabled, the conversation follows. Answers, canned lines and repeated prompts all track the language of the call. A short English question does not receive a Turkish answer.
Proper names, brands and locations can be supplied as transcription hints. These hints do not add new facts; they help the recogniser spell what was heard.
Measured latency
On the priority line, measured silence-to-first-audio is p50 3.2 seconds, p95 6.6 seconds; the answer path's first byte is p50 1.8 seconds. These figures vary with hardware, network and call length.
Spoken answers cannot be very short. A short sentence followed by another short sentence makes the generated audio sound broken, so answers are not shortened below a meaningful length. Waiting lines are only used above a measured threshold, and brief silence is treated as normal.
Designing spoken answers
A listener cannot scan a long list on screen. Give the direct answer first, offer only a few options and ask permission before adding detail. Rather than reading URLs, long codes or tables aloud, offer to send the information in writing or reach a representative.
Cost and availability
The voice channel is billed per minute and costs considerably more than written chat. Minutes are shared between voice calls on the website and phone calls. The key is not in your panel but on Dolphy's side, so enabling the voice channel requires talking to our team.
Privacy and handoff
On the priority line no audio is recorded; the conversation transcript is stored. If recording is enabled, purpose, retention and access should be explicit and the user should be informed. Transcripts and summaries can also contain personal data.
When evidence is missing, a sensitive decision is requested or the user asks for a person, the agent should say so and use a safe handoff path. Automated guidance is not a replacement for an authorised professional in urgent or high-risk decisions.