Voice Chat
Voice chat is the always-on, telephone-style conversation mode users get by swiping right from text chat. The audio stream is bidirectional and continuous — the user just talks and the AI talks back — and the conversation is grounded in the same documents and project configuration the text chat uses.
How Users Access It
Voice mode is reached by swiping right on the chat page. The chat container is a horizontally-paged scroll view: the left page is the text chat you already know; the right page is the voice page.
The first swipe to the voice page opens the WebSocket session and starts the mic. Swiping back to text pauses the mic but keeps the session warm — re-entering within ~10 seconds reuses the same connection. After ~10 seconds idle, the session closes; the next swipe opens a fresh one.
There is no “tap-to-talk” or push-to-talk button. The mic is on the entire time the user is on the voice page; the AI uses server-side voice activity detection to know when the user has finished speaking.
Enabling Voice
Voice is gated at two layers, both of which must allow it for the mic to appear:
| Layer | Where | What it does |
|---|---|---|
| Subscription plan | See Billing | The project’s plan must include voice. If the plan disables it, the WebSocket connection is rejected with close code 4003 (VOICE_NOT_ALLOWED) regardless of any other setting. |
| SDK prop | <Qafka voiceEnabled={false} /> | Per-app override. Default true. Use it when a project is voice-eligible but a particular build (e.g. accessibility-restricted variant) shouldn’t expose it. |
If either layer disables voice, the SDK hides the voice page entirely — no swipe affordance, no mic icon, no surprise tap.
Voice State Machine
The voice page has five states the SDK exposes and the voice components render:
| State | Meaning |
|---|---|
idle | No active connection. The voice page is visible but the mic isn’t capturing. |
connecting | WebSocket handshake + audio engine warmup. Brief. |
listening | Mic is open, capturing user audio, waiting for VAD to detect end of speech. |
thinking | User has finished speaking; the AI is processing (possibly invoking a tool). No audio plays. |
speaking | The AI is speaking back. Audio is rendering through the device speaker; transcript updates token-by-token. |
Barge-in
The user can interrupt the AI mid-sentence. By default, voice sessions run full-duplex: the device’s native echo cancellation is trusted to strip the AI’s own playback out of the mic, so the mic stays open the whole time — including while the AI is speaking — which is what makes barge-in feel instant.
If the on-device echo cancellation proves unreliable during a session (the AI’s own voice starting to get picked up and transcribed as if it were the user), the SDK automatically downgrades that session to half-duplex: the mic stops forwarding audio while the AI is speaking, so playback can’t leak back in as a false user turn. This is automatic and per-session — there’s no prop to control it, and a fresh session starts full-duplex again.
Barge-in is still useful for “no, wait, I meant…” follow-ups; less useful when the AI is reading a long answer the user actually wants — keep replies concise via your voice instructions.
Voice Instructions
Voice mode uses a separate per-project instructions field, edited from the AI Behavior page’s Voice tab. It’s appended to the system prompt only when voice mode is active, and exists because voice prompting needs are different from text:
- Pronunciation rules (“Acme X3 should be read as ‘ex-three’, not spelled out”)
- Length limits (“Keep answers to 2–3 sentences”)
- Tone (“Friendly but professional”)
- Redaction (“Never read out full account numbers”)
The text-chat critical instructions still apply on top of voice instructions — if you’ve already written a rule there, you don’t need to repeat it. Voice instructions are for things only relevant to spoken interaction.
Voice UI Customization
The voice page’s visuals (indicator, background, transcript, mute button) are swappable component slots, the transcript area’s behavior is controlled by the voiceTranscript prop ('centered' | 'chat' | 'off'), and the user can mute their own mic independently of the SDK’s internal mic gating. The session itself is also controllable through the widget’s imperative handle (connect/disconnect, pause/resume mic, mute/unmute/toggleMute). All of this is covered in full on the React Native Widget page — this page focuses on voice’s runtime behavior, not its UI surface.
For everyday tone/color tuning, edit the chat theme instead — it’s automatically applied to the default voice components without writing a single line of code.
Conversation Persistence
A voice conversation is the same database object as a text conversation — same conversationId, same place in the Conversations list, same retention. Voice messages are internally distinguishable from text messages within a session.
The user can swipe back and forth between text and voice within a single session and the AI keeps the conversation context — both transports write to and read from the same message history.
Tools, Navigation, and Context in Voice
Voice is no longer limited to Q&A against your documents, but tool support is narrower than in text:
- Tools — only
customandcustom-with-aiexecution-mode tools are registered in voice sessions;onToolSuggestedfires for them the same way it does in text.server,staticData, andcalendarmode tools aren’t exposed to the AI in voice sessions at all — they’re simply not declared, so the AI has no way to call them there. See Handling Tools › Tools in Voice for the full breakdown. - Navigation — built-in navigation suggestions work in voice sessions. See Navigation for how AI-triggered and user-triggered navigation differ.
- Runtime context — the
contextprop is forwarded into the voice session and injected into the voice system prompt, the same as text.
Voice also has its own built-in capabilities beyond navigation: searching your documents for facts (the same grounding text chat uses) and rendering the data chip list handled by the dataChipList slot.
What’s Out of Scope
Voice is intentionally narrower than text in a few places, by design:
- No file upload in voice (React Native) — file inputs are a text-mode interaction in the mobile SDK.
- No camera / video input — audio only.
- No audio recording / replay — only transcripts persist; raw audio is not stored.
- External suggestions (WhatsApp, phone, app store, etc.) are text-only for now — voice can’t surface them.
Close Codes
If the voice WebSocket closes unexpectedly, the close code tells you why:
| Code | Meaning |
|---|---|
4001 | Authentication failed — missing/invalid auth frame, invalid API key, or a session-token problem (missing, mismatched, or a bundle ID mismatch). |
4002 | The connection failed to establish after authentication succeeded. |
4003 | Voice isn’t allowed — the project’s plan doesn’t include voice, or the project/API key behind the connection couldn’t be resolved. |
4008 | Connection capacity limit hit — too many concurrent connections for this server, this IP, or this API key. |
4010 | The project is suspended (project_unavailable) — same underlying state as Error Handling › project_unavailable, but unlike the text-chat case, this does not trigger onUnavailable: the voice page just resets to idle like any other unhandled close code. |
4011 | The project’s voice minutes are exhausted for the current period. |
Of these, only 4011 gets a dedicated, localized message surfaced by the SDK today: the app is scrolled back to text chat and the message is dropped into the transcript as a chat bubble. The other codes just reset the voice page to idle with no message shown — the <Qafka /> component doesn’t currently expose a callback for the raw close code/reason.
Connection Lifecycle
| Event | What happens |
|---|---|
| User swipes to voice page | Open WebSocket if none is open, then start the audio pipeline, then start the mic |
| User swipes back to text page | Pause the mic; keep the WebSocket open for ~10s in case they swipe back |
| 10s idle on text page | Close the WebSocket; next swipe to voice opens a fresh session |
| 60s with no activity while connected | The server closes the WebSocket on its own idle timer; the voice page resets to idle |
Voice minutes exhausted (4011) | Session closes, a localized message is shown, and the SDK falls back to text |
| App goes to background | Mic and WebSocket close immediately |
| App returns to foreground on the voice page | Open a new WebSocket and start the mic |
| Network drop | Close cleanly and surface the error to the voice page |
Server-Side Components
The voice session is a Qafka-managed WebSocket between the SDK and our backend. The backend proxies to the underlying model provider itself — provider credentials stay server-side, and clients never talk to that provider directly. RAG retrieval, conversation persistence, and system prompt assembly all happen on the backend. Voice has its own connection and usage controls rather than reusing text chat’s: connection-capacity limits (close code 4008) and per-period voice-minute metering (close code 4011) — see Close Codes.