Real-Time Voice AI Needs a Conversation Policy
The latest voice models can listen while they speak, handle interruptions, and keep a conversation moving while tools work in the background.
That changes what founders can build.
It also creates a new product risk. A voice agent can sound confident long before it has earned the right to act.
Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15. The models can execute tools in the background while continuing a conversation. OpenAI released GPT-Live-1 in the API five days earlier, with full-duplex conversation and controls for how a voice agent speaks and acts.
The technical layer is improving quickly. The founder opportunity is no longer just access to a natural voice. It is the operating policy behind that voice.
Before you tune personality, define the conversation policy.
A conversation policy is more than a prompt
A prompt tells the model how to respond. A conversation policy tells the product how to behave when a real person is talking.
It should answer practical questions:
- When may the agent interrupt?
- How long should it wait before filling silence?
- What should it say while a tool is working?
- Which facts must it repeat for confirmation?
- Which actions can it complete on its own?
- When must it stop and transfer to a person?
These rules matter because speech moves faster than a form. A user may correct a date halfway through a sentence. They may say “yes” to acknowledge that they heard the agent, not to authorize a purchase. They may change the request after a tool call has already started.
A fluent model does not remove that ambiguity. It can make the ambiguity harder to notice.
Your domain experience helps you recognize the moments that matter. A logistics operator knows the difference between reporting a delay and changing a delivery commitment. A healthcare administrator knows when a caller is asking for general information and when the conversation has crossed into a decision that requires licensed judgment. A home-services owner knows which symptoms justify an emergency handoff.
That knowledge belongs in product rules, not only in a system prompt.
Design turn-taking around the job
Google’s Live API supports barge-in, which lets users interrupt the model. It also supports proactive audio, which gives developers control over when the model should respond. GPT-Live-1 is designed to listen and speak at the same time, rather than forcing every exchange into rigid turns.
Those capabilities make natural conversation possible. They do not tell you what good turn-taking looks like for your workflow.
Start by mapping four moments:
- The opening. State what the agent can help with and what it cannot do.
- The intake. Ask one question at a time when accuracy matters. Allow broader answers when discovery matters.
- The action. Confirm the required facts before the product writes, books, sends, charges, or changes anything.
- The close. Explain what happened, what will happen next, and how the user can correct a mistake.
Then test the messy versions. The caller interrupts. Two people speak. A name is spelled twice in different ways. The user goes silent. A tool takes ten seconds. The requested action falls outside the approved workflow.
A good voice product does not hide these situations with filler. It handles them predictably.
Background work needs visible progress
Gemini 3.8 Live can continue speaking while tools and API calls run in the background. Its Extended Thinking model can use short verbal acknowledgements and progress narration during longer tasks. OpenAI describes GPT-Live-1 as a front-end voice layer that can delegate deeper reasoning and actions to other models and tools.
This creates a useful product pattern. The voice layer can keep the user oriented while another system does the work.
But progress narration needs boundaries. Otherwise the agent may describe an outcome before the tool has returned it.
Use three states:
- Received. “I have the address and preferred time.”
- Working. “I’m checking availability now.”
- Confirmed. “The appointment is booked for Tuesday at 2 p.m.”
Do not let “working” sound like “done.”
This is especially important when the background action affects a customer record, calendar, payment, shipment, or external message. The agent should report the tool result, not predict it.
For a first product, log each state change. Store the user request, the confirmed inputs, the tool call, the tool response, and the final spoken confirmation. That record will show you whether failures come from listening, reasoning, tool execution, or communication.
Authority should be explicit
The most important question in a voice product is not “How human does it sound?”
It is “What is it allowed to do?”
Create an authority table before launch:
- Inform. The agent may answer from an approved source.
- Draft. The agent may prepare an action for review.
- Confirm. The agent may act after repeating the critical details and receiving clear approval.
- Transfer. The agent must hand off when policy, risk, or uncertainty crosses a defined threshold.
The exact table depends on the domain. That is why experienced professionals have an advantage. You already know which mistakes are inconvenient and which ones damage trust.
Do not treat every “yes” as consent. Confirmation should name the action and its important details. “Should I send this proposal to Maria at Acme for $12,000?” is clearer than “Ready to proceed?”
Do not let a transfer erase the work already done. Pass the summary, collected facts, unresolved issue, and relevant transcript to the person taking over. The customer should not have to restart the conversation.
Build the policy before the personality
The new voice APIs make it easier to create an impressive demo. Google’s Live API includes interruption handling, tool use, transcription, and controls for proactive audio. OpenAI prices the GPT-Live-1 front-end voice layer at $0.05 per minute, before the cost of the backend model and tools you pair with it.
Access is becoming easier. Product judgment is not.
If you are building a voice product from your domain expertise, start with one repeated conversation and write the policy before choosing the voice:
- Define the finished outcome.
- Mark the facts that require confirmation.
- Set the agent’s authority level for each action.
- Script what it says during background work.
- Define the transfer conditions.
- Test interruptions, silence, corrections, and tool failures.
- Review the logs with someone who handles the workflow today.
That policy becomes a durable product asset. Models will change. Voices will improve. The rules for completing the job safely and clearly will remain closer to the customer problem.
Your domain expertise gives you those rules. Turn them into a product before someone else turns the same workflow into a generic demo.
If you want help selecting the right workflow, validating it with buyers, and launching a focused AI product in 12 weeks, book a strategy call. That is the work we do inside AI Product Accelerator.