Imported from manishkjs/gemini_live_pipecat (
docs-site/public/gemini-live/SKILL.md). Install upstream withnpx skills add manishkjs/gemini_live_pipecat --skill gemini-live. Copyright stays with the author.
Gemini Live API & Pipecat Production Engineering Reference (Complete 25-Section Handbook)
Installation: Save this file to
.agents/skills/gemini-live/SKILL.mdor~/.gemini/config/skills/gemini-live/SKILL.mdso your coding agent automatically enforces all 25 Gemini Live protocol, audio, tool-calling, and token-optimization invariants.
1. Protocol, Endpoints, Model Compatibility & Framing Matrix
| Parameter | Google AI Studio (Developer API) | Vertex AI (Production GA & Preview) |
|---|---|---|
| Base Host | generativelanguage.googleapis.com |
{region}-aiplatform.googleapis.com (for example us-central1-aiplatform.googleapis.com) |
| API Version | v1alpha or v1beta |
v1beta1 (do not use v1 for Live Bidi endpoints) |
| WebSocket Path | /ws/google.ai.generativelanguage.{ver}.GenerativeService.BidiGenerateContent |
/ws/google.cloud.aiplatform.v1beta1.LlmBidiService/BidiGenerateContent |
| Auth Method | ?key=$GEMINI_API_KEY |
Authorization: Bearer $(gcloud auth print-access-token) (Application Default Credentials) |
| Headers | Content-Type: application/json |
Content-Type: application/json |
| Supported Models | models/gemini-3.1-flash-live-preview |
projects/{PROJECT_ID}/locations/us-central1/publishers/google/models/gemini-live-2.5-flash-native-audio |
| Usage Metadata Fidelity | Includes a per-turn unattributed prompt_token_count residual (93 to 221 tok/turn) and omits thoughts_token_count from total_token_count |
Exact reconciliation (prompt_token_count == Prompt_TEXT + Prompt_AUDIO, residual = 0) and includes thoughts_token_count in total_token_count |
Protocol, Framing & Model Compatibility Invariants (Gemini 2.5 vs. Gemini 3.1 Live)
When building or migrating between gemini-live-2.5-flash-native-audio and gemini-3.1-flash-live-preview, enforce these seven protocol invariants:
response_modalitiesMust Be Singleton["AUDIO"]:- Never pass
["AUDIO", "TEXT"]inLiveConnectConfig. Requesting dual modalities triggers an immediate WebSocket close with1007 Invalid Argument. Retrieve text transcripts viainput_audio_transcriptionandoutput_audio_transcription.
- Never pass
proactive_audio&enable_affective_dialogCompatibility Guard:gemini-live-2.5-flash-native-audiosupportsproactive_audio=Trueandenable_affective_dialog=True.- Setting either flag on
gemini-3.1-flash-live-previewtriggers a1007WebSocket handshake failure. Gate both parameters behind"2.5" in model_id.
- Wire Format Changes (
realtime_inputvs.client_content):- On Gemini 3.1 Live,
send_realtime_input(media=...)is deprecated in favor ofsend_realtime_input(audio=Blob(data=pcm, mime_type="audio/pcm;rate=16000"))andsend_realtime_input(text=...). send_client_contenton Gemini 3.1 Live only supports appending turns (turn_complete=Trueorturn_complete=False) and cannot replace prior conversation history mid-session.
- On Gemini 3.1 Live,
- Thinking Configuration (
thinking_levelvs.thinking_budget):- Gemini 2.5 Flash Native Audio: Accepts
ThinkingConfig(thinking_budget=0, include_thoughts=False)to disable thinking completely. - Gemini 3.1 Flash Live: Uses discrete
ThinkingLevel("MINIMAL","LOW","MEDIUM","HIGH"). Never passthinking_budget=0to Gemini 3.1 Live; it degrades multi-step tool parameter copying and slot accuracy in multi-turn tool workflows. Always configureThinkingConfig(thinking_level="MINIMAL", include_thoughts=False).
- Gemini 2.5 Flash Native Audio: Accepts
- S2ST Interpreter Mode vs. Conversational Agent Mode:
- Do not combine
TranslationConfig(continuous speech-to-speech simultaneous interpretation) with conversational system instructions or function calling. Keep unidirectional S2ST interpreter pipelines and interactive bidirectional voice agents strictly separate.
- Do not combine
- Session Lifecycle Rotation (10 to 15 min):
- The server rotates WebSocket connections after 10 to 15 minutes with close code
1008after emitting asession_resumption_updateevent. Reconnect automatically usingsession_resumption_handle=new_handle.
- The server rotates WebSocket connections after 10 to 15 minutes with close code
- Audio Formats & Setup Handshake ACK:
- Input Audio (User to Gemini): PCM 16kHz 16-bit little-endian mono (
audio/pcm;rate=16000). - Output Audio (Gemini to Client): PCM 24kHz 16-bit little-endian mono (
audio/pcm;rate=24000). - Setup Handshake ACK: Vertex AI emits
{"setupComplete": {"sessionId": "<UUID>"}}before streaming audio.
- Input Audio (User to Gemini): PCM 16kHz 16-bit little-endian mono (
2. Speech-to-Text Routing & Cascaded Model Hierarchy
A. Speech-to-Text (STT) & Reasoning Config Invariants
- Live Transcription (
input_audio_transcription&output_audio_transcription):- Enable both
input_audio_transcription=AudioTranscriptionConfig()andoutput_audio_transcription=AudioTranscriptionConfig()insideLiveConnectConfigso your backend receives real-time user and assistant transcripts alongside native audio streams. - AI Studio
AudioTranscriptionConfigInvariant: Thelanguage_codesfield insideAudioTranscriptionConfigis only supported on Vertex AI (USE_VERTEXAI=true). Passinglanguage_codesin Google AI Studio mode (GEMINI_API_KEY) raises:ValueError: language_codes parameter is only supported in Vertex AI mode, not in Gemini Developer API mode.When operating in AI Studio mode, instantiateAudioTranscriptionConfig()with zero arguments. - Model-Generation Thinking Config Invariant (
thinking_budget=0vs.ThinkingLevel.MINIMAL):from google.genai import types if "2.5" in model_id: thinking_cfg = types.ThinkingConfig(thinking_budget=0, include_thoughts=False) else: level = getattr(types.ThinkingLevel, "MINIMAL", "minimal") thinking_cfg = types.ThinkingConfig(thinking_level=level, include_thoughts=False) - Pipecat 1.2+
LLMUserAggregatorTurn Invariant:LLMUserAggregatorignoresInterimTranscriptionFrame. Custom STT services must emitTranscriptionFrameand call_handle_transcription(transcript, is_final=True)to commit speech text into the LLM turn context. - Pipecat 1.2+ VAD Config Invariant:
FastAPIWebsocketParamsdoes not acceptvad_analyzer(passing it raises aTypeError). ConnectVADProcessordirectly inside the frame pipeline.
- Enable both
chirp_3(Cloud Speech-to-Text v2 Multilingual Sidecar):- Location Invariant: Must use
location="us"(US Multi-Region) with regional endpointus-speech.googleapis.com. Callingus-central1fails with400 Expected resource location to be us. - Languages: Pass
language_codes=["en-US", "hi-IN"]for simultaneous multilingual recognition. - Streaming: Set
enable_interim_results=Truefor low-latency partial transcripts.
- Location Invariant: Must use
chirp_2: Hosted inlocation="us-central1"(us-central1-speech.googleapis.com).latest_long/latest_short/telephony: Available inus-central1andglobal.
B. Cascaded Model Tier Standards (STT -> LLM -> TTS Alternative Pipeline)
When operating a modular cascaded pipeline alongside native Speech-to-Speech:
- Default STT: Cloud Speech-to-Text v2 (
chirp_3inlocation="us") or Gemini Live STT (response_modalities=["TEXT"]). - Default LLM:
gemini-2.5-flash-liteorgemini-2.5-flash(thinking_budget=0for low-latency dialogue). - Default TTS: Cloud Text-to-Speech
chirp_3HD voices or Gemini TTS (gemini-2.5-flash-preview-tts).
3. IAM Permissions & Local Dev Tunneling
Required Roles by Service
Grant the runtime service account the following least-privilege IAM roles:
- Vertex AI Live & LLMs:
roles/aiplatform.user(aiplatform.endpoints.predict). - Speech-to-Text v2:
roles/speech.client(speech.recognizers.recognize). - Text-to-Speech:
roles/texttospeech.client(texttospeech.synthesize). - Session / Memory Storage (Optional):
roles/datastore.userorroles/storage.objectAdmin. - ADC Quota Project: Local developer credentials require
gcloud auth application-default set-quota-project <YOUR_GCP_PROJECT_ID>.
SSH Tunneling Invariant (Remote Linux VM -> Local Browser)
- Avoid HTTP reverse proxies that strip WebSocket
Upgradeheaders or blockwss://audio streams. - Forward the application port over SSH so your local browser connects to
localhost:
Openssh -L 7860:localhost:7860 user@your-dev-vm.example.comhttp://localhost:7860/in Chrome, which treatslocalhostas a secure context (window.isSecureContext === true) fornavigator.mediaDevices.getUserMedia()microphone capture.
4. Duplex Voice Pipeline, Idle Routing & Silero VAD Architecture
A. Silero VAD Processor Invariant (Pipecat 1.2+)
FastAPIWebsocketParamsLimitation: In Pipecat 1.2+, WebSocket transports (FastAPIWebsocketTransport) do not run VAD internally. Passingvad_analyzertoFastAPIWebsocketParamsis ignored or raises aTypeError.- Canonical VAD Pipeline Pattern:
- Instantiate
SileroVADAnalyzerwith conversational parameters:from pipecat.audio.vad.silero import SileroVADAnalyzer from pipecat.audio.vad.vad_analyzer import VADParams from pipecat.processors.audio.vad_processor import VADProcessor vad_analyzer = SileroVADAnalyzer( params=VADParams( confidence=0.7, # Strict speech probability threshold (filters breaths/clicks) start_secs=0.2, # 200ms speech onset confirmation stop_secs=0.4, # 400ms natural conversational pause threshold min_volume=0.6, # Minimum volume threshold to ignore ambient room noise ) ) vad_processor = VADProcessor(vad_analyzer=vad_analyzer) - Insert
vad_processorimmediately afterstart_triggerinpipeline_elementsto emitVADUserStartedSpeakingFrameandVADUserStoppedSpeakingFramedownstream:pipeline_elements = [ transport.input(), start_trigger, vad_processor, stt, TranscriptionBroadcaster(participant="User"), context_aggregator.user(), llm, TranscriptionBroadcaster(participant="Bot"), tts, context_aggregator.assistant(), transport.output() ] - Pass
vad_analyzertoLLMUserAggregatorParamsso turn-taking boundaries synchronize with LLM context aggregation:user_params = LLMUserAggregatorParams(vad_analyzer=vad_analyzer) context_aggregator = LLMContextAggregatorPair(context, user_params=user_params)
- Instantiate
B. Greeting Turn Invariant (Anti-Duplicate Audio)
- In
StartTriggerProcessor.process_frame(), push ONLYLLMRunFrame(). - NEVER push both
LLMContextFrameandLLMRunFrame().LLMUserAggregatoremits context automatically upon receivingLLMRunFrame(); pushing both triggers duplicate parallel greeting speech.
C. Pipecat Upstream Frame Routing & User Idle Management
- Upstream Silence Blindness Hazard: Pipecat transports push audio downstream to the LLM, and the LLM emits
TranscriptionFrame,UserStartedSpeakingFrame, andInterruptionFramedownstream. AnyFrameProcessor(such asUserIdleProcessor) placed upstream of the LLM never sees downstream transcription frames. A short idle timeout (such as 5s) will mistakenly treat active user speech as silence and disconnect the call. - The Solution: Wire direct activity reporting from the LLM service to the idle processor:
# agent_live.py if hasattr(self, "user_idle_processor") and self.user_idle_processor: self.user_idle_processor.record_activity("speech_transcription") - Cadence Guidelines: Use a 15s idle timer with check-in prompts at 15s, 30s, and 45s, disconnecting only after 60s of uninterrupted silence.
D. Pipecat WebSocket Receive Loop Invariant: Never Chain server_content Fields with elif
- The Terminal Frame Co-Occurrence Trap: In
google-genaiAsyncSession.receive(), Gemini Live frequently sends a terminal WebSocket frame where bothmessage.server_content.model_turn(the final audio chunk) andmessage.server_content.turn_complete = True(oroutput_transcription) are populated on the same message. - Anti-Pattern (
elifChain): Chainingif sc.interrupted: ... elif sc.model_turn: ... elif sc.turn_complete: ... elif sc.output_transcription:causes_handle_msg_turn_completeand_handle_msg_output_transcriptionto be skipped whenevermodel_turnis present on the final frame. Consequently, Pipecat never emitsLLMFullResponseEndFramefor that turn, stallingLLMAssistantResponseAggregatorand mis-bucketing per-turn usage rollups. - Canonical Pattern (Independent
ifBlocks):sc = message.server_content if sc: if sc.interrupted: await host.broadcast_interruption() if sc.model_turn: await host._handle_msg_model_turn(message) if sc.input_transcription: await host._handle_msg_input_transcription(message) if sc.output_transcription: await host._handle_msg_output_transcription(message) if sc.grounding_metadata: await host._handle_msg_grounding_metadata(message) if sc.turn_complete: await host._handle_msg_turn_complete(message) if message.tool_call: await host._handle_msg_tool_call(message)
E. Pipecat ContextWindowCompressionConfig & SlidingWindow Wiring Invariant
- Silent No-Op Anti-Pattern: Passing a plain dictionary (
{"enabled": True, "trigger_tokens": ...}) or storingtarget_tokensinsidekwargs["extra"]onGeminiLiveInputParamsis silently ignored by Pipecat, leaving long calls with unbounded audio and tool-result history. - Canonical Pattern: Always construct a typed
types.ContextWindowCompressionConfigwithtrigger_tokens=5000(the Vertex AI minimum trigger floor) andtypes.SlidingWindow(target_tokens=3500):from google.genai import types if compression_enabled and trigger_tokens: target_tokens = target_tokens or 3500 kwargs["context_window_compression"] = types.ContextWindowCompressionConfig( trigger_tokens=trigger_tokens, sliding_window=types.SlidingWindow(target_tokens=target_tokens), )
5. Live Observability & Diagnostic Buffer
- Zero-Disk In-Memory Ring Buffer: Intercept Loguru and Python standard logging into a bounded
deque(maxlen=1500)indiagnostic_buffer.py, exposed viaGET /api/logsandPOST /api/logs/clear(polled every 1.5s by the debug UI). - Token Fragment Filtering: Never log single-word streaming fragments:
- User: Flush complete sentences upon end-of-sentence punctuation (
.,เฅค,?,!) or VAD stop. - Bot: Accumulate streaming output chunks in
_bot_turn_text_buffer; log the complete turn upon_handle_msg_turn_completeor_handle_msg_interrupted.
- User: Flush complete sentences upon end-of-sentence punctuation (
6. Latency Measurement & Telemetry Invariants
A. First-Packet Live TTFB (Gemini Live)
Calculate Time-to-First-Byte (TTFB) on whichever arrives first from the model turn (text transcription or PCM audio chunk):
if getattr(self, "_current_turn_ttft", None) is None and getattr(self, "_my_ttfb_start", None) is not None:
self._current_turn_ttft = time.time() - self._my_ttfb_start
self._my_ttfb_start = None
B. Speech-to-Text Latency & Turnaround Formulas
-
Cloud Speech v2 (
CustomGoogleSTTService):- The 2ms Audio Stream Trap: Continuous raw audio frames flow every 20ms during silence. Calculating
now - last_audio_timeyields bogus ~2ms readings. - Exact Speech Offset Formula:
STT_Latency = t_now - (T_stream_start + result_end_offset.total_seconds()) datetime.timedeltaSDK Invariant: In the Google Cloud Speech v2 Python SDK,result_end_offsetis a nativedatetime.timedeltaobject (dur.total_seconds()), not a protobuf duration with.nanos.
- The 2ms Audio Stream Trap: Continuous raw audio frames flow every 20ms during silence. Calculating
-
Streaming Live STT (
CustomGeminiTranscribeLiveService):- Real-Time Turnaround Formula:
- If user has finished speaking:
STT_Latency = t_now - t_user_stopped_speaking(typically 80ms to 350ms). - If speech is ongoing in real time:
STT_Latency = t_now - t_last_audio_sent(typically 50ms to 150ms).
- If user has finished speaking:
- The Utterance Duration Anti-Pattern: Never calculate
t_now - t_user_started_speakingas STT latency. That measures the total duration of the spoken sentence (such as 7.8s to 9.1s) rather than speech processing turnaround delay. - Turn State Reset Invariant: Immediately upon emitting
TranscriptionFrameand calling_handle_transcription(is_final=True), resetself._user_started_speaking_time = Noneandself._user_stopped_speaking_time = None.
- Real-Time Turnaround Formula:
C. Strict Metric Binding, Turn Buffering & Token Event Normalization
- Label Binding:
isCascade = this.connectedBotType === "tts-llm-stt".- Gemini Live: Displays
Live TTFB: {ms}ms. - STT-LLM-TTS: Displays
STT: {ms}ms | LLM TTFB: {ms}ms | TTS: {ms}ms.
- Gemini Live: Displays
- Async Turn Buffering (
pendingLLMLatency): Store early-arriving LLM latencies in a buffer and bind them when creating the bot DOM bubble so they never attach to the previous turn's bubble. - Cross-Turn Metric Flushing & Isolation:
resetMetrics()must purge all latency buffers on every connection start.- Clear
this.lastTurnSTTLatency = nullimmediately after rendering on a bubble so previous turn metrics never bleed into subsequent turns. - Cap TTFB turnaround timers at
< 15.0sto prevent idle timeout skew.
- Token Accumulation & Modality Enum Normalization:
- Accumulate token and cost metrics strictly once per turn on the discrete
usageevent (turnComplete), never inside streaming audio chunk callbacks (onMessage) which fire 15 to 30 times per turn. - Normalize Python GenAI SDK
MediaModalityenums (str(m.modality).split(".")[-1].upper()) before JSON serialization so keys like"mediamodality.text"never break modality rate-card lookups in TypeScript.
- Accumulate token and cost metrics strictly once per turn on the discrete
7. Tool-Calling Architecture, AntiCancel Shield & CallSlots Pattern
A. Session Resumption Handles vs. Stable Session IDs
- Ephemeral Handles: The Live API mints a new
session_resumption_handleafter every turn. - Stable Session ID: All turn handles map to a single immutable Session ID (
setupComplete.sessionId). Index logs, CRM records, and observability traces by the stable Session ID.
B. Eliminating Redundant Tool-Call Loops
- Application-Layer State Machines: Move deterministic state routing into the Python or client application layer.
- Lean System Instructions: Keep persona, voice style, tone, and safety in the root System Instruction; strip out heavy multi-branch decision trees.
- Dynamic Tool Allow-Lists: Expose only the minimal subset of tools valid for the active conversation state.
C. Non-Blocking Function Calling (behavior: "NON_BLOCKING")
- Setup Declaration: Declare
"behavior": "NON_BLOCKING"infunction_declarationsso the model can speak a natural filler phrase ("Let me check your account details...") while your backend executes the API call. - Avoid Prompt Hijacking: Never intercept tool calls to dispatch
send_client_content("speak exactly...")withturn_complete=True. - Correct Pattern:
- Guide the persona in the System Instruction to acknowledge the lookup naturally before emitting the tool call.
- Execute the tool asynchronously in the background via
asyncio.create_task(...). - Send
tool_response(FunctionResponse) withscheduling="WHEN_IDLE"over the live WebSocket upon completion so the tool result is voiced as soon as the current filler audio finishes without clipping the model's sentence.
D. Tool Response Scheduling Protocol (WHEN_IDLE vs. SILENT vs. INTERRUPT)
WHEN_IDLE(Default for User-Facing Lookups): Queues the tool result until the model finishes speaking its current filler sentence, then immediately generates the spoken answer.SILENT(For Background State Sync Only): Absorbs the tool payload into session context without generating new speech.- Caution on Gemini 3.1 Flash Live: If the model speaks a filler sentence ("One moment while I check..."), finishes its turn (
turnComplete), and then receives aFunctionResponsewithscheduling="SILENT", the model will remain silent until the user speaks again. Always useWHEN_IDLEwhenever the caller expects an immediate spoken answer after a filler phrase.
- Caution on Gemini 3.1 Flash Live: If the model speaks a filler sentence ("One moment while I check..."), finishes its turn (
func_response = types.FunctionResponse(
id=function_id,
name=function_name,
response=result_payload,
scheduling=types.FunctionResponseScheduling.WHEN_IDLE,
)
E. Pipecat AntiCancel Tool Shield, 8.0s Watchdog & Terminal speaking_Hold
When a user coughs or says "okay" while an asynchronous tool call (such as an order lookup or appointment booking) is in flight, Pipecat's default interruption handler cancels the running asyncio task, dropping the API write mid-flight. Protect tool execution with three guards:
- Override
_cancel_function_call(self, function_name: str | None): Log[AntiCancel] Refusing to cancel in-flight function calland return cleanly without cancelling the backgroundasyncio.Task. - Frame Interception with 8.0s Watchdog (
process_frame): While_tools_in_flight()is active (_active_tools_in_flight > 0withinTOOL_LOCK_MAX_HOLD_SECS = 8.0s), suppressInterruptionFrame,UserStartedSpeakingFrame, andFunctionCallCancelFrame. Never suppressCancelFrame(the pipeline disconnect/shutdown signal, or containers will leak on disconnect). - Client-Side
speaking_HoldDuring Terminal Actions (transfer_call/end_call): When the model invokes a terminal tool while speaking a farewell ("Please hold while I connect you to a specialist..."), do not tear down the audio stream while farewell speech is still buffered in the browser. Enter a client-sidespeaking_Holdstate: pause mic forwarding (send_realtime_input), wait for the client to emitplayback_drainedonce theAudioWorkletring buffer empties, and only then execute the handoff or hangup.
F. The CallSlots Server-State Tool Handler Pattern (Eliminating the 3-Part Tool Token Tax)
On Gemini 3.1 Flash Live, LiveConnectConfig.tools schemas and tool-call responses in session history are re-billed on every subsequent turn. To prevent a multi-tool workflow from adding 2,000+ tokens/turn and 4,800+ tokens of tool-response history bloat:
- Never Auto-Inject Tool Lists (
live_tool_names) intosystem_instruction:- Tool declarations belong exclusively in
LiveConnectConfig.tools. Auto-generating an"Available Live API Tools"directory intosystem_instructionduplicates billing by+650 to 800 tokens/turn. Setlive_tool_names=None.
- Tool declarations belong exclusively in
- Strictly Whitelist
FunctionDeclarationFields (Never Leak HTTP Metadata or Constants):- When converting REST tool configs into
google.genai.types.FunctionDeclarationobjects, stripmethod,url,headers, and static server fields. - Move all static or session-scoped parameters into the server Python handler rather than asking the model to generate constants.
- When converting REST tool configs into
- Cache Lookup Objects by Monotonic Short IDs (
EC-1,EC-2) & Return Compact~25-TokenSummaries:- Returning a raw
~800-tokenJSON array from a lookup tool (such as a branch or clinic finder) leaves those 800 tokens in the Live conversation context for every remaining turn (800 tok * 6 remaining turns = ~4,800 tok/call). - Instead, cache the full dictionary objects in server session state (
CallSlots.centres) under monotonic query-scoped short IDs (EC-1..EC-3for pincode 1;EC-4..EC-6if the caller asks about a second pincode) and return a~25-tokenspeakable string ("EC-1: Koramangala (2.1 km) | EC-2: Indiranagar (4.0 km)"). - Always guard lookups with
store = self.centres.get(centre_id)so a hallucinated ID (EC-9) returns a speakable recovery prompt rather than raising an unhandledKeyError.
- Returning a raw
- Collapse Deterministic Post-Action Tool Chains into the Python Handler:
- If booking an appointment is always followed by sending a messaging confirmation (
232 tok/turnschema) and returning branch contact details (183 tok/turnschema), execute both follow-up API calls directly insidehandle_create_appointment_booking()in Python. This removes 2 tool schemas fromLiveConnectConfig.toolsand collapses 3 sequential LLM tool turns into 1 Python round-trip.
- If booking an appointment is always followed by sending a messaging confirmation (
8. Guided State Machine & Phase Engine Architecture
A. Core Architectural Pattern
Structure complex enterprise calls into discrete, predictable Intent Phases:
- Definite Intent Principle: Define
definite_intent(business goal),conversational_boundaries(forbidden actions), andallowed_toolsfor each phase. - Phase-Specific Tool Whitelisting: Restrict execution of action or calculation tools to designated phases. Phases for discovery, consent, and education keep
allowed_tools: []. - Zero Search Tools for Core Domain Directives: Ground core brand credentials, regulatory disclosures, and safety rules directly in conversational prompt cards with zero tool calls.
@dataclass
class PhaseDefinition:
phase_id: int
title: str
definite_intent: str # Explicit purpose of this phase
allowed_tools: List[str] # Whitelisted tools for this state
conversational_directive: str # Spoken behavioral instructions and guidelines
fast_path_keywords: List[str] # 0ms regex triggers for immediate state switching
B. Dual-Tier Dynamic Phase Router
+---------------------------------------------------+
| User Utterance (Live Transcription) |
+---------------------------------------------------+
|
v
+---------------------------------------------------+
| Tier 1: Fast-Path Regex (0ms) |
+---------------------------------------------------+
| |
[Match Found] [No Regex Match]
| |
| v
| +-----------------------------------+
| | Tier 2: Async Semantic Classifier |
| | (Flash-Lite, <120ms) |
| +-----------------------------------+
| |
| [Confidence >= 0.70]
| |
+---------------+---------------+
|
v
+---------------------------------------------------+
| await transition_to(target_phase) |
| Update Active Phase State & Metrics |
+---------------------------------------------------+
- Tier 1 (Instant 0ms Fast-Path): Regex evaluation on transcribed user sentences for high-confidence explicit intents (
$0cost,0mslatency). - Tier 2 (Async Background LLM Classifier): Non-blocking semantic classification using
gemini-2.5-flash-litewhen Tier 1 finds no match. - Hysteresis & Confidence Gate: Only transition states if
confidence >= 0.70and the target phase transition is valid.
C. JIT Prompt Yielding (clientContent) & Mid-Speech Collision Invariant
- Critical Protocol Truth (
BidiGenerateContentClientContentInterruption Hazard): SendingBidiGenerateContentClientContentduring active model generation unconditionally interrupts the model's in-flight speech (turn_complete=Falseonly avoids starting a new response turn; it does not prevent cutting off active audio generation). - Mid-Speech Guard & Queue Flushing: Always queue phase transitions that occur while
self._is_bot_speakingis true, and flush them insideon_bot_stopped_speakingwithturn_complete=False:
async def transition_to(self, target_phase: int, trigger_reason: str):
async with self._lock:
if self._is_bot_speaking:
self._pending_phase = target_phase
self._pending_reason = trigger_reason
else:
await self._yield_prompt_to_gemini(target_phase, trigger_reason)
async def on_bot_stopped_speaking(self):
self._is_bot_speaking = False
if self._pending_phase is not None:
phase, reason = self._pending_phase, self._pending_reason
self._pending_phase, self._pending_reason = None, None
await self._yield_prompt_to_gemini(phase, reason)
D. Direct Transcription Hook & Dual-Stream UI Invariant
Override _push_user_transcription inside the LLM service to stream User bubbles to the client UI, record observability traces, and trigger the state machine:
async def _push_user_transcription(self, text: str, result=None):
await super()._push_user_transcription(text, result)
if text and text.strip():
clean_text = text.strip()
await self.push_frame(OutputTransportMessageFrame(message={
"label": "rtvi-ai", "type": "server-message",
"data": {"type": "transcription", "participant": "User", "text": clean_text}
}))
GLOBAL_LANGSMITH_TRACER.record_user_turn(clean_text)
if hasattr(self, "phase_tracker") and self.phase_tracker:
await self.phase_tracker.handle_user_transcript(clean_text)
E. Deterministic Transition Gates & Consistency Ledger
- Stage Skip Guard: Block anomalous forward jumps greater than
Nstages (such as max+3) in a single turn. - Prerequisite Fact Gating: Gated phases cannot be entered until required facts are recorded in the session
FactStore. - Numeric Consistency Ledger: Record all bot-quoted numbers and validate bounds before quoting to guarantee 100% cross-turn consistency.
9. State Machine Disambiguation & Anti-Looping Discipline
A. Disambiguating Turn 1 Greetings vs. Mid-Call Farewells
- The Greeting Loop Anti-Pattern: If callback requests or wrap-up phrases route back to Phase 1, the bot re-introduces itself from the start of the script.
- Strict Funnel Invariant:
- Phase 1 (Opening & Availability): Restricted strictly to the Turn 1 opening. Never route back to Phase 1 mid-call.
- Phase 9 (Commitment & Close): Any mid-call concluding signal, farewell, or callback request (
"bye","alvida","thank you bye","chalo bye","baad mein baat karte hain") must route to Phase 9.
- Tier-1 Fast-Path Regex for Farewells (0ms, $0):
if self.current_phase > 1 and re.search(r" (bye|boy|by|alvida|thank you|thanks|chalo bye|ok bye|theek hai bye|wrap up|chalta hu|chalti hu|rakhta hu|rakhti hu|baad mein baat|later) ", lower): await self.transition_to(9, trigger_reason="Tier-1 Regex: User wrapping up call / farewell")
B. Dynamic Commitment Validation (No Rigid Scripts)
- In Phase 9, dynamically remind the customer of the specific option explored in the current session (such as the 12-month return plan, minimum deposit, or appointment slot) and confirm whether they are ready to proceed or prefer a scheduled follow-up. Close cleanly with zero loops.
10. Speaker Persona & Grammatical Agreement across Prompt Cards
A. The JIT Prompt Card Persona Drift Trap
When Prompt Cards are injected Just-In-Time via send_client_content(turn_complete=False), the prompt card becomes the most recent instruction in context. If speaker identity and grammatical gender constraints are omitted from the card, multilingual models frequently drift into default masculine Hindi verb forms ("เคฌเคคเคพ เคฐเคนเคพ เคนเฅเค", "เคฆเฅเคคเคพ เคนเฅเค").
B. Mandatory Prompt Card Header Standard
Every dynamically yielded Prompt Card must include the speaker identity and grammatical agreement rules using natural-language prefixes (never square brackets, which native audio models can vocalize aloud):
directive_text = (
f"System update for agent - Active Phase {target_phase} ({card['title']}):\n"
f"Speaker Persona: Pragya (Female Senior Wealth Manager at Cymbal Lending).\n"
f"Mandatory Female Grammar: Always speak in 100% consistent feminine Hindi grammar for yourself.\n"
f"โข REQUIRED FEMININE VERBS: 'เคฎเฅเค เคฌเคคเคพ เคฐเคนเฅ เคนเฅเค', 'เคเคฐเคคเฅ เคนเฅเค', 'เคฆเฅเคคเฅ เคนเฅเค', 'เคฎเคฆเคฆ เคเคฐเฅเคเคเฅ', 'เคธเคฎเค เคเค'\n"
f"โข FORBIDDEN MASCULINE VERBS: Never use masculine verb forms ('เคฐเคนเคพ เคนเฅเค', 'เคเคฐเคคเคพ เคนเฅเค', 'เคฆเฅเคคเคพ เคนเฅเค', 'เคเคฐเฅเคเคเคพ').\n\n"
f"{card['directive']}\n\n"
f"Context: {trigger_reason}\n"
f"Rule: Always use Devanagari for Hindi words and Latin script for English financial terms."
)
11. Sub-Millisecond Near-Process Dual-Layer RAG Caching
A. The Remote Vector Search Dead Air Hazard
- The Anti-Pattern: Calling remote vector database search APIs inside an active voice turn takes 3,200ms to 3,800ms, introducing over 3 seconds of dead air.
- Dual-Layer Near-Process Hierarchy:
- L1 In-Memory Process RAM (
<0.05ms): Compile domain Q&A pairs into an Okapi BM25 inverted index in process memory on startup for sub-millisecond retrieval. - L2 Cloud Memorystore Valkey / Redis (
~1.0ms): Use a regional GCP Memorystore instance inus-central1for cross-instance shared state and normalized query hash keys (cymbal:sheet_rag:*). - Graceful Fallback: If L2 Redis is unreachable, fall back immediately to the local L1 RAM BM25 index without blocking the live voice stream.
- L1 In-Memory Process RAM (
- Proactive vs. Reactive RAG Injection:
- Proactive Background RAG (Between Turns): Inject retrieved context silently via
LiveClientContent(turns=[Content(role="user", parts=[Part.from_text("System context update: ...")])], turn_complete=False). Becauseturn_complete=False, the model absorbs the context without generating unsolicited speech. - Reactive On-Demand RAG (User Asks a Specific Question): Expose a non-blocking
search_knowledge_basetool returningFunctionResponse(..., scheduling="WHEN_IDLE").
- Proactive Background RAG (Between Turns): Inject retrieved context silently via
12. Dual-Brain Frontcar / Downcar Architecture & Enterprise Memory Bank
A. Frontcar vs. Downcar Decoupling
- Frontcar (Fast Path - Synchronous Duplex Voice): Ultra-lean real-time pipeline (
gemini-3.1-flash-live-previeworgemini-live-2.5-flash-native-audio) focused on conversational persona, audio pacing, active phase intent, and non-blocking tool execution. At session connect (t = 0ms), inject the user's compact Markdown profile (< 500tokens) intosystem_instruction. Zero heavy database writes occur on the audio thread. - Downcar (Slow Path - Asynchronous Post-Session Worker): Triggered on
on_client_disconnected. Usesgemini-2.5-flash-lite(response_mime_type="application/json") to parse session transcripts, extract canonical user facts and action items, generate episodic summaries, and persist them to your cloud memory store.
B. Dual-Threshold Vector Memory Standard
- Deduplication Gate (
SIMILARITY_THRESHOLD = 0.83): Enforcecosine_similarity >= 0.83(plus category match and MD5content_hash) when updating or inserting memory facts so distinct concepts never overwrite each other. - Retrieval Gate (
RETRIEVAL_THRESHOLD = 0.40): For semantic recall (retrieve_memory), usethreshold = 0.40. Conversational questions embed at0.65 to 0.78cosine similarity against declarative facts; using0.83for retrieval causes false-negative recall misses. - Stateless Container Rule: Never store user facts in local container files (
/tmpor local SQLite) on Cloud Run. Persist to a managed cloud store (Firestore,Cloud SQL, orVertex AI Agent Engine Memory Bank).
C. Spoken Lexical User Identity Normalization
Standardize spoken introductions by stripping conversational framing ("My name is..."), removing boundary-aware honorifics (Dr., Mr., Smt.), and sanitizing symbols while preserving alphanumeric IDs (user_<name>).
13. Token Modality, Cost Accounting & Context Compression Architecture
A. Published Rate Card (Vertex AI & Google AI Studio)
| Token Category | Gemini 2.5 Flash Native Audio (per 1M Tokens) | Gemini 3.1 Flash Live Preview (per 1M Tokens) | Modality & Billing Mechanics |
|---|---|---|---|
| Prompt TEXT | $0.50 |
$0.75 |
System instructions, tool declarations (on 3.1 Live), text turns, and tool responses |
| Prompt AUDIO | $3.00 |
$3.00 |
~25 tokens/sec (~1,500 tok/min; budget ~27.8 tok/sec / ~1,668 tok/min / ~$0.0050/min as a conservative planning figure; includes both user speech and carried played model audio) |
| Candidate TEXT / Thoughts | $2.00 |
$4.50 |
Output text transcripts and thinking tokens (thoughts_token_count) |
| Candidate AUDIO | $12.00 |
$12.00 |
~25 tokens/sec of generated model speech (~1,500 tokens/min = $0.018/min) |
B. Dual-Engine Optimization Architecture: 5,000-Token Trigger Floor vs. Lean SI & JIT Prompt Cards
- The 5,000-Token Trigger Invariant: Vertex AI Live enforces a minimum threshold of 5,000 tokens (
trigger_tokens=5000,target_tokens=3500) for sliding-window context compression. - The Lean SI Mismatch:
- In an optimized voice pipeline, the root System Instruction (SI) is intentionally compact (
~225 to 500tokens) containing only core persona, tone, grammatical rules, and safety boundaries. - Because the Live protocol is append-only, every injected prompt card and every turn of user and bot speech accumulates in session history until the
5,000-token trigger is reached (typically around Turn 7 to 12).
- In an optimized voice pipeline, the root System Instruction (SI) is intentionally compact (
- The Dual-Engine Solution:
- Early Turns (Lean Static SI + Dynamic JIT Prompt Cards):
- Keep the root SI under
500tokens (120tokens for the opening card,~500tokens average across phases). - Yield Phase Directives and compliance scripts Just-In-Time via
send_client_content(turn_complete=False)only when entering a new conversation phase. - Never stuff static multi-page decision trees or unused tool schemas into the root SI.
- Keep the root SI under
- Long Calls (Sliding-Window Compression
5000 / 3500+FactStore):- Once total context reaches
5,000tokens,SlidingWindow(target_tokens=3500)automatically evicts the oldest dialogue turns (including their accumulated user and model audio tokens) while preserving the initialsystem_instruction. FactStore(< 60tokens) restores structured continuity across compaction boundaries.
- Once total context reaches
- Early Turns (Lean Static SI + Dynamic JIT Prompt Cards):
C. Bidirectional Audio Accumulation & Sliding-Window Eviction Dynamics
- Why
Prompt AUDIOGrows Rapidly (UserInputAudio + PlayedModelOutputAudio):- Unlike text-only APIs where prior assistant replies are compact text tokens, a native Speech-to-Speech session retains both the user's input speech audio and the model's played output audio in the active conversation history (
Prompt AUDIO, billed at$3.00 / 1Mon every subsequent turn), alongside their text transcripts inPrompt TEXT:Prompt_AUDIO(turn k) = sum_{i=1..k} UserInputAudio(i) + sum_{i=1..k-1} PlayedModelOutputAudio(i) - Production Telemetry Proof (Vertex AI
gemini-live-2.5-flash-native-audio, 3 Tools Declared):- Turn 1:
Prompt = 279(236 TEXT=211SI +25user ASR text,43 AUDIOuser greeting),Response = 223(181 AUDIOplayed bot speech,42 TEXTbot transcript). - Turn 2:
Prompt = 556(AUDIO = 271=43Turn 1 user audio +181Turn 1 played bot audio +47Turn 2 user audio;TEXT = 285=236Turn 1 text +42Turn 1 bot text +7Turn 2 user ASR text),Response = 212(170 AUDIO,42 TEXT).
- Turn 1:
- Unlike text-only APIs where prior assistant replies are compact text tokens, a native Speech-to-Speech session retains both the user's input speech audio and the model's played output audio in the active conversation history (
- Barge-In Truncation Reduces Subsequent
Prompt AUDIO:- When a caller interrupts the model mid-sentence (barge-in), the server truncates the unplayed tail of the model's audio turn from session history based on playback duration. On the next turn, only the model audio tokens actually played before the interruption are carried forward into
Prompt AUDIO. - Automated Benchmark Trap: Because the server finishes streaming audio packets over the WebSocket faster than real-time playback duration (
generated_audio_duration), an automated test script that sendsclient_content(turn_complete=True)immediately upon receivingturnCompletewithout waiting for real-time audio playout triggers server-side barge-in truncation, slicing the model's audio out of history and under-measuring real productionPrompt AUDIOgrowth.
- When a caller interrupts the model mid-sentence (barge-in), the server truncates the unplayed tail of the model's audio turn from session history based on playback duration. On the next turn, only the model audio tokens actually played before the interruption are carried forward into
- Sliding-Window Audio Eviction:
- When
ContextWindowCompressionConfig(trigger_tokens=5000, sliding_window=SlidingWindow(target_tokens=3500))fires, the server evicts the oldest conversation turns (for example contractingPrompt AUDIOfrom2,689 -> 1,794and2,015 -> 799tokens) while keeping the rootsystem_instructionintact. - Because audio input costs
$3.00 / 1Mvs.$0.50 to $0.75 / 1Mfor text input (4x to 6xrate ratio), evicting1,216audio tokens saves the cost equivalent of4,864 to 7,296text tokens.
- When
D. FactStore Restoration Protocol & 10-Turn Cap Invariant
- The Context Amnesia Footgun: When sliding-window compaction evicts early dialogue turns, the model can forget facts stated in Turns 1 to 3 (such as customer name, pin code, or loan amount) unless those facts are preserved cleanly.
- The
FactStoreInvariant:- Detect compaction via
current_prompt < (last_prompt - 150). - Maintain a compact
< 60-tokenJSONFactStoreof extracted entities plus a capped window of recent transcript lines, and inject them viasend_client_content(turns=[Content(...)], turn_complete=False).
- Detect compaction via
- The 10-Turn Cap Rule:
- Never re-inject the entire session transcript on a 30-turn call (which would inject
1,250+text tokens and defeat compression). - Strictly cap injected dialogue to the last 10 turns (
history[-10:]):capped_history = history[-10:] if len(history) > 10 else history injected_lines = [f"{turn['role']}: {turn['text']}" for turn in capped_history]
- Never re-inject the entire session transcript on a 30-turn call (which would inject
- Prompt Restoration Template (Zero Bracket Tags):
System update for agent - Conversation Transcript Log (Last 10 Turns): <injected_turns> Instructions for agent: - Continue the conversation with the caller from the latest turn. - Retain all customer details, numbers, preferences, and agreements stated above. - Do not repeat greetings, do not re-introduce yourself, and do not verbally acknowledge this transcript update.
E. Total Transcription History Backend Observability Invariant
- Zero Silent Drops: Log every dialogue turn in real time and in session summaries within the backend logger:
- Per-Turn Logs: Log
[Transcript User (Turn N)]: <text>and[Transcript Assistant (Turn N)]: <text>. - At Compaction Injection: Log the complete session transcription history alongside the 10-turn restoration payload.
- At Client Disconnect: In
on_client_disconnected, emit the numbered session transcript (1. User: ...,2. Assistant: ...).
- Per-Turn Logs: Log
F. Audio vs. Text Token Dynamics & Compression Alert Guarding
- Composite Prompt Tokens: On Vertex AI,
prompt_token_countequalsPrompt_TEXT + Prompt_AUDIO(residual = 0). On Google AI Studio (gemini-3.1-flash-live-preview),prompt_token_countexceedsPrompt_TEXT + Prompt_AUDIOby93 to 221tokens per turn (see ยง13.H). Never compute billed volume by summingprompt_tokens_detailsalone; always readprompt_token_countand bill any positive differencemax(0, prompt_token_count - sum(details))at the text input rate. - Guard Observability Alerts: Configure monitoring dashboards so context-compression alerts only fire when
prompt_token_countcrosses>= 4,500tokens followed by a sustained drop (> 1,000tokens).
G. Tool Declaration & Tool Response Token Accounting (Prompt TEXT)
- Gemini 2.5 Flash Native Audio (
0Tool Schema Tokens inprompt_token_count): Ongemini-live-2.5-flash-native-audio, the turn-by-turnprompt_token_countcountssystem_instructiontext tokens (225tok) and omits declaredtools(function_declarations) schemas (0billed tool schema tokens per turn), billing tool payloads only when tools are invoked and theirFunctionResponsepayloads enter conversation history. - Gemini 3.x Live Architecture (Full Developer-Turn Tool Schema Billing + History Carry): On 3.x Live pipelines where
function_declarationsare serialized into the developer system turn (as well as every returnedFunctionResponsein session history), declared tool schemas (~70 to 250tokens per concise tool, or2,300 to 3,500+tokens/turn for 16 verbose tools) and tool outputs are re-billed on every subsequent turn. - Pruning Invariant: Register only the 1 to 3 essential real-time tools required for the active workflow and apply the
CallSlotspattern (ยง7.F) so neither tool schemas nor800-token JSON lookup responses bloat your5,000-token context window.
H. Cross-Endpoint usage_metadata Reconciliation Invariant (Vertex AI vs. AI Studio)
Paired multi-turn sessions on identical Pipecat clients and prompts establish exact reconciliation behaviors across endpoints:
- Google AI Studio (
gemini-3.1-flash-live-preview) UnattributedPromptResidual:T0 (greeting): Prompt: 795 (TEXT 501, AUDIO 201) -> details sum 702 residual 93 T1 : Prompt: 1091 (TEXT 538, AUDIO 418) -> details sum 956 residual 135 T2 : Prompt: 1343 (TEXT 578, AUDIO 594) -> details sum 1172 residual 171 T3 : Prompt: 1692 (TEXT 633, AUDIO 838) -> details sum 1471 residual 221 TOTAL RESIDUAL 620 = 12.6%- On AI Studio,
response_tokens_detailsreportsAUDIOonly (0 TEXTout), whileprompt_token_countincludes internal turn formatting and the carried output text transcript that are not broken out inprompt_tokens_details. - AI Studio also sets
total_token_count = prompt_token_count + response_token_count, omittingthoughts_token_countfromtotal_token_counteven whenthoughts_token_countis populated.
- On AI Studio,
- Vertex AI (
gemini-live-2.5-flash-native-audio) Zero Residual:- On Vertex AI,
prompt_token_count == Prompt_TEXT + Prompt_AUDIO(residual = 0on every turn),response_tokens_detailsexplicitly breaks out bothAUDIOandTEXT, andtotal_token_count == prompt_token_count + response_token_count + thoughts_token_count.
- On Vertex AI,
- Cost-Calculator Invariant:
- Any Live pricing module must compute
residual_input = max(0, prompt_token_count - sum(prompt_tokens_details))and includethoughts_token_countexplicitly when computing total turn cost.
- Any Live pricing module must compute
I. Pre-Speech Greeting Audio Metering on AI Studio (gemini-3.1-flash-live-preview)
On Google AI Studio, gemini-3.1-flash-live-preview bills ~196 to 201 audio prompt tokens on Turn 1 before the user has spoken if the client microphone stream is already open during the bot's greeting playback:
05:58:48.8 WebSocket connected (mic streaming silence)
05:58:49.8 Bot started speaking (greeting)
05:58:56.4 usage_metadata: Prompt: 795 (TEXT: 501, AUDIO: 201) # 7.6s of open mic
Because the Live API re-bills active context on every subsequent turn, those 201 Turn-1 silence tokens compound to 804+ extra audio tokens over 4 turns.
- Mitigation: Gate client microphone forwarding until
TTSStoppedFrame(or client greeting playback drain) fires on Turn 1. When Turn 1 is triggered cleanly without open-mic audio before the greeting completes, Turn 1Prompt AUDIOis0.
J. Zero Implicit Context Caching Discount (cached_content_token_count = 0)
Across all production Gemini Live endpoints (gemini-live-2.5-flash-native-audio on Vertex AI and gemini-3.1-flash-live-preview on Google AI Studio), cached_content_token_count is 0 (None) on every turn. System instructions, declared tool schemas (on 3.1 Live), and carried conversation history in the active window are billed at standard input rates on every turn:
| Model | Turn 1 Prompt TEXT for an Identical 225-Token System Instruction (0 Tools) |
|---|---|
gemini-live-2.5-flash-native-audio |
225 tokens (the system instruction with zero extra scaffolding) |
gemini-3.1-flash-live-preview |
501 tokens (225 instruction + 276 model preamble scaffolding) |
What can look at first glance like "2.5 caching the system prompt" is actually 2.5-flash-native-audio having a smaller fixed preamble (225 vs. 501 tokens) and omitting tool schemas from prompt_token_count. Both models re-bill their active context window on every turn (~2.9x cumulative re-bill amplification over 4 to 5 turns).
K. Thinking Token Accounting (thoughts_token_count)
- On
gemini-live-2.5-flash-native-audiowiththinking_budget=0andgemini-3.1-flash-live-previewwithThinkingLevel.MINIMAL(include_thoughts=False),thoughts_token_countis0(None) on standard conversational turns. - Always keep
thinking_level="MINIMAL"andinclude_thoughts=Falseon 3.1 Live for real-time voice agents; higher thinking levels (LOW,MEDIUM, orHIGH) add both turn latency and output-rate thinking token spend ($4.50 / 1M) across multi-turn calls.
L. Parametric Per-Call Cost Model & Turn-by-Turn Calculator
Because both user speech (u_aud) and played model output audio (b_aud_kept) accumulate in Prompt AUDIO on every turn until SlidingWindow(5000/3500) triggers, and both user ASR text (u_txt) and model output text (b_txt) accumulate in Prompt TEXT:
def calculate_live_session_cost(
si_tokens: int = 500,
tool_tokens: int = 200,
turns: int = 6,
user_sec_per_turn: float = 4.0,
bot_sec_per_turn: float = 6.0,
barge_in_play_ratio: float = 1.0,
model: str = "gemini-3.1-flash-live-preview",
sliding_window_trigger: int = 5000,
sliding_window_target: int = 3500,
) -> dict:
# Turn-by-turn token and USD calculator matching Gemini Live production billing.
u_aud = int(round(user_sec_per_turn * 25))
b_aud_gen = int(round(bot_sec_per_turn * 25))
b_aud_kept = int(round(b_aud_gen * barge_in_play_ratio))
u_txt = int(round(u_aud * 0.15))
b_txt = int(round(b_aud_gen * 0.25))
if model == "gemini-live-2.5-flash-native-audio":
r_txt_in, r_txt_out, r_aud_in, r_aud_out = 0.50, 2.00, 3.00, 12.00
base_txt = si_tokens # 2.5 omits static tool_tokens from prompt_token_count
thoughts_per_turn = 0
else:
r_txt_in, r_txt_out, r_aud_in, r_aud_out = 0.75, 4.50, 3.00, 12.00
base_txt = si_tokens + 276 + tool_tokens # 3.1 adds +276 preamble scaffolding
thoughts_per_turn = 0 # 0 with ThinkingLevel.MINIMAL
cum_txt_in = 0
cum_aud_in = 0
hist_turns = [] # list of (turn_txt, turn_aud)
for _ in range(1, turns + 1):
hist_txt = sum(x[0] for x in hist_turns)
hist_aud = sum(x[1] for x in hist_turns)
cur_txt = base_txt + hist_txt + u_txt
cur_aud = hist_aud + u_aud
if sliding_window_trigger and (cur_txt + cur_aud) > sliding_window_trigger:
while hist_turns and ( base_txt + sum(x[0] for x in hist_turns) + u_txt + sum(x[1] for x in hist_turns) + u_aud ) > sliding_window_target:
hist_turns.pop(0)
cur_txt = base_txt + sum(x[0] for x in hist_turns) + u_txt
cur_aud = sum(x[1] for x in hist_turns) + u_aud
cum_txt_in += cur_txt
cum_aud_in += cur_aud
hist_turns.append((u_txt + b_txt, u_aud + b_aud_kept))
resp_aud = turns * b_aud_gen
resp_txt = turns * (b_txt + thoughts_per_turn)
usd = (
cum_txt_in * r_txt_in
+ cum_aud_in * r_aud_in
+ resp_aud * r_aud_out
+ resp_txt * r_txt_out
) / 1e6
return {
"prompt_text_tokens": cum_txt_in,
"prompt_audio_tokens": cum_aud_in,
"response_audio_tokens": resp_aud,
"response_text_and_thought_tokens": resp_txt,
"total_usd": round(usd, 5),
}
M. Duplex Voice Compounding Economics: The "Carried Audio Tax" & Who Drives Call Cost
Empirical measurements on an 18-turn uncompressed session (54,315 total billed tokens, $0.1178 on Gemini 2.5 Flash Native Audio) illustrate why carried audio history dominates multi-turn voice spend:
+------------------------------------------------------------------------+
| Audio Prompt Input (Carried Audio History) : 25,365 tok -> $0.0761 (64.6%) |
| Audio Response Output (Generated Speech) : 2,268 tok -> $0.0272 (23.1%) |
| Text Prompt Input (System & Cards) : 25,904 tok -> $0.0130 (11.0%) |
| Text Response Output (Transcripts) : 778 tok -> $0.0016 (1.3%) |
| -------------------------------------------------------------------- |
| TOTAL : 54,315 tok -> $0.1178 (100%) |
+------------------------------------------------------------------------+
- Why
$12.00 / 1MAudio Output Is Only 23.1% of Total Spend:- Although generated bot speech has the highest unit price (
$12.00 / 1Mtokens =$0.018 / min),Candidate AUDIOis billed only once on the turn it is spoken.
- Although generated bot speech has the highest unit price (
- Why
Prompt AUDIO($3.00 / 1M) Drives 64.6% of the Call Bill:- Every turn's user audio and played bot audio remain in
Prompt AUDIOand are re-billed at$3.00 / 1Mon every subsequent turn untilSlidingWindow(5000/3500)evicts old turns. - By Turn 14 of an uncompressed call (
2,363accumulated audio tokens in prompt), even a one-word user acknowledgment ("Yes") re-processes all2,363prior audio tokens at$3.00 / 1M. - Note why the 18-turn trace above cost
$0.1178: it already used a compact~500-token root prompt (25,904cumulative text tokens across 18 turns) rather than a monolithic4,500-token static SOP (81,000cumulative text tokens). - Combining all three optimization levers (Pillar 1
SlidingWindow(5000/3500)for long calls + Pillar 2 JIT Prompt Cards capping active instructions at~500tokens + Pillar 3 rolling audio-to-text pruning /FactStore< 60tokens for calls under 15 turns) reduces a 3-minute, 18-turn enterprise call from$0.246($0.082/minwith a 4,500-token static prompt and unpruned audio) down to$0.093($0.031/min, a 62% reduction).
- Every turn's user audio and played bot audio remain in
N. Interruption Economics & Barge-In Math (Output Savings vs. Extra-Turn Prompt Fee)
When a user interrupts the bot mid-speech (server_content.interrupted == True), two opposing cost effects occur:
- What You Save (Aborted Output Audio + Truncated History):
- The server stops generating remaining audio (
saving $12.00 / 1Mon ungenerated output tokens) and truncates the unplayed tail of the bot's turn so unplayed audio tokens are not carried forward into futurePrompt AUDIOturns. - Example: Cutting off a
300-token bot monologue after50played tokens saves250output audio tokens ($0.0030) plus250 * $3.00 / 1M = $0.00075on every subsequent turn.
- The server stops generating remaining audio (
- The Hidden Trap (The "Micro-Interruption Prompt Fee"):
- Input prompt evaluation for the interrupted turn has already occurred and been billed before the first audio chunk was generated.
- If the interruption was a false barge-in (a cough, throat clear, or backchannel "uh-huh" that triggers an extra conversational turn), the model must run a brand-new prompt evaluation across the entire active context window (
3,500 to 4,000tokens):Extra Turn Prompt Cost = (2,300 audio tok * $3.00/1M) + (1,700 text tok * $0.50/1M) = $0.00775 Net Impact of False Barge-In = $0.00300 saved - $0.00775 extra turn = -$0.00475 net loss
- Golden Rule of Interruption:
- Decisive Interruption (Saves Money): Caller cuts off an unwanted explanation and pivots directly to the next step ("Skip that, book the 10 AM slot"), replacing a turn and truncating carried bot audio.
- Fragmented False Barge-In (Burns Money): Coughs or backchannel fillers (
"uh-huh") split one turn into two, re-billing the entire context history. TuneStartSensitivity.START_SENSITIVITY_LOWandSileroVADAnalyzer(confidence=0.7, start_secs=0.2)to filter out non-speech bursts.
O. Model Comparison & Production Decision Matrix (2.5 Native Audio vs. 3.1 Flash Live)
| Dimension | gemini-live-2.5-flash-native-audio (GA) |
gemini-3.1-flash-live-preview (Preview) |
|---|---|---|
| Deployment Endpoints | Vertex AI (us-central1, v1beta1) |
Google AI Studio (v1alpha / v1beta) |
| Pricing Rates (In / Out) | Text: $0.50 / $2.00, Audio: $3.00 / $12.00 |
Text: $0.75 / $4.50, Audio: $3.00 / $12.00 |
| Fixed Prompt Scaffolding (T0) | 225 tokens (0 extra preamble scaffolding) |
501 tokens (+276 model preamble scaffolding) |
| Tool Schema & Response Accounting | 0 billed tokens for static function_declarations; bills FunctionResponse payloads in history |
Bills +276 preamble tokens, all declared tools schemas, and every FunctionResponse payload in history on every turn (CallSlots pruning required) |
| Thinking Configuration | ThinkingConfig(thinking_budget=0) |
ThinkingConfig(thinking_level="MINIMAL") (never pass thinking_budget=0) |
| Proactive Audio & Affective Dialog | Supported (proactive_audio=True, enable_affective_dialog=True) |
Unsupported (causes 1007 handshake error if set) |
| Usage Metadata Reconciliation | 100% exact (residual = 0 on every turn) |
93 to 221 tok/turn unattributed residual on AI Studio (prompt_token_count - sum(details)) |
| Ideal Use Case | Cost-sensitive high-volume voice bots and affective dialogue | Low-latency multi-step reasoning, expressive multilingual voice agents, and dynamic tool workflows |
P. Enterprise Voice Agent Design Principles
- Target 6 to 8 Turns for Transactional Funnels:
- A focused 6-turn booking or verification call stays well below the 5,000-token compaction trigger and costs
~$0.025 to $0.035in total.
- A focused 6-turn booking or verification call stays well below the 5,000-token compaction trigger and costs
- Enable
SlidingWindow(5000, 3500)on Every Session:- Keeps active context between
3,500and5,000tokens so 15-to-30-minute support calls scale linearly (~$0.06/minwithSlidingWindowalone, or~$0.031/minwhen combined with JIT prompt cards and rolling audio-to-text pruning as shown in ยง13.M) rather than quadratically.
- Keeps active context between
- Decouple UI Progress Telemetry from Model Tool Calling:
- Never rely on the voice model to call a
get_phase_cardortrack_progresstool just to update the browser UI. - Run server-side regex and async Flash-Lite transcription tracking (
PhaseTracker) on incoming user transcripts for0msvoice latency,$0Live tool token overhead, and deterministic UI state updates.
- Never rely on the voice model to call a
- Gate the Microphone on Turn 1:
- Keep client mic audio gated until the opening greeting finishes playing (
TTSStoppedFrame) so ambient room noise during the greeting is never metered into Turn 1Prompt AUDIO.
- Keep client mic audio gated until the opening greeting finishes playing (
Q. The Turn-0 Cold-Start Trap: Why Initial Tokens Ballooned to 1,254
Truncated - read the full file at https://github.com/manishkjs/gemini_live_pipecat/blob/83ae466442067cf0d8a5845c1c4462115353fd6d/docs-site/public/gemini-live/SKILL.md.
