DEV Community

Evan Lin
Evan Lin

Posted on Originally published at evanlin.com on

[AI in Action] Gemini 3.8 Live Upgrade: Paying Off Technical Debt and How My Test Script Fooled Me Three Times

image-20260916214342753

Recap

I have a LINE Bot that I use every day, linebot-helper-python. It provides summaries and social media copy for URLs, allows continuous questioning for YouTube links, and handles bookmarks and location queries. One of its features is a LIFF voice assistant: it opens a webpage, connects to the Gemini Live API via WebSocket, and supports push-to-talk or hands-free conversation. The transcribed content is then summarized into a message and pushed back to the LINE chatroom.

In mid-September, Google released Gemini 3.8 Live. After reading the announcement, what I wanted to do was slightly different from what they were trying to sell.

The announcement focused on capabilities: Speech to Speech Index ranked first (82.6 points), Big Bench Audio at 97.7%, automatic detection of 97 languages with the ability to switch mid-conversation, tool calling executing in the background without interrupting the flow, and real-time visual input. There is also a Gemini 3.8 Live Extended Thinking model that can reason while speaking.

My first thought was: Now I can finally remove that preview exception.


The Real Motivation: A Preview Model That Tripped Me Up Twice

In my previous post about agentic video, I mentioned that my intelligent dialogue feature was broken for a while, and no one noticed. The reason was that loader/chat_session.py had gemini-3-pro-preview hardcoded. After that model was removed from Vertex AI, the actual path returned a 404, and the exception handling in that file simply raised the error.

After fixing it, I added a set of guards using tests to enforce that "Model IDs can only appear in config/agent_config.py" and "No preview or experimental models allowed." But those guards had two exceptions:

# Live API only supports gemini-3.1-flash-live-preview; TTS is a separate model family.
# These two places must maintain preview models and are outside the scope of the guards.
EXEMPT_FILES = {"services/voice_live.py", "tools/tts_tool.py"}
Enter fullscreen mode Exit fullscreen mode

The reason for the voice exception was "The Live API only has this one model available." At the time, that was true. But once that sentence is written into a comment, no one checks if it's still valid—the last loophole for the same type of bug sat there in plain sight for months.

The significance of 3.8 Live, therefore, isn't just the capability upgrade. Its model name does not contain -preview, which means that exception can be removed.


Ask the API First, Not the Documentation

This time, I didn't read the documentation first. Instead, I asked the API directly what actually exists:

curl -s "https://generativelanguage.googleapis.com/v1beta/models?key=$KEY&pageSize=200" \
  | jq -r '.models[]?.name' | rg -i 'live'


models/gemini-3.5-transcribe-live
models/gemini-3.1-flash-live-preview
models/gemini-3.8-live
models/gemini-3.8-live-extended-thinking
models/gemini-3.5-live-translate-preview
Enter fullscreen mode Exit fullscreen mode

One command gave me five confirmed model names, which is much more reliable than digging through documentation. gemini-3.8-live is right there, with a clean name.

I also happened to find gemini-3.5-live-translate-preview, which I didn't know about. I was originally thinking about "whether to implement a real-time interpretation mode," and it seems that's a specialized model rather than a brute-forced general model. I'll look into that later.

Next, I queried the metadata:

3.1-flash-live-preview 3.8-live
inputTokenLimit 131072 131072
outputTokenLimit 65536 65536
supportedGenerationMethods bidiGenerateContent bidiGenerateContent

The numbers are identical. At this point, I had initial confidence that it would be a "painless swap," but this was just metadata, not behavior.


Pitfall 1: Warning Against a Non-existent Risk

When I first evaluated this, I solemnly noted: "The announcement didn't mention Vertex AI, and the decision for this project is to switch entirely to Vertex, so it's uncertain if 3.8 Live will be available."

It sounded very professional. Then I looked at my own code:

client = live_genai.Client(
    api_key=GOOGLE_AI_API_KEY,
    vertexai=False,
    http_options={"api_version": "v1beta"},
)
Enter fullscreen mode Exit fullscreen mode

vertexai=False. The voice path has used the Gemini API with an AI Studio key from the very beginning; it doesn't use Vertex at all—and 3.8 Live was released on the Gemini API.

I was warning my own repo about a risk that didn't exist in my repo. The project as a whole indeed "switched entirely to Vertex AI," but voice was the exception, and I was the one who wrote that exception.

Cause and Solution: Project-level decision records can become a memory shortcut. "We use Vertex for everything" is a correct summary, but summaries erase exceptions. Checking the source code takes thirty seconds and is much cheaper than reasoning based on impressions.


Pitfall 2: The Control Group Saved Me

After confirming the model existed, I wrote a throwaway script for testing. The key design was: Don't write a separate config; directly import the existing build_live_config() and build_voice_tools() from the project. I wanted to test "Can the configuration I'm currently running work?", not "Can a separate configuration I wrote work?".

The first run looked like this:

=== gemini-3.8-live ===
  [handsfree] {'connected': True, 'audio': True, ..., 'error': None}
  [PTT] {'connected': True, 'audio': False, ...,
               'error': 'APIError: 1007 None. Precondition check failed.'}
Enter fullscreen mode Exit fullscreen mode

The push-to-talk (PTT) mode was blocked. If I had only tested 3.8, I would have concluded: "3.8 Live doesn't support PTT; cannot upgrade."

But I also ran the old model:

=== gemini-3.1-flash-live-preview ===
  [PTT] {..., 'error': 'APIError: 1007 None. Precondition check failed.'}
Enter fullscreen mode Exit fullscreen mode

Exactly the same error. Since both models failed in the same way, it wasn't a model issue; it was a script issue.

The actual error was: In PTT mode, the production environment sends PCM audio, but I sent text for convenience. Inserting text input between activity_start and activity_end is an invalid combination.

After switching to real 16kHz PCM:

=== gemini-3.8-live ===
  [PTT] {'connected': True, 'audio': True, 'out_tx': True, 'in_tx': True,
         'resume': True, 'events': ['audio','in_tx','out_tx','resume','turn_complete']}
Enter fullscreen mode Exit fullscreen mode

Everything was there, and it matched the old model item by item.

Cause and Solution: The only thing I did right this time was keeping a control group. When testing something new, testing the old thing in the same way costs almost nothing, but it allows you to distinguish between "the new thing is broken" and "my test is written incorrectly." These two things look identical.


Pitfall 3: Sine Waves Aren't Human Voices

In the second version of the script, I used math.sin to generate a 180Hz PCM segment as audio input. PTT mode tested smoothly, but both models timed out in hands-free mode.

I almost wrote this off as "hands-free behavior pending confirmation." But a moment's thought revealed why: hands-free mode relies on Gemini's own Voice Activity Detection (VAD), and a pure sine wave isn't speech. The VAD correctly judged that "no one is talking," so it never triggered.

So it wasn't a "detected problem"; it was not detected at all. Two timeouts looked like data, but they were actually two empty spaces.

This later evolved into a checklist. I separated what the script tested from what it didn't:

Tested: Connection, config acceptance, PTT audio round-trip, bidirectional transcripts, resumption handle, turn_complete, tool calling trigger with correct parameters.

Not tested: Real human Chinese speech recognition, automatic VAD in hands-free mode, interrupting while speaking, reconnection after a ten-minute connection recycle, whether the resumption handle actually restores context, Google Search grounding actually taking effect, the second half of slow task delegation pushed back to LINE, voice quality, and latency.

Of the eight items, eight were not tested. The script ran beautifully, but it touched the protocol layer, not what a user would encounter.


Comparison Table

After going through the process, the final comparison looks like this. Both columns are results from hitting the real API using the project's existing config generator:

Validation Item 3.1-flash-live-preview 3.8-live
PTT (activity signal + PCM chunks) Audio / Bidirectional Transcripts / resume / turn_complete Same
session_resumption Yes Yes
context_window_compression Accepted Accepted
Bidirectional AudioTranscriptionConfig Yes Yes
Voice Aoede Yes Yes
google_search and function_declarations in separate Tools Accepted Accepted
Tool calling actually triggered Both tool parameters correct Same
api_version="v1beta" Required Still required

In other words, gemini-3.8-live can be swapped directly into the existing configuration.


Extended Thinking Has a Required Field

I also tested gemini-3.8-live-extended-thinking, and the connection was immediately blocked:

APIError: 1007 None. Thinking level must be specified for this model.
Enter fullscreen mode Exit fullscreen mode

After adding thinking_config, both LOW and HIGH could connect. So it's not broken; it just has an extra required parameter.

I decided not to adopt it because the current build_live_config() doesn't send thinking_config. If I just swapped the model name, the voice assistant would crash at the connection step. Using it would require modifying the config generator, which is a separate task.

But I couldn't "keep" the knowledge that it shouldn't be set, so I wrote it as a test:

# Live API (bidiGenerateContent) support list. Confirmed by connection test on 2026-09-16.
# Deliberately excludes gemini-3.8-live-extended-thinking: this model mandates thinking_level,
# and connections without thinking_config will be blocked (1007).
LIVE_CAPABLE = {"gemini-3.8-live"}
Enter fullscreen mode Exit fullscreen mode

Last time I wrote "Live API only supports a certain model" in a comment, that sentence expired for months without anyone noticing. This time, I wrote it as an assertion that will fail.


What Changed

The actual code changes were small, about eighty lines across seven files:

  • config/agent_config.py: Added VOICE_MODEL module constant and AgentConfig.voice_model. A detail here—I deliberately didn't just put it in get_agent_config(), because that function requires GOOGLE_CLOUD_PROJECT, and services/voice_live.py needs the model ID at import time. If I went that route, imports would fail if the project environment variable wasn't set.
  • services/voice_live.py: Removed the model literal entirely and pulled it from the config.
  • tests/test_model_config.py: Reduced EXEMPT_FILES from two files to one; included voice_model in the existing "no preview" and "must be in the verified capable list" guards.

The last guard originally looked like this:

def test_voice_live_model_unchanged():
    """Live API only supports gemini-3.1-flash-live-preview; changing it will break the voice assistant."""
    assert VOICE_MODEL == "gemini-3.1-flash-live-preview"
Enter fullscreen mode Exit fullscreen mode

It asserted a string. Strings expire, and when they do, the test is still green—it only gets in your way when you want to upgrade; it doesn't save you when a model is taken down.

By changing it to a list of capable models, it guards against the category of "this value must be a verified Live model," rather than a specific value.

The exception for tools/tts_tool.py remains. I checked, and all three TTS models on the API (2.5-flash-preview-tts, 2.5-pro-preview-tts, 3.1-flash-tts-preview) are still in preview, with no GA versions to swap to. The exceptions were reduced from two to one, but not zeroed out.


Deployment Order: Config First, Code Second

Environment variables were set before merging:

gcloud run services update linebot-helper-python --region us-central1 \
  --update-env-vars VOICE_MODEL=gemini-3.8-live
Enter fullscreen mode Exit fullscreen mode

This step had absolutely no effect at the time. The live version was still running the old image, which had the model hardcoded in voice_live.py and wouldn't even read this variable.

But the order was correct. Merging triggers Cloud Build for automatic deployment. As soon as the new image goes up, it reads the already-in-place configuration, ensuring there isn't a window where "the code is up but the config isn't."

This approach has a side effect that I find more valuable than the upgrade itself: Rollback doesn't require re-deployment.

gcloud run services update linebot-helper-python --region us-central1 \
  --update-env-vars VOICE_MODEL=gemini-3.1-flash-live-preview
Enter fullscreen mode Exit fullscreen mode

If real-world testing feels off, a single command switches it back without waiting for a build. Changing the model ID from hardcoded to an environment variable bought me this flexibility.


Then I Tested It for Real, and the Results Overturned My Observations

There was one thing in the script I always found suspicious. During the tool calling round, the old model would speak a filler phrase before waiting for the tool result, while 3.8 did not.

My judgment at the time was cautious: this might be the "tools execute in the background without interrupting the flow" mentioned in the announcement, or it might just be my loop closing early. Automation couldn't tell; it required a human talking to know if it was an improvement or an awkward silence.

After merging and deploying, I spoke a few sentences via LIFF. The answer: There is no difference from before; it's very smooth.

So that difference doesn't exist to a human ear. It was a byproduct of my loop returning as soon as it got turn_complete. I was observing the shape of my own script.

In this article, my script lied to me three times: the PTT incident made the new model look broken, the sine wave incident made "not tested" look like "tested," and this time made a non-existent difference look like a new capability. None of them were model issues.


296 Passing Tests Don't Prove Voice Works

This is the part I find most worth writing down.

After the upgrade, the tests went from 293 to 296, all green. But this number has almost nothing to do with "whether voice works on 3.8."

tests/test_voice_live.py has 32 tests covering PTT disabling auto-VAD, activity signals, transcript forwarding, tool calling execution, resumption handles, connection recycling, and interruptions. It looks comprehensive. But they all use FakeLiveSession and don't connect to the internet. If the model changes from 3.1 to 3.8, not a single one of these 32 tests will change color.

They test "Is what I'm sending correct?", not "What is the other side returning?".

The three layers of coverage look like this:

Layer What it tests Who runs it Repeatable
Repo test suite (296) Outgoing config and message format CI on every push Yes
Throwaway probe (3) Protocol layer compatibility with real API Ran only once No
Real human conversation Understanding, smoothness Me No

The middle layer is the only automated test that actually touched 3.8, and it's not in the repo, CI doesn't run it, and no one will know the next time a model is taken down.

The newly added guard only blocks "setting it to a known bad value"; it doesn't block "Google taking 3.8-live down"—which is what actually happened last time.

That loophole is still open. Turning the probe into a smoke test that requires a key is the solution; I haven't done it yet.


The code is at kkdai/linebot-helper-python, and this change is in PR #24. The official announcement is Gemini 3.8 Live, which mentions that all AI-generated audio carries a SynthID watermark—I did not verify this.

The announcement spent the most space on real-time visual input, which I didn't touch at all this time—currently, the LIFF getUserMedia is hardcoded to video: false. That will probably be the next post.

Top comments (0)