Skip to content

[firebase_ai] LiveSession.receive() ends on gemini-3.8-live's voiceActivity message (unhandled format) — every spoken turn kills the stream #18749

Description

@luiszapata6

Is there an existing issue for this?

  • I have searched the existing issues.

Which plugins are affected?

AI

Which platforms are affected?

iOS, Android

Description

gemini-3.8-live (Agent Platform, us-central1) sends a top-level server message when the user starts and stops speaking:

{"voiceActivity": {"type": "ACTIVITY_START", "audioOffset": "2.520s"}}
{"voiceActivity": {"type": "ACTIVITY_END",   "audioOffset": "3.640s"}}

firebase_ai 4.0.0 does not model it, so _parseServerMessage throws
FirebaseAISdkException('Unhandled format for LiveServerMessage: {voiceActivity: ...}').
_listenToWebSocket forwards that as an error on _messageController — fine so far, the socket stays open — but LiveSession.receive() is an async* over await for (final result in _messageController.stream), so the error ends the generator: the listener gets onError and then onDone, while the WebSocket and the controller underneath are still alive and the model goes on answering.

In practice every spoken turn on 3.8 ends the app's subscription at the first word (ACTIVITY_START), ~1 ms apart in our logs:

svc onError FirebaseAISdkException ... Unhandled format for LiveServerMessage: {voiceActivity: {type: ACTIVITY_START, audioOffset: 2.520s}}
svc onDone SOCKET CLOSED

An app that treats onDone as "session closed" (which is what it means on every other path) tears the session down or resumes onto a fresh one and loses the turn being spoken. Our workaround is to re-subscribe with session.receive() on the same session after that specific exception and drop the dead generator's done, which works because the controller is broadcast.

Two things would fix this properly:

  1. Parse the message (voiceActivity with type and audioOffset). It is the documented replacement for the deprecated speechState on Gemini 3.x Live models (see gemini-3.8-live emits no VoiceActivity (documented replacement for deprecated speechState) googleapis/python-genai#2981 and Emit input speech events on Gemini 3.x Live from voice_activity messages pydantic/pydantic-ai#9034) and is also the only server-side speech-start/end signal apps can use for barge-in UI.
  2. Make receive() resilient to a single unparseable frame — e.g. forward the error without ending the stream (_messageController.stream directly, or a handleError that keeps the generator alive) — so a future unknown message type degrades to one dropped frame instead of a dead session.

Reproducing the issue

  1. FirebaseAI.agentPlatform(location: 'us-central1').liveGenerativeModel(model: 'gemini-3.8-live', liveGenerationConfig: LiveGenerationConfig(responseModalities: [ResponseModalities.audio]))
  2. final session = await model.connect(); session.receive().listen(print, onError: print, onDone: () => print('done'));
  3. Stream microphone audio with sendAudioRealtime and say anything.
  4. On the first spoken syllable: FirebaseAISdkException: Unhandled format for LiveServerMessage: {voiceActivity: {type: ACTIVITY_START, ...}} followed immediately by done. Text input never triggers it (no VAD), and gemini-live-2.5-flash-native-audio never sends the message.

Firebase Core version

4.14.0

Flutter Version

3.47.2 (stable)

Relevant Log Output

13:36:41.446 onError FirebaseAISdkException: Unhandled format for LiveServerMessage: {voiceActivity: {type: ACTIVITY_START, audioOffset: 2.360s}} This indicates a problem with the Firebase AI Logic SDK...
13:36:41.447 onDone

Flutter dependencies

firebase_ai: 4.0.0
firebase_core: 4.14.0

Additional context and comments

Relevant SDK code: lib/src/live_session.dart (_listenToWebSocket, receive) and lib/src/live_api.dart (_parseServerMessage).

Activity

  1. luiszapata6 commented on Oct 5, 2026

    @luiszapata6
    Author

    Thanks @Seungpyo1007 — from reading #18752, yield* is exactly what we need: the parse error reaches onError and the stream keeps delivering, so the turn after a voiceActivity frame is no longer lost.

    For context, until this ships we work around it by re-listening to receive() whenever the error is the "Unhandled format for LiveServerMessage" one. That re-subscription can miss frames that arrive in the same burst as the unparsed message (e.g. a toolCall or turnComplete right after voiceActivity), which yield* avoids, so we'd drop our workaround once this is released.

    On parsing voiceActivity: we don't need it ourselves; keeping it in step with the other Firebase AI SDKs is fine from our side. The stream staying alive is what matters.

    We'll test the fix against gemini-3.8-live on device as soon as it is in a published firebase_ai version.

  2. added
    Needs AttentionThis issue needs maintainer attention.
    platform: androidIssues / PRs which are specifically for Android.
    platform: iosIssues / PRs which are specifically for iOS.
    plugin: ailabel issues for firebase_ai plugin
    type: bugSomething isn't working
    on Oct 6, 2026
  3. SelaseKay commented on Oct 8, 2026

    @SelaseKay
    Contributor

    Hi @luiszapata6, thanks for the report. I'm looking into this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Needs AttentionThis issue needs maintainer attention.platform: androidIssues / PRs which are specifically for Android.platform: iosIssues / PRs which are specifically for iOS.plugin: ailabel issues for firebase_ai plugintype: bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions