Skip to content

gemini-3.8-live emits no VoiceActivity (documented replacement for deprecated speechState) #2981

Description

@r09528039

The Live API reference marks speechState as deprecated: "Use VoiceActivity instead". gemini-3.1-flash-live-preview emits voiceActivity (top-level, ACTIVITY_START / ACTIVITY_END with audioOffset) on both v1alpha and v1beta. gemini-3.8-live and gemini-3.8-live-extended-thinking emit nothing during user speech.

VoiceActivity is referenced by the deprecation note but not defined on the reference page nor listed in the BidiGenerateContentServerMessage union.

Is the absence of any server-side voice-activity signal on 3.8 intended?

Activity

  1. added
    type: questionRequest for information or clarification. Not an issue.
    priority: p3Desirable enhancement or fix. May not be included in next release.
    on Sep 18, 2026
  2. Venkaiahbabuneelam commented on Sep 18, 2026

    @Venkaiahbabuneelam

    Hi @r09528039,

    Thanks for the clear explanation! You are right that gemini-3.8-live models do not send VoiceActivity signals at this time.

    We are escalating this issue internally to investigate further. If you need voice activity detection right now, please use gemini-3.1-flash-live-preview as a workaround.

    We’ll keep you updated on progress. Thanks!

  3. r09528039 commented on Sep 29, 2026

    @r09528039
    Author

    Hi @Venkaiahbabuneelam ,

    It looks like gemini-3.8-live now sends voiceActivity (ACTIVITY_START / ACTIVITY_END), which seems like it could fix my issue. Has this been resolved on your end?

    Thanks!

  4. ajliouat commented on Oct 1, 2026

    @ajliouat

    Data point for the question above. On 29 and 30 September, with google-genai 2.25.0 (Python) against gemini-3.8-live on the Gemini API (AI Studio) and VAD at its default (no realtime_input_config), voiceActivity frames arrived in all 20 sessions with streamed 16 kHz PCM speech. ACTIVITY_START came 139 to 154 ms after the first audio chunk of an utterance while the model was silent, and up to 261 ms while it was speaking; ACTIVITY_END came about 1.2 to 1.3 s after the last speech chunk.

    On the wire the message is {"voiceActivity":{"type":"ACTIVITY_START","audioOffset":"0.160s"}}, and the SDK exposed it as voice_activity_type. One detail: a second utterance starting 0.3 s after the previous one got no ACTIVITY_START of its own; the server merged both into one.

    Raw JSONL timelines and the clips: https://github.com/frontier-on-cloud/gemini-live-stop-test (results/audio_*.jsonl, events raw_vad and voice_activity).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

priority: p3Desirable enhancement or fix. May not be included in next release.type: questionRequest for information or clarification. Not an issue.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions