Skip to content

feat(aic): update ai-coustics integration to SDK 0.21 (VAD probability + Tyto) - #1

Closed
gkmngrgn wants to merge 2 commits into
mainfrom
goekmengoergen/sys-778-pipecat-update-ai-coustics-integration-to-sdk-021-aic-sdk
Closed

gkmngrgn wants to merge 2 commits into
mainfrom
goekmengoergen/sys-778-pipecat-update-ai-coustics-integration-to-sdk-021-aic-sdk

Conversation

@gkmngrgn

@gkmngrgn gkmngrgn commented Jun 24, 2026 •

Copy link
Copy Markdown

Summary

Updates the ai-coustics integration to the newly released SDK 0.21 (Python wrapper aic-sdk 2.5.0, native 0.21.0). Two independent commits.

Commit 1 — VAD probability + new Quail models

  • Bump the aic-sdk pin ~=2.3.0 → ~=2.5.0. The old pin resolved to >=2.3.0,<2.4.0, which blocked 2.4/2.5 — nobody on Pipecat could install the newest aic-sdk.
  • AICQuailVADAnalyzer.voice_confidence() now returns the model's continuous raw probability (VadContext.raw_vad_probability(), new in 0.21) instead of the binary is_speech_detected() flag, so Pipecat's own VADParams (confidence/start_secs/stop_secs) performs the speech gating — the intended VADAnalyzer contract.
  • Removed the now-redundant SDK-side sensitivity/speech_hold_duration/minimum_speech_duration params (they only affected the unused post-processed path; the analyzer is unreleased, so no deprecation needed).
  • Default VAD model → quail-vf-vad-2.0-s-16khz; example enhancement model → quail-vf-2.2-l-16khz.

Commit 2 — Tyto audio-quality analyzer

  • New AICTytoAnalyzer FrameProcessor backed by the ai-coustics Tyto model (aic-sdk 2.4.0+). Passively taps input audio, buffers into the SDK Collector, and runs Analyzer.analyze_buffered() off the event loop on a configurable interval.
  • Emits a new AICAudioQualityMetricsData — seven 0.0–1.0 scores (risk_score, speaker_reverb, speaker_loudness, interfering_speech, media_speech, noise, packet_loss) predicting downstream STT/VAD/turn-taking degradation — via a MetricsFrame and an on_audio_analysis event.
  • RTVIObserver forwards the scores to clients under an audio_quality metrics key.
  • Adds a voice-aicoustics-audio-quality example and unit tests.

Test plan

  • uv run ruff check / uv run ruff format --check — clean.
  • uv run pytest for the aic suites + RTVI + deprecation markers — 185 passed.
  • uv run towncrier build --draft — both changelog entries render.

Remaining before ready (why this is a draft)

  • Changelog fragment names: currently orphan +… files (render with an empty (PR )). Rename to <PR#>.changed.md / <PR#>.added.md once this PR number is known.
  • Tyto streaming smoke test with a real AIC_SDK_LICENSE: unit tests mock the SDK. Need to confirm on real hardware that Collector.buffer() accepts variable-size chunks under allow_variable_frames=True, and that the input sample rate is resampled to the model rate internally. Run examples/voice/voice-aicoustics-audio-quality.py.
  • Confirm whether Tyto warrants its own PR vs. shipping alongside the VAD work.

gkmngrgn added 2 commits June 24, 2026 00:25
Bump the aic-sdk pin to ~=2.5.0 (native SDK 0.21). The previous ~=2.3.0
pin resolved to >=2.3.0,<2.4.0, which blocked the 2.4/2.5 releases.

AICQuailVADAnalyzer.voice_confidence() now returns the model's continuous
raw probability (VadContext.raw_vad_probability(), new in 0.21) instead of
the binary is_speech_detected() flag, so Pipecat's own VADParams
(confidence/start_secs/stop_secs) performs the speech gating — the intended
VADAnalyzer contract. The SDK-side sensitivity / speech_hold_duration /
minimum_speech_duration params, which only affected the now-unused
post-processed path, are removed (the analyzer is unreleased).

Switch the default VAD model to the new quail-vf-vad-2.0-s-16khz, and bump
the example enhancement model to quail-vf-2.2-l-16khz.
Add AICTytoAnalyzer, a FrameProcessor backed by the ai-coustics Tyto
analysis model (aic-sdk 2.4.0+). It passively taps the pipeline's input
audio, buffers it into the SDK Collector, and runs Analyzer.analyze_buffered()
off the event loop on a configurable interval (the analysis is not
real-time safe). Each pass emits a new AICAudioQualityMetricsData — seven
0.0–1.0 scores (risk_score, speaker_reverb, speaker_loudness,
interfering_speech, media_speech, noise, packet_loss) that predict
downstream STT/VAD/turn-taking degradation — via a MetricsFrame and an
on_audio_analysis event.

RTVIObserver forwards the scores to clients under an audio_quality metrics
key. Adds a voice-aicoustics-audio-quality example and unit tests with
mocked Collector/Analyzer.
@gkmngrgn gkmngrgn closed this Jun 24, 2026
@gkmngrgn
gkmngrgn deleted the goekmengoergen/sys-778-pipecat-update-ai-coustics-integration-to-sdk-021-aic-sdk branch June 24, 2026 07:28
gkmngrgn pushed a commit that referenced this pull request Aug 6, 2026
gkmngrgn pushed a commit that referenced this pull request Sep 7, 2026
Address review on PR #1:

- Client mode no longer ignores --moq-bind. Pass it to moq.Client as the
  dial source address, so choosing a non-ephemeral local port is valid.
  The plumbing is now a general local bind, so rename serve_bind -> bind
  (the --moq-bind flag was already mode-neutral): serve mode keeps the
  concrete listen-address default, client mode defaults to ephemeral.

- Trim the example to a serve-mode default plus one client-mode line,
  dropping the public cdn.moq.dev/anon showcase (serve is fine as the
  example's default).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
steckes pushed a commit that referenced this pull request Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant