Skip to content

[WIP] Add intelligent voice interruption capability to Jeeva robot - #3

Merged
SathiyaSenpai merged 1 commit into
mainfrom
copilot/add-intelligent-voice-interruption
Mar 9, 2026
Merged

SathiyaSenpai merged 1 commit into
mainfrom
copilot/add-intelligent-voice-interruption

Conversation

Copilot AI commented Mar 9, 2026

Copy link
Copy Markdown
Contributor

Thanks for asking me to work on this. I will get started on it and keep this PR's description up to date as I form a plan and make progress.

Original prompt

Problem

The Jeeva companion robot currently lacks two critical capabilities that distinguish it from a dumb home speaker:

1. Intelligent Voice Interruption (Barge-In)

When the robot is speaking and a user interrupts, the robot should:

  • Immediately detect the interruption (voice activity during TTS playback)
  • Gracefully stop speaking mid-sentence — not abruptly cut audio, but fade out naturally
  • Listen to what the user is saying and respond to the interruption contextually
  • Remember what it was saying so it can offer to continue if the interruption was brief (e.g., "sorry, go on")
  • Prioritize the user's speech — the robot serves the human, not the other way around

This requires:

  • Voice Activity Detection (VAD) running on ESP32 concurrently with audio playback
  • A barge-in protocol between ESP32 (detects interruption) and Jetson (controls TTS/LLM pipeline)
  • State machine: SPEAKING → INTERRUPTED → LISTENING → THINKING → SPEAKING
  • The INMP441 I2S mic on ESP32 must be actively monitoring even while the MAX98357A speaker is playing audio

2. Multi-Speaker Voice Differentiation

The robot should know WHO is talking:

  • Speaker diarization — distinguish between different voices (e.g., elderly user vs. nurse vs. family visitor)
  • Speaker identification — associate a voice with a registered user profile (name, health data, medication schedule, language preference)
  • Behavioral adaptation — respond differently based on who is speaking:
    • Primary user (elderly resident): warm, patient, health-aware responses
    • Caregiver/nurse: clinical, concise, report-oriented responses
    • Unknown voice: polite but guarded, no health data disclosure without authorization
  • Voice enrollment — simple process for registering new voices ("Jeeva, learn my voice" → 10-second voice sample → speaker embedding stored)
  • When multiple people are talking simultaneously, the robot should:
    • Focus on the closest/loudest speaker (or the registered primary user)
    • Politely indicate it can only listen to one person at a time
    • Not get confused and hallucinate mixed transcriptions

3. Truly Intelligent Companion Behavior

The robot must be intelligent in EVERY aspect — not just answering questions:

Proactive Intelligence:

  • Detect user mood from voice prosody (pitch, speed, energy, pauses) and adapt behavior
  • Notice patterns: "You seem quieter than usual today, is everything alright?"
  • Track conversation topics over days/weeks and bring them up naturally: "How did your grandson's exam go? You mentioned it last Tuesday"
  • Anticipate needs based on time + context: morning greeting, medication time awareness, bedtime routine
  • Environmental awareness: if room is very quiet for too long + no movement detected, gentle check-in

Conversational Intelligence:

  • Maintain conversation context across multiple turns (not just current session — across days/weeks via SQLite memory)
  • Handle topic switches gracefully without losing previous context
  • Know when to be silent (companionship ≠ constant talking)
  • Detect and respond to emotional distress, loneliness cues, confusion
  • Code-switch naturally between Tamil/Hindi/English mid-conversation (matching the user's language switching pattern)

Social Intelligence:

  • Remember birthdays, anniversaries, important dates mentioned in conversation
  • Understand family relationships ("your daughter Priya who lives in Bangalore")
  • Adapt formality level to the speaker (respectful Tamil with elders, casual with children)
  • Recognize when the user is talking to someone else (not addressing the robot) and stay quiet

Implementation Scope

This PR should add the following to the repository:

ESP32 Firmware Changes (esp32/src/)

  1. voice_activity_detector.h — Real-time VAD on INMP441 I2S stream using energy + zero-crossing-rate, runs concurrently with audio playback. Sends BARGE_IN interrupt to Jetson via UART when speech detected during robot speaking state.
  2. speaker_tracker.h — Extracts basic voice features (fundamental frequency F0, spectral centroid, energy envelope) from audio frames on ESP32 for quick speaker-change detection. Sends raw audio chunks to Jetson for full speaker embedding computation.
  3. interrupt_state_machine.h — State machine managing: IDLE → WAKE_WORD_DETECTED → LISTENING → PROCESSING → SPEAKING → INTERRUPTED → LISTENING. Clean transitions with timeout handling.
  4. Update main.cpp — Integrate VAD, speaker tracking, and interrupt state machine into the main FreeRTOS loop with proper task priorities.
  5. Update config.h — Add configuration constants for VAD thresholds, barge-in sensitivity, speaker tracking parameters.

Jetson Python Services (new jetson/ directory)

  1. jetson/speaker_recognition.py — Speaker embedding extraction using resemblyzer or speechbrain (runs on Jetson GPU). Speaker enrollment, identification, and diarization for up to 20 registered ...

This pull request was created from Copilot chat.


✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.

@SathiyaSenpai
SathiyaSenpai marked this pull request as ready for review March 9, 2026 11:23
Copilot AI review requested due to automatic review settings March 9, 2026 11:23
@SathiyaSenpai
SathiyaSenpai merged commit 769bdff into main Mar 9, 2026
1 check failed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot wasn't able to review any files in this pull request.


💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

You can also share your feedback on Copilot code review. Take the survey.

Copilot stopped work on behalf of SathiyaSenpai due to an error March 9, 2026 11:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants