Skip to content

fix: match audio device, ASR language, and VAD timing to Jetson Thor setup - #16

Open
Aaron-VG wants to merge 1 commit into
mainfrom
fix/kinect-mic-english-vad-timing
Open

Aaron-VG wants to merge 1 commit into
mainfrom
fix/kinect-mic-english-vad-timing

Conversation

@Aaron-VG

Copy link
Copy Markdown

Summary

Local config tweaks made while integrating this package into coghri_robotless_stack on a Jetson Thor:

  • audio_stream_manager's device-name match was "jabra" - switched to "Kinect" to select the Azure Kinect Microphone Array present on this rig instead.
  • speech_recognition's language allowlist was "es,ca" - switched to "en" for English-speaking dev/test use.
  • min_silence_duration was 0.5s - a normal mid-sentence breath/pause could exceed that, splitting one spoken sentence into multiple ASR segments that each independently triggered downstream reactive logic. Bumped to 0.9s. (A debounce was also added on the consumer side in coghri_reactive_tasks_generator as a second line of defense.)

Test plan

  • Verified live on a Jetson Thor: correct mic selected, English transcription, and no more sentence fragmentation into multiple ASR segments under normal speech pacing.

…setup

audio_stream_manager was set to select the input device by matching
"jabra" in the device name - wrong for a Jetson Thor rig using an Azure
Kinect Microphone Array instead. Switched to matching "Kinect".

speech_recognition's language allowlist was "es,ca" - switched to "en"
for English-speaking test/dev use.

min_silence_duration was 0.5s, short enough that a normal mid-sentence
breath/pause could exceed it and split one spoken sentence into several
separate ASR segments, each independently triggering downstream reactive
logic. Bumped to 0.9s.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants