Conversation
…setup audio_stream_manager was set to select the input device by matching "jabra" in the device name - wrong for a Jetson Thor rig using an Azure Kinect Microphone Array instead. Switched to matching "Kinect". speech_recognition's language allowlist was "es,ca" - switched to "en" for English-speaking test/dev use. min_silence_duration was 0.5s, short enough that a normal mid-sentence breath/pause could exceed it and split one spoken sentence into several separate ASR segments, each independently triggering downstream reactive logic. Bumped to 0.9s.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Local config tweaks made while integrating this package into
coghri_robotless_stackon a Jetson Thor:audio_stream_manager's device-name match was"jabra"- switched to"Kinect"to select the Azure Kinect Microphone Array present on this rig instead.speech_recognition's language allowlist was"es,ca"- switched to"en"for English-speaking dev/test use.min_silence_durationwas0.5s - a normal mid-sentence breath/pause could exceed that, splitting one spoken sentence into multiple ASR segments that each independently triggered downstream reactive logic. Bumped to0.9s. (A debounce was also added on the consumer side incoghri_reactive_tasks_generatoras a second line of defense.)Test plan