Skip to content

Improve voice reliability and add recording, research, and local-agent workflows - #4

Open
coltonehrman wants to merge 1 commit into
wassgha:mainfrom
coltonehrman:codex/comprehensive-contribution
Open

coltonehrman wants to merge 1 commit into
wassgha:mainfrom
coltonehrman:codex/comprehensive-contribution

Conversation

@coltonehrman

@coltonehrman coltonehrman commented Sep 9, 2026 •

Copy link
Copy Markdown

Hi! First, thank you for building OpenDex—and apologies for arriving with such a large first contribution without discussing it with you beforehand. I started using it, then started building and fixing things to my heart's content, and the scope grew from there. I realize that leaves you with a substantial review, so I've tried to make the changes, checks, and limitations as clear as possible.

This brings together voice/desktop reliability fixes and several new workflows. I'm sharing it for feedback on both the implementation and whether these additions fit your direction for the project. There's no expectation that you accept the whole package as-is.

What changes

  • Voice reliability: microphone/wake handoff, confirmed spoken turns, response and tool cancellation, direct OpenAI realtime, and visible recovery for provider errors—including errors received before the renderer subscribes.
  • Desktop interaction: screen-access recovery, fresh screen descriptions and zoom, native macOS app/window controls, clearer compact transcripts and progress, and detachable desktop widgets.
  • New workflows: explicit interaction recordings, research plans with source-linked findings, persistent usage/cost estimates, tool-backed demos, source walkthroughs, and repeatable browser benchmarks.
  • Local development tools: opt-in diagnostics and Codex task discovery/actions/self-enhancement, plus permission-gated repository Git operations.
  • Permissions and validation: shared concurrent skill prompts, per-action Git confirmation, capability-aware tool availability, and a single regression-test command wired into CI.

How to review

Start with the contribution review guide. It maps every feature area to its implementation and tests, explains defaults and data storage, and includes a manual acceptance checklist. Detailed guides cover recordings, research, usage, diagnostics/local agents, and benchmarks.

The branch is based on v1.1.14 (3e89834), with one consolidated commit: 275 files, 15,221 additions, 776 deletions. The shared integration points are main IPC, preload, the voice hook, and the skill registry. Development journals and personal agent instructions have been removed from the submission; no local recordings, runtime histories, credentials, or generated builds are included.

Validation

On macOS arm64 with Node 25.6.1 and pnpm 10.8.1:

  • pnpm install --frozen-lockfile, pnpm typecheck, and pnpm build passed, including native compilation.
  • pnpm test: 307 passed, 0 failed, 0 skipped.
  • Both isolated Electron checks passed: node scripts/test-usage-desktop.mjs and node scripts/test-widgets-desktop.mjs.

CI now runs the regression suite alongside typecheck/build. Linux/Node 22 CI has not yet run on this contribution. The build reports mixed-import/chunk-size warnings and deprecated macOS application lookup. No paid-provider smoke test or final live microphone acceptance was run for this submission. Desktop fixtures validate their specific UI/IPC paths, not real provider billing or acoustic behavior.

Decisions and limits before merging

  • Transcript storage: diagnostics currently save rotating conversation transcripts locally by default, separately from the opt-in Diagnostics skill and anonymous analytics. There is no transcript opt-out/delete UI yet. Permissioned diagnosis can send that history to the chosen model. This needs an explicit product/privacy decision before release.
  • Platform support: the new native addon requires Xcode command-line tools on macOS. The Codex desktop adapter uses version-specific local IPC and rejects untested builds. Signed packages, cross-architecture builds, and Windows/Linux runtime behavior remain unverified.
  • Behavior and cost: confirming transcripts adds interruption latency; research enforces visible-link navigation and can use up to 96 steps. Usage totals include estimates and unknown charges, not invoice reconciliation or a spending limit.

Thank you for taking a look. I'd appreciate your direction on what belongs in OpenDex and what you'd like revised before this moves beyond draft.

Add recording, research, usage, local-agent, widget, and benchmark workflows with regression coverage and maintainer documentation.
@coltonehrman
coltonehrman marked this pull request as ready for review September 9, 2026 14:38

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant