This document describes the complete testing infrastructure for the universal-agent-runtime project, designed to achieve and maintain 100% code coverage across all components.
The testing infrastructure provides:
- 100% Code Coverage: Comprehensive testing for both Rust backend and TypeScript frontend
- Real Service Integration: Tests run against actual PostgreSQL, SurrealDB, Redis, and LLM services
- Multi-Level Testing: Unit, integration, API, and end-to-end tests
- Local verification: Development and release checks run on operator-controlled local hosts
- Docker Integration: Consistent testing environments via Docker Compose
- Rich Reporting: HTML, JSON, XML, and LCOV coverage reports
-
Unit Tests
- Rust:
cargo test --lib --bins --tests - TypeScript:
bun test - Fast, isolated component testing
- Rust:
-
Integration Tests
- Rust: Tests in
tests/directory with_integrationsuffix - Database integration with real PostgreSQL and SurrealDB
- Redis caching and session management
- Single-threaded execution for data consistency
- Rust: Tests in
-
API Tests
- HTTP endpoint testing via
tests::apimodule - Request/response validation
- Authentication and authorization flows
- Error handling and edge cases
- HTTP endpoint testing via
-
End-to-End Tests
- Playwright automation testing full user workflows
- Cross-browser compatibility (Chromium focus)
- UI interaction validation
- Real application deployment testing
-
Performance Tests
- Load testing and benchmarking
- Response time validation
- Resource usage monitoring
- Regression detection
Central configuration for test execution parameters, service settings, and coverage thresholds.
Orchestrates test services with proper health checks:
- PostgreSQL with pgvector extension
- SurrealDB in memory mode
- Redis Stack for caching
- Unstructured API for document processing
- Main application with test environment
- Rust: local
cargo-llvm-cov; seedocs/coverage-baseline.mdfor the 60% threshold and per-file baseline - TypeScript: local
vitest --coverageusing the v8 provider infrontend/vitest.config.ts - Playwright: V8 coverage integration
- Local guard:
tools/coverage-drift.shreports per-file drift vs.docs/coverage-baseline.md;.grcovrcwas removed (unused, superseded bycargo-llvm-cov)
./tools/test-all.sh --quick- Smoke tests and unit tests only
- Fast feedback for development
./tools/test-all.sh --full- All test categories including E2E
- Complete coverage analysis
- Performance benchmarks
./tools/test-all.sh --ci- Non-parallel local execution for constrained hosts
- Comprehensive logging
- Strict coverage requirements
./tools/coverage.sh --unified./tools/coverage.sh --rust-only./tools/coverage.sh --typescript-only./tools/coverage.sh --unified --opendocker-compose -f docker-compose.test.yaml up --build
docker-compose -f docker-compose.test.yaml run test-runner# Start dependencies only
docker-compose -f docker-compose.test.yaml up -d postgres redis surreal
# Run tests against external services
./tools/test-all.sh --fullRun unit, integration, end-to-end, browser, conformance, regression, security, performance, load, stress, soak, and release-certification checks locally. Formatting, linting, typechecking, and coverage also run locally. GitHub Actions are reserved for deployment execution and deployment-specific validation; a build, package, image, or release step must not invoke a test suite directly or as a side effect.
The repository enforces this boundary with
pnpm github-actions-policy:validate and the pre-commit hook in lefthook.yml.
The tiered local commands in .claude/rules/rust.md and
.claude/rules/typescript.md remain authoritative for when each check runs.
# Rust toolchain with components
rustup component add rustfmt clippy llvm-tools-preview
# Coverage tools
cargo install cargo-llvm-cov
# Node.js and Bun
curl -fsSL https://bun.sh/install | bash
# Docker and Docker Compose
# Install via official documentation
# Playwright
npx playwright install --with-deps chromiumInstall the required tools on each supported local build/certification host. Release receipts record the exact source commit and tool outputs used.
# Database connections
DATABASE_URL=postgres://postgres:postgres@localhost:5431/uar_test
REDIS_URL=redis://localhost:6378
SURREAL_URL=ws://localhost:8001/rpc
# Test configuration
CONFIG_FILE=test-config.yaml
APP_ENVIRONMENT=test
TEST_MODE=true
COVERAGE=true
# Coverage instrumentation
CARGO_INCREMENTAL=0
RUSTFLAGS="-C instrument-coverage"
LLVM_PROFILE_FILE="tests/coverage/rust/coverage-%p-%m.profraw"OPENAI_API_KEY=your_api_key_here
TAVILY_API_KEY=your_api_key_heretests/
├── integration/ # Rust integration tests
│ ├── api/ # API endpoint tests
│ ├── database/ # Database integration tests
│ └── services/ # Service integration tests
├── e2e/ # End-to-end tests
│ ├── specs/ # Playwright test specifications
│ ├── pages/ # Page object models
│ └── utils/ # Test utilities
├── performance/ # Performance and benchmark tests
├── coverage/ # Coverage reports
│ ├── rust/ # Rust coverage data
│ ├── typescript/ # TypeScript coverage data
│ ├── e2e/ # E2E coverage data
│ └── unified/ # Combined coverage reports
├── config/ # Test-specific configuration
└── fixtures/ # Test data and fixtures
tools/
├── test-all.sh # Main test execution script
├── coverage.sh # Coverage report generator
└── ...
Primary Tool: cargo-llvm-cov
Alternative: cargo-tarpaulin
Formats: LCOV, HTML
Configuration: .cargo/config.toml
Primary Tool: Bun's built-in coverage
Alternative: c8, nyc
Formats: HTML, JSON, LCOV, text
Integration: Playwright V8 coverage for E2E
Tool: Custom HTML dashboard Features:
- Multi-format report aggregation
- Interactive coverage browsers
- Historical trend analysis
- Quality gate validation
- Direct links to detailed reports
- Tool:
cargo-mutants - Run locally:
cargo mutants --no-shuffle --features server-full - Reports:
docs/mutation-history/ - Summary:
bash tools/mutation-summarize.sh <report-dir>
- Tool:
cargo-fuzz - Targets:
fuzz/fuzz_targets/{chunker,rag_verification,mcp_message_parser,json_schema_validator}.rs - Run locally:
cargo +nightly fuzz run chunker(requires nightly Rust andcargo-fuzz)
- Tool:
proptest - Coverage: settings store serde roundtrip, retrieval RRF invariants, governance policy hot-reload semantics
- Run locally: included in
cargo test
- Hot-path unwrap/expect guard:
src/uar/api/,src/uar/runtime/, andsrc/server.rsdenyclippy::unwrap_used/clippy::expect_usedat the module level. Runcargo clippy --features server-full --no-depsto verify. - Central error type tests:
src/uar/error.rsverifiesUarErrorstable codes, JSON response shape, andtracing-errorspan capture. - Sentry feature build:
cargo check --no-default-features --features server-full,sentryensures the optional Sentry integration compiles. - Documentation: see
docs/observability.mdfor operator-facing tracing and Sentry setup guidance.
- Format Compliance:
cargo fmt,prettier - Linting:
clippy,eslint - Type Safety: TypeScript strict mode
- Security:
cargo audit,bun audit - Coverage Thresholds: Configurable per environment
- Performance: Response time and throughput validation
- Docker Health: Service connectivity and readiness
- Conventional Commits:
commitlint+lefthookfor the JS workspace - Mutation Testing:
cargo-mutantsnightly - Fuzz and Property Tests:
cargo-fuzzandproptest - Hot-Path Error Handling:
clippy::unwrap_used/clippy::expect_useddenied insrc/uar/api/,src/uar/runtime/, andsrc/server.rs
- Code Review: Required for all PRs
- Architecture Review: For significant changes
- Security Review: For authentication/authorization changes
- Performance Review: For database schema or API changes
# Check service health
docker-compose -f docker-compose.test.yaml ps
# View service logs
docker-compose -f docker-compose.test.yaml logs postgres
docker-compose -f docker-compose.test.yaml logs redis
docker-compose -f docker-compose.test.yaml logs surreal
# Reset containers
docker-compose -f docker-compose.test.yaml down -v
docker-compose -f docker-compose.test.yaml up -d --build# Verify tool installation
cargo llvm-cov --version
# Check profraw files generated
find . -name "*.profraw" -type f
# Clear coverage cache
rm -rf tests/coverage/
rm -f *.profraw# Run specific test category
cargo test --test config_integration -- --nocapture
cargo test tests::api -- --test-threads=1 --nocapture
# Enable debug logging
RUST_LOG=debug cargo test
# Playwright debugging
npx playwright test --debug
npx playwright test --headed --slow-mo=1000# Check resource usage
docker stats
# Analyze slow tests
cargo test -- --nocapture | grep -E "(test|elapsed)"
# Profile test execution
time ./tools/test-all.sh --quick- Isolation: Each test should be independent
- Determinism: Tests should produce consistent results
- Clarity: Test names should describe expected behavior
- Coverage: Test both happy paths and error conditions
- Performance: Keep unit tests fast (<10ms each)
- Completeness: Test all code paths including error handling
- Quality: Focus on meaningful tests, not just coverage percentage
- Maintenance: Keep tests updated with code changes
- Documentation: Use tests as behavioral documentation
- Fast feedback: Run the current tier's focused local checks while editing.
- Comprehensive validation: Run the phase-completion tier locally.
- Release quality: Retain source-bound local certification receipts.
- Monitoring: Track local test execution times and success rates.
- Test Execution Time: Track performance trends
- Coverage Percentage: Monitor quality improvements
- Failure Rate: Identify unstable tests
- Resource Usage: Optimize local certification-host costs
- Coverage Reports:
tests/coverage/unified/index.html - Test Results: source-bound local verification receipts
- Performance Trends: Artifact-based historical data
- Coverage Drop: Below threshold warnings
- Test Failures: local command failure and retained logs
- Performance Regression: Automated benchmark comparisons
- Choose appropriate test category (unit/integration/e2e)
- Follow naming conventions (
*_test.rs,*.test.ts,*.spec.ts) - Ensure tests are deterministic and isolated
- Update coverage thresholds if needed
- Document complex test scenarios
- Update configuration files as needed
- Test changes locally before committing
- Update documentation for new features
- Verify the local GitHub Actions policy guard still passes
- Consider backward compatibility
- Profile test execution to identify bottlenecks
- Optimize Docker image sizes and startup times
- Use caching effectively on local verification hosts
- Balance thoroughness with execution speed
- Monitor resource usage trends