Finds and validates Brazilian sensitive data and leaked secrets — offline, with check digits, before they reach your logs, your CI, or your AI agent's context.
M4 done — the CLI, CI and the agent path all work. vedei scan, vedei git and vedei diff report to table, JSON, JSONL or SARIF; the GitHub Action uploads to Code Scanning; vedei hook keeps secrets out of a coding agent's context at 8.4 ms p95; vedei transcript scan and scrub handle what already leaked. Multilingual false-positive calibration is M5, and it is the thing most likely to annoy you before then — see the note under the roadmap.
go install github.com/mkmuniz/vedei/cmd/vedei@latest
go install github.com/mkmuniz/vedei/cmd/vedei-hook@latestvedei scan . # a tree, honouring .gitignore
vedei scan . --format sarif -o vedei.sarif
vedei git # the history, including deleted files
vedei diff --staged # what is about to be committed
vedei transcript scan # what already leaked into your session logs
echo 'cpf 529.982.247-25' | vedei stream
# cpf ***.***.***-25 [vedei: cpf redacted] (exit 3)Exit codes are distinct on purpose: 0 clean, 3 findings, 1 the scan itself failed, with 1 outranking 3. A pipeline that cannot tell a leak from a broken scan will read a permission error as a clean run.
In CI:
- uses: mkmuniz/vedei@v1
with:
mode: diff # only what changed; "history" on a scheduleWire the hook into your agent: hooks/. Two binaries, because the split is what makes the hook fast — vedei carries 417 secret rules, vedei-hook is 5 MB and only talks to the daemon.
.vedeiignore takes a finding fingerprint, derived from the type and the normalized value and never from the location — so an entry survives the file being moved or renamed. This repository's own file is .vedeiignore: one fingerprint per deliberate test value, each with the reason written next to it. Paths are almost never silenced: ignoring *_test.go wholesale would also silence a credential genuinely committed into a test, which is a place credentials genuinely end up. The single exception is corpus/cases/, the labeled corpus — every line there is a labeled example by construction, and a fingerprint per case would not scale.
| Data | Offline validation | Algorithm |
|---|---|---|
| CPF | ✅ | Mod 11, two check digits; repeated-sequence rejection |
| CIN | ✅ | Uses the CPF as the national number — same validation |
| CNPJ numeric | ✅ | Mod 11, weights 5,4,3,2,9,8,7,6,5,4,3,2 |
| CNPJ alphanumeric | ✅ | Mod 11 with ASCII − 48 for letters — first issued on 2026-07-31 (Receita Federal) |
| CNH | ✅ | Mod 11, 11-digit variant |
| PIS / NIS / NIT | ✅ | Mod 11, weights 3,2,9,8,7,6,5,4,3,2 |
| Título de eleitor | ✅ | Two check digits, mod 11, embedded state code |
| CNS (SUS card) | ✅ | Weighted sum ≡ 0 (mod 11) |
| Card PAN | ✅ | Luhn + BIN range |
| Pix key | ✅ | Per type: CPF/CNPJ, email, phone, EVP (UUID v4) |
| Pix E2EID | ✅ | E + ISPB(8) + YYYYMMDDHHMM + 11 alphanumerics |
| Branch / account | Check digit varies per institution | |
| RG | ❌ | No national standard. Contextual heuristic only, low confidence |
| Secrets (463 rules) | via provider | Delegated to betterleaks |
Most detectors match CNPJ with \d{14}. That has been wrong since July 2026.
ISPB registry from guibranco/BancosBrasileiros — 400+ institutions, updated daily.
| Surface | What it does | Status |
|---|---|---|
| AI agent context | PostToolUse hook for Claude Code and Codex. cat .env reaches the model redacted |
🎯 MVP |
| Agent transcripts | Audits what already leaked into ~/.claude/projects/*.jsonl |
🎯 MVP |
| CI/CD | CLI, GitHub Action, SARIF, pre-commit, pre-receive | planned |
| Application runtime | slog.Handler and HTTP middleware that redact before writing |
planned |
| Observability | OpenTelemetry collector processor for logs and traces | planned |
Secret scanning is solved and crowded — betterleaks ships 463 rules, kingfisher 1,051. Rebuilding that corpus loses. So vedei imports betterleaks (MIT) and spends its effort where global tools are weakest: Brazilian sensitive data, validated by check digit rather than matched by shape — CPF, CNPJ including the alphanumeric format, CNH, Pix keys and E2EIDs, card PANs with Brazilian brands. And it enforces at the boundary where data leaves: the agent's context, the CI run, the log line. One Go binary, no runtime, no network.
The second reason is architectural. betterleaks validates a secret by asking the provider over HTTP, which needs a rate limiter, a timeout budget, and care about where those requests go. Brazilian personal data doesn't work that way: a CPF is a mod-11 check digit, a card is Luhn, a Pix key is its type. All offline. No network, no rate limits, no SSRF surface. Simpler and safer, not just different.
A .env has protections — gitignore, file permissions, a vault. ~/.claude/projects/*.jsonl has none: plaintext, no encryption, no rotation, no expiry, no security tool watching it. Whoever reads the disk — malware, a stolen laptop, an exfiltrated backup — reads every secret the agent ever touched.
vedei reduces what lands there and shows what already did. It does not replace disk encryption.
Documented upstream: #44868, #50014, #95680.
Which agents' transcripts are read. vedei transcript scan and scrub cover Claude Code (~/.claude/projects/) and Codex (~/.codex/sessions/). Two other agents keep the same kind of record and are not read yet:
| agent | where it keeps sessions | why not yet |
|---|---|---|
| Gemini CLI | ~/.gemini/tmp/<project hash>/chats/, JSON files |
the format is documented but untested against real sessions; reading it blind would report coverage it does not have |
| Cursor | state.vscdb, a SQLite database under the Cursor user directory |
reading it needs a SQLite driver, and scrub cannot safely rewrite a database |
Both are plaintext on disk too. Until they are supported, treat their directories as holding whatever their agents read.
| Secret | Personal data | |
|---|---|---|
| Answers | Is it live? | Is it structurally real? |
| Mechanism | HTTP to the provider | Local arithmetic |
| Network | required | none |
| Latency | 50–500 ms | microseconds |
| Next step | revoke / rotate | redact / remove |
The line vedei will not cross: personal data is validated for structure, never by querying an official registry. vedei says 123.456.789-09 is a structurally valid CPF. Never whose.
- A detected value never reaches a language model. Enforced by a blocking CI test.
- Personal data is validated structurally, never against a registry.
- Model output is a signal, never a filter. It may reorder a queue; it may not discard a finding.
- Fail open on redaction, fail closed on reporting. Breaking a session is worse than missing a redaction; reporting "clean" on a broken scan is worse than a red build.
- Distinct exit codes:
0clean,3findings,1error. - No false positives on structurally invalid values. If the check digit fails, it is not a CPF.
- Thresholds calibrated per language. English-tuned heuristics misbehave on Portuguese source.
Full records in docs/adr/.
| Milestone | Status | |
|---|---|---|
| M0 | Foundation | ✅ |
| M1 | Brazilian detectors + offline validation | ✅ |
| M2 | Secret engine via betterleaks | ✅ |
| M3 | MVP — AI agent surface | ✅ |
| M4 | CLI, CI/CD, SARIF, transcript scrub | ✅ |
| M5 | Multilingual false-positive corpus + pt-BR calibration | 🚧 |
| M6 | Runtime SDK | ⬜ |
| M7 | AI layer | ⬜ |
| M8 | Observability collector | ⬜ |
| M9 | Secret validation + destination allowlist | ⬜ |
| M10 | v1.0 | ⬜ |
Features, implementation steps, test strategy and exit criteria per stage: MILESTONES.md. Component design and performance budgets: ARCHITECTURE.md.
On M5. The original hypothesis was a precision bias: betterleaks discards a candidate as prose when it contains a word from a 33,775-word English dictionary, so defaultpassword is dropped and senhapadrao falls through to a ratio threshold. Measured end to end on v1.8.1, the full filter chain rejects prose placeholders equally in English, Portuguese and Spanish. The real bias was in recall: the rules key on English names, so a real credential was found 100% of the time under an English key, 57% under a Spanish one and 29% under a Portuguese one — senha and segredo were never recognized. vedei now translates key names before a second pass, reaching 100% in all three. The labeled corpus behind these numbers is in corpus/: 119 cases, precision and recall 1.000 per type, enforced in CI.
detect/br/ # the core: pure validators, zero dependencies
detect/engine.go # the Engine interface — insulation from betterleaks' v2 API
detect/aws/ # AWS access key ids — the one secret rule vedei owns
engine/secrets/ # betterleaks wrapped behind that interface, plus detect/aws
redact/ # format-preserving redaction
report/ # table, JSON, JSONL, SARIF
scan/ # directory walker, gitignore matcher, git history
stream/ # stdin -> stdout, the MVP path
daemon/ # Unix socket server and client — the 8.4 ms path
transcript/ # agent session log reader
hooks/ # agent hooks, git hooks, launchd and systemd units
corpus/ # labeled precision/recall harness (make corpus)
Vedei is Portuguese: the first person past tense of vedar — to seal, to block, to keep from passing. Eu vedei: "I sealed it."
It is what the person running the tool says once it has done its job. The CPF did not reach the model, the key did not reach the log, the card number did not reach the commit. Sealed.
- New detector? Read
CONTRIBUTING.md— property-based tests are mandatory for check-digit code. - Found a vulnerability?
SECURITY.md. Do not open a public issue. - MIT, same as betterleaks.