Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@
"name": "sensegrep",
"source": "./plugin/sensegrep-plugin",
"description": "Semantic code search for AI agents. Search code by meaning, not text patterns. Adds sensegrep MCP tools + smart usage instructions to Claude Code.",
"version": "1.16.0",
"version": "1.17.0",
"author": {
"name": "sensegrep"
},
Expand Down
2 changes: 1 addition & 1 deletion .cursor-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@
"name": "sensegrep",
"source": "plugin/sensegrep-cursor",
"description": "Semantic code search for AI agents. Search code by meaning, not text patterns.",
"version": "1.16.0"
"version": "1.17.0"
}
]
}
3 changes: 2 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
node_modules/
bun.lock
bun.lock
.sensegrep-build.lock

# build outputs
dist/
Expand Down
36 changes: 33 additions & 3 deletions docs/cli-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -347,6 +347,36 @@ sensegrep selftest --strict --json

Hybrid retrieval defaults to parallel lexical and semantic collection. Search investigates at least 200 candidates before selecting the requested number of results; broad queries may take longer than adaptive mode. Use --hybrid-mode adaptive to explicitly opt into lexical skipping.

Duplicate scans first detect exact/normalized copies. Small cosine-vector candidate sets (up to 512) are compared in memory. Continuations preserve previous pairs in temporary snapshot-bound checkpoints; changed snapshots/candidate sets require restarting without --resume-cursor. A max-candidates cap still means incomplete coverage even when that subset finishes.

Context accepts --max-output-bytes in addition to --max-tokens. Compact, relevant complete symbols are preferred over truncating an oversized first result; truncation remains explicit in JSON.
Duplicate scans first detect exact/normalized copies. Cosine-vector candidate sets up to 4096 are compared in memory, with precomputed norms and cached symmetric distances; larger sets and other distance metrics use the vector store. Diagnostic output reports load, preparation, and neighbor-scan timings plus the selected strategy. Continuations preserve previous pairs in temporary snapshot-bound checkpoints; changed snapshots/candidate sets require restarting without --resume-cursor. A max-candidates cap still means incomplete coverage even when that subset finishes.

Search and context accept `--max-output-bytes` in addition to `--max-tokens`.
For search, the byte limit must be at least 256. JSON byte limits include the envelope,
UTF-8 encoding, pretty printing, and the trailing newline. If evidence does not fit,
`status: "incomplete"` and explicit truncation are returned. Token counts remain estimates.

Token-budget selection considers the candidate pool before applying the result limit.
It balances relevance, new query-term coverage, source roles, and references from strong
candidates. Tests and type contracts receive less context space unless requested;
`--purpose test` explicitly favors test evidence. These are retrieval heuristics, not
proof that all relevant behavior has been found. Whole symbols are preferred, with an
explicit partial snippet only when no complete candidate fits.

Search/context output reports `answerSufficiency: "not-assessed"`. Diagnostic cards use
`rankingStrength` for the high/medium/low ranking heuristic. Full internal results retain
`confidence` as a deprecated compatibility alias; neither value is a calibrated answer
probability. Execution completion is distinct from answer sufficiency.

Graph references include source/target locations, resolution method, and call line when
available. Calls inside nested indexed symbols belong to the smallest containing symbol.
`graphCoverage` is the fraction of resolved extracted edges (including synthetic imports
and table references), not recall against all real references. Use compiler-aware tools
for exhaustive refactoring. Dynamic calls and unsupported alias/re-export forms may remain
unresolved.

Clusters require similarity to every member to prevent transitive similarity chains.
Titles prefer symbol/file terms over common testing/framework imports. Colliding cluster
titles include a source location for disambiguation.

Chunking policy version 6 preserves structural metadata even below the old minimum file
size. `sensegrep index --no-watch` detects the old policy and rebuilds the index atomically;
existing indexes remain readable until that explicit indexing command runs.
71 changes: 71 additions & 0 deletions docs/evaluation-fixes-2026-09-24.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
# Correções da avaliação de 24/09/2026

Implementação local na branch `works/fix-evaluation-regressions`, baseada na versão
1.16.0. Estes resultados se referem à CLI compilada do repositório, não ao pacote
global publicado.

## Mudanças

- Arquivos pequenos passam pelo parser e preservam linguagem, símbolos e chamadas.
O fallback também conserva linguagem e trechos curtos. Política de chunking 6:
a próxima indexação explícita reconhece a necessidade de reconstrução.
- O contexto aplica o orçamento antes de limitar o número final de resultados.
A seleção considera relevância, cobertura adicional de termos, papel do arquivo
e referências textuais em candidatos relevantes. O ganho por tamanho é limitado.
Testes e contratos continuam disponíveis e podem ser priorizados explicitamente.
- `rankingStrength` substitui a apresentação ambígua de `confidence` nos cards
diagnósticos; a saída interna completa preserva o alias para compatibilidade.
`answerSufficiency: "not-assessed"` separa conclusão da execução de suficiência.
- O grafo atribui chamadas ao menor símbolo indexado que as contém. Referências
expõem origem, destino, método de resolução e linha da chamada quando disponível.
O cache foi versionado e o denominador da cobertura está documentado.
- Duplicatas com até 4096 candidatos e distância cosseno usam vetores normalizados
e distâncias simétricas em memória. Há verificação de prazo, cancelamento,
continuação e tempos separados de leitura, preparação e busca de vizinhos.
- Clusters exigem compatibilidade com todos os membros, evitando agrupamentos
gigantes formados por uma cadeia de semelhanças. Nomes preferem símbolos/arquivos;
imports comuns de frameworks e testes têm menos influência.
- `search --max-output-bytes` limita o JSON completo, incluindo formatação e UTF-8.
O mínimo aceito é 256 bytes; limites inválidos retornam JSON de erro e exit code 2.

## Validação

- Build e type-check dos workspaces passaram; versões dos pacotes consistentes.
- Suíte completa: 44 arquivos, 232 testes aprovados. Depois, duas regressões de
seleção por papel/referência foram adicionadas; a execução direcionada passou
com 25 testes. Total de testes na árvore final: 234.
- Fixture real indexada com Ollama, incluindo arquivos menores que 200 caracteres:
filtros Python, Java e Vue retornaram resultados; o grafo encontrou
`purchase -> authorizeDebit -> readBalance` via alias importado e distinguiu
chamada direta de chamada agendada.
- Contexto de concorrência de voz com 1200 tokens passou a incluir
`markContextConsumed`, ausente na avaliação anterior. Contexto de autenticação
manteve os helpers de refresh de sessão. Ambos indicam orçamento incompleto.
- Busca de duplicatas com os mesmos parâmetros da avaliação anterior:
1628 candidatos, limite 1500, prazo 5 segundos. Antes: 44 processados; depois:
873 processados (aproximadamente 20 vezes mais nesta medição).
- Continuação: 627 candidatos restantes processados, sem timeout. Os 24 grupos
acumulados foram iguais aos 24 da execução única de 1500 candidatos.
O resultado continua incompleto devido aos 128 candidatos excluídos pelo limite.
- Escopo pequeno de duplicatas: 219/219 candidatos e os mesmos quatro grupos;
1,524 s de processo, contra 1,511 s na avaliação anterior. O ganho é no escopo amplo.
- Consulta ampla de pagamentos: maior cluster retornado caiu de 52 para 18 membros;
os títulos distinguem processamento de webhook, reconciliação e mutações de status.
- JSON minimal/content/diagnostic/full com pretty printing respeitou 1800 bytes;
content respeitou também 1500 bytes. Quando os diagnósticos não cabem, a saída
degrada para um envelope incompleto e válido.
- A consulta negativa sobre Kubernetes continua podendo retornar um vizinho
semântico; a saída agora declara explicitamente que suficiência não foi avaliada.

## Limites e aplicação

Os tempos são observações locais, não um benchmark controlado de hardware.
O ganho de duplicatas vem do algoritmo de comparação, não de uso adicional da GPU.
Ranking e contexto continuam heurísticos; não garantem encontrar toda regra de negócio.
O grafo não substitui resolução pelo compilador para chamadas dinâmicas, re-exports
e outras formas que não resolve. A contagem de tokens de saída continua estimada.

O pacote npm, a CLI global e a skill instalada permanecem na versão publicada.
Depois de disponibilizar esta implementação, execute `sensegrep index --no-watch`
nos projetos para reconstruir índices com a política antiga. Não houve reconstrução
do índice principal durante estes testes; apenas a fixture temporária foi indexada.
12 changes: 6 additions & 6 deletions package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

11 changes: 11 additions & 0 deletions packages/cli/CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,16 @@
# @sensegrep/cli

## 1.17.0

### Minor Changes

- [`bde06a7`](https://github.com/Stahldavid/sensegrep/commit/bde06a75e8327c4a5c0a4c90db9b03f06c00a3b6) Thanks [@Stahldavid](https://github.com/Stahldavid)! - Preserve structural metadata in small files, improve token-budget context selection and cluster coherence, expose graph resolution evidence, accelerate bounded duplicate scans, and support serialized byte budgets for search. Distinguish ranking strength from answer sufficiency. Chunking policy changes trigger an atomic rebuild on the next explicit indexing command.

### Patch Changes

- Updated dependencies [[`bde06a7`](https://github.com/Stahldavid/sensegrep/commit/bde06a75e8327c4a5c0a4c90db9b03f06c00a3b6)]:
- @sensegrep/core@1.17.0

## 1.16.0

### Minor Changes
Expand Down
4 changes: 3 additions & 1 deletion packages/cli/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,9 @@ sensegrep semantic-kinds --json

`--json` writes parseable JSON to stdout; progress and warnings are written to stderr.
Argument errors under `--json` also use stdout JSON and exit code 2. `--max-output-bytes`
is enforced against the complete serialized payload.
is supported by search/context/audit/literal and enforced against the complete serialized JSON payload. Search requires at least 256 bytes.
Search/context explicitly report `answerSufficiency: "not-assessed"`; diagnostic
`rankingStrength` describes retrieval ranking, not proof that the question is answered.
Hybrid retrieval runs semantic and lexical work concurrently and uses one batched index read
for lexical matches. `--hybrid-mode parallel` is the default; `--no-hybrid` is available for
latency-sensitive semantic-only discovery. Deterministic query embeddings are cached locally
Expand Down
4 changes: 2 additions & 2 deletions packages/cli/package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@sensegrep/cli",
"version": "1.16.0",
"version": "1.17.0",
"type": "module",
"bin": {
"sensegrep": "dist/main.js"
Expand All @@ -25,7 +25,7 @@
"node": ">=20"
},
"dependencies": {
"@sensegrep/core": "^1.16.0"
"@sensegrep/core": "^1.17.0"
},
"publishConfig": {
"access": "public"
Expand Down
1 change: 1 addition & 0 deletions packages/cli/src/args.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ describe("CLI arguments", () => {
})

it("accepts embedding timeout and global audit budgets", () => {
expect(validateKnownFlags("search", { "max-output-bytes": "4800" })).toBeUndefined()
expect(validateKnownFlags("context", { "max-output-bytes": "4800" })).toBeUndefined()
expect(validateKnownFlags("search", { "embedding-timeout": "1000" })).toBeUndefined()
expect(validateKnownFlags("audit", {
Expand Down
2 changes: 1 addition & 1 deletion packages/cli/src/args.ts
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ const ALLOWED_FLAGS_BY_COMMAND: Record<string, Set<string>> = {
]),
verify: new Set([...GLOBAL_FLAGS, "strict"]),
status: new Set([...GLOBAL_FLAGS, "verbose", "verify"]),
search: new Set([...GLOBAL_FLAGS, ...EMBEDDING_FLAGS, ...INDEX_RUN_FLAGS, ...SEARCH_FILTER_FLAGS]),
search: new Set([...GLOBAL_FLAGS, ...EMBEDDING_FLAGS, ...INDEX_RUN_FLAGS, ...SEARCH_FILTER_FLAGS, "max-output-bytes", "maxOutputBytes"]),
literal: new Set([...GLOBAL_FLAGS, "query", "include", "exclude", "limit", "regex", "ignore-case", "ignoreCase", "filesystem", "max-output-bytes", "maxOutputBytes", "include-rendered-output", "dry-run"]),
context: new Set([...GLOBAL_FLAGS, ...EMBEDDING_FLAGS, ...INDEX_RUN_FLAGS, ...SEARCH_FILTER_FLAGS, "require-coverage", "requireCoverage", "max-output-bytes", "maxOutputBytes"]),
audit: new Set([
Expand Down
4 changes: 4 additions & 0 deletions packages/cli/src/search-commands.ts
Original file line number Diff line number Diff line change
Expand Up @@ -143,6 +143,7 @@ export function buildCommonSearchParams(query: string, flags: Flags, defaults: O
if (flags["no-shake"] !== undefined) params.shake = false
assignNumberParam(params, flags, "minScore", ["min-score", "minScore"])
assignNumberParam(params, flags, "maxTokens", ["max-tokens", "maxTokens"])
assignNumberParam(params, flags, "maxOutputBytes", ["max-output-bytes", "maxOutputBytes"])
if (flags.hybrid !== undefined) params.hybrid = toBool(flags.hybrid) ?? true
if (flags["no-hybrid"] !== undefined) params.hybrid = false
assignStringParam(params, flags, "hybridMode", ["hybrid-mode", "hybridMode"])
Expand Down Expand Up @@ -193,6 +194,9 @@ export async function executeSearchLikeTool(input: {
includeRendered,
includeFilterExplanations: input.params.explainFilters === true,
})
if (typeof input.params.maxOutputBytes === "number") {
payload.budget = { ...payload.budget, maxBytes: input.params.maxOutputBytes }
}
const finalPayload = enforceActualOutputBudget(payload)
if (input.params.requireCoverage === true && finalPayload.coverageSatisfied === false) process.exitCode = 2
writeJson(finalPayload)
Expand Down
2 changes: 1 addition & 1 deletion packages/cli/src/usage.ts
Original file line number Diff line number Diff line change
Expand Up @@ -77,7 +77,7 @@ Search options:
--continue-uncovered Add token-bounded batches until changed-file textual coverage is complete
--batch-tokens <n> Per-batch audit budget (default: 4000)
--max-total-tokens <n> Global audit token budget, including continuation batches
--max-output-bytes <n> Global serialized audit evidence budget
--max-output-bytes <n> Serialized JSON budget for search/context/audit (search minimum: 256)
--max-batches <n> Maximum number of continuation batches
--profile <name> Select a side-by-side named index profile
--embed-model <name> Override remote embedding model
Expand Down
6 changes: 6 additions & 0 deletions packages/core/CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# @sensegrep/core

## 1.17.0

### Minor Changes

- [`bde06a7`](https://github.com/Stahldavid/sensegrep/commit/bde06a75e8327c4a5c0a4c90db9b03f06c00a3b6) Thanks [@Stahldavid](https://github.com/Stahldavid)! - Preserve structural metadata in small files, improve token-budget context selection and cluster coherence, expose graph resolution evidence, accelerate bounded duplicate scans, and support serialized byte budgets for search. Distinguish ranking strength from answer sufficiency. Chunking policy changes trigger an atomic rebuild on the next explicit indexing command.

## 1.16.0

### Minor Changes
Expand Down
2 changes: 1 addition & 1 deletion packages/core/package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@sensegrep/core",
"version": "1.16.0",
"version": "1.17.0",
"type": "module",
"main": "./dist/index.js",
"types": "./dist/index.d.ts",
Expand Down
2 changes: 1 addition & 1 deletion packages/core/src/semantic/chunk-limits.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@ import { tokenCounterIdentity } from "./token-count.js"
import { getEmbeddingConfig, type EmbeddingConfig } from "./embedding-config.js"

const CHARS_PER_TOKEN = 4
const CHUNKING_SIGNATURE_VERSION = 5
const CHUNKING_SIGNATURE_VERSION = 6

export type GeneralChunkLimits = {
max: number
Expand Down
17 changes: 17 additions & 0 deletions packages/core/src/semantic/chunking.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,23 @@ import { describe, expect, it } from "vitest"
import { getGeneralChunkLimits } from "./chunk-limits.js"
import { Chunking } from "./chunking.js"

describe("small source files", () => {
const cases = [
["caller.ts", "import { debit as allowDebit } from './rules';\nexport function purchase() { return allowDebit(); }", "typescript", "purchase"],
["mail.py", "class MailQueue:\n def retry_failed_delivery(self, attempts):\n return 'retry' if attempts < 3 else 'manual_review'\n", "python", "retry_failed_delivery"],
["Invoice.java", "class Invoice { public boolean isOverdue(int days) { return days > 30; } }", "java", "Invoice"],
["Status.vue", '<script setup lang="ts">\nconst label = "Ready"\n</script>\n<template><span>{{ label }}</span></template>', "vue", undefined],
] as const
it.each(cases)("preserves metadata in %s without padding", async (file, content, language, symbol) => {
expect(content.length).toBeLessThan(200)
for (const chunks of [await Chunking.chunkAsync(content, file), (await Chunking.analyzeAsync(content, file)).chunks]) {
expect(chunks.length).toBeGreaterThan(0)
expect(chunks.some((chunk) => chunk.language === language)).toBe(true)
if (symbol) expect(chunks.some((chunk) => chunk.symbolName === symbol)).toBe(true)
}
})
})

describe("Chunking oversized content", () => {
it("splits very large single-line code into safe chunks", () => {
const content = `const payload = "${"a".repeat(getGeneralChunkLimits().max * 2)}";`
Expand Down
Loading
Loading