Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
352 changes: 183 additions & 169 deletions docs/wiki/Audit-Skill.md

Large diffs are not rendered by default.

28 changes: 18 additions & 10 deletions docs/wiki/CI-Recipes.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,9 +5,9 @@ GitLab CI, and a generic shell script — plus how `omnirank fix` fits the same
a second, distinct gate. Each recipe turns a specific class of problem into a non-zero
exit that fails the job.

OmniRank is not published to PyPI as of v0.3.0, so every example below installs straight
from git. The `pip install "package @ git+URL#subdirectory=..."` syntax used here was
verified to work against the real repository.
OmniRank is not published to PyPI, so every example below installs straight from git. The
`pip install "package @ git+URL#subdirectory=..."` syntax used here was verified to work
against the real repository.

## GitHub Actions

Expand Down Expand Up @@ -110,15 +110,18 @@ run cleanup after a failure, as shown here.

Not every gate name in the schema's `audit.failOn` enum can actually cause `--fail-on` to
fail a build, and putting all 42 in the list creates false confidence rather than more
protection. Verified directly against every gate's severity in `scripts/py/omnirank/gates/`:
protection. Verified directly against every gate's severity in `scripts/py/omnirank/gates/`
(`tests/test_repo_docs.py::test_documented_gate_counts_match_the_registry` re-derives
these two numbers from `registry.py` and the config schema on every CI run, so they cannot
drift silently):

| Group | Gates | Effect on `--fail-on` |
|---|---|---|
| Warning-only (16) | `og`, `hreflang`, `image-dims`, `citation-licence`, `lastmod-inflation`, `faq` (downgraded from error in v0.2.1), `duplicate-title`, `duplicate-description`, `canonical-cluster`, `hreflang-reciprocity`, `page-weight`, `compression`, `render-blocking`, `heading-order`, `image-alt`, `schema-required` (last three new in v0.4.0) | Every finding these gates can produce is `severity: "warning"`; `has_failures()` only counts errors. Listing them has zero effect on the exit code, ever. |
| Info-only (4, all new in v0.4.0) | `hsts`, `nosniff`, `csp`, `referrer-policy` | Every finding these gates can produce is `severity: "info"`. OmniRank reports your security headers as inventory facts and never grades them — see [[Audit-Skill#security-v040]]. |
| Mixed, but never error (1, new in v0.4.0) | `link-text` | Its `.empty` id is `warning`, its `.generic` id is `info` — neither is ever `error`, so the gate as a whole cannot gate a build. |
| Warning-only (16) | `og`, `hreflang`, `image-dims`, `citation-licence`, `lastmod-inflation`, `faq` (downgraded from error in v0.2.1), `duplicate-title`, `duplicate-description`, `canonical-cluster`, `hreflang-reciprocity`, `page-weight`, `compression`, `render-blocking`, `image-alt`, `heading-order`, `schema-required` | Every finding these gates can produce is `severity: "warning"`; `has_failures()` only counts errors. Listing them has zero effect on the exit code, ever. |
| Info-only (4) — new in v0.4.0 | `hsts`, `nosniff`, `csp`, `referrer-policy` | The four `security` header gates. `info`-severity findings cost zero points and can never fail a build under any circumstance — see [[Security-Layer]]. |
| Mixed warning/info (1) | `link-text` | Emits `seo.link-text.empty` (warning) and `seo.link-text.generic` (info). Listing it catches nothing at error severity either way. |
| Removed from the enum entirely (v0.2.1) | `crawl-hygiene` | Its dedicated check (`hygiene.check_removed()`) needs a removed-URL list no config field supplies, so it could never fire from a plain run — v0.2.1 dropped it from the schema rather than ship a dead gate name. |
| Can actually fail a build (21) | `h1`, `canonical`, `title-length` (missing only), `description-length` (missing only), `answer-block`, `speakable`, `llms-txt`, `llms-full`, `facts-json`, `ai-allowlist`, `schema`, `schema-fabrication`, `noindex-in-sitemap`, `response-time`, `sitemap-health`, and — new in v0.4.0 — `lang`, `robots-sitemap`, `canonical-target` (its `.redirects` id is a warning; `.noindexed`/`.not-found` are errors), `hreflang-noindex`, `mixed-content` (active-subresource id only), `https-redirect` | These can produce an error-severity finding and gate a build |
| Can actually fail a build (21) | `h1`, `canonical`, `title-length` (missing only), `description-length` (missing only), `answer-block`, `speakable`, `llms-txt`, `llms-full`, `facts-json`, `ai-allowlist`, `schema`, `schema-fabrication`, `noindex-in-sitemap`, `response-time`, `sitemap-health`, `mixed-content`, `https-redirect`, `robots-sitemap`, `canonical-target`, `hreflang-noindex`, `lang` | These can produce an error-severity finding and gate a build. The last six are new in v0.4.0 — three `security`, three contradiction gates, plus `lang`. |

**The practical guidance:** pick gates that map to problems severe enough to block a
merge, not the full list. A reasonable starting set for most sites is `h1 canonical
Expand All @@ -127,8 +130,13 @@ schema` (structural SEO baseline) plus, once `geo-artifacts` is wired into your
production — this is the check that would have caught the OpenNext/CloudFront 403 trap
described in [[GEO-Artifacts-Skill#what-is-the-opennextcloudfront-403-trap]] before a
human noticed). Add `answer-block` once you have deliberately built AEO-oriented pages —
gating on it before you have any answer blocks just fails every build. Leave the 21
warning- or info-only gates out of `--fail-on` entirely.
gating on it before you have any answer blocks just fails every build. Consider adding
`robots-sitemap`, `canonical-target` and `hreflang-noindex` once you trust your sitemap
and canonical hygiene — these are the new v0.4.0 contradiction gates, and every finding
they can produce is provable from the site's own declarations, so false positives should
be rare; see [[Contradictions]] before relying on that for `robots-sitemap` specifically,
since its coverage varies by Python version. Leave the 21 warning-only/info-only/mixed
gates out of `--fail-on` entirely — they can never move the exit code.

Findings from gates you did **not** list in `--fail-on` are still computed and still land
in the JSON report and terminal summary — they just do not fail the build. Nothing is
Expand Down
11 changes: 5 additions & 6 deletions docs/wiki/Claude-Code-Setup.md
Original file line number Diff line number Diff line change
Expand Up @@ -126,12 +126,11 @@ See [[Fix-Preview]] for the full flag reference and what its output looks like.

Both `SKILL.md` files list explicit "When NOT to use" cases. Asking Claude to "write our
JSON-LD" or "emit schema for this entity type" will not trigger either shipped skill —
that is `aeo-onpage`, which is on the roadmap and not present in v0.3.0. Asking it to
"submit this URL to Google" will not trigger anything either — that is the unshipped
`indexing` skill. Asking it to "apply the fix" or "write the canonical tag for me" will
also not do anything by itself: `omnirank fix` only prints a diff, and no skill in
v0.3.0 applies an edit to your source tree. See [[Roadmap]] for the full list of what
does not exist yet.
that is `aeo-onpage`, which is on the roadmap and not built yet. Asking it to "submit this
URL to Google" will not trigger anything either — that is the unshipped `indexing` skill.
Asking it to "apply the fix" or "write the canonical tag for me" will also not do
anything by itself: `omnirank fix` only prints a diff, and no shipped skill applies an
edit to your source tree. See [[Roadmap]] for the full list of what does not exist yet.

## See also

Expand Down
29 changes: 15 additions & 14 deletions docs/wiki/Configuration-Reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ being silently ignored.

## What does "Consumed" vs "Schema-only" mean?

Not every field the schema accepts is read by v0.3.0's shipped code. **Consumed** means a
Not every field the schema accepts is read by v0.4.0's shipped code. **Consumed** means a
shipped code path reads the field. **Schema-only** means the field is validated, stored,
and forward-compatible with a roadmap skill, but nothing in `audit`, `geo-artifacts`, or
`fix` reads it yet. Writing a schema-only field is not wasted — validation still checks
Expand Down Expand Up @@ -186,29 +186,30 @@ Free-form string map. **Schema-only** — validated, not read.
| `sampleSize` | integer, minimum `0` | no | `200` | Maximum URLs pulled from the sitemap for a crawl. `0` means no limit. |
| `failOn` | array of gate-name enum values | no | `[]` | Gate names that make `omnirank audit` exit `1` when they carry an error-severity finding. Overridden by `--fail-on` whenever that flag is present at all, even with zero names. |

`failOn`'s allowed values grew to **42 gate names** as of v0.4.0: `h1`,
`canonical`, `title-length`, `description-length`, `hreflang`, `og`, `image-dims`,
`answer-block`, `faq`, `speakable`, `llms-txt`, `llms-full`, `facts-json`,
`failOn`'s allowed values grew to **42 gate names** as of v0.4.0 (up from 28 at
v0.2.0/v0.2.1): `h1`, `canonical`, `title-length`, `description-length`, `hreflang`, `og`,
`image-dims`, `answer-block`, `faq`, `speakable`, `llms-txt`, `llms-full`, `facts-json`,
`ai-allowlist`, `citation-licence`, `sitemap-health`, `lastmod-inflation`, `schema`,
`schema-fabrication`, `duplicate-title`, `duplicate-description`, `noindex-in-sitemap`,
`canonical-cluster`, `hreflang-reciprocity`, `response-time`, `page-weight`,
`compression`, `render-blocking` (28 through v0.3.0), plus 14 in v0.4.0: `hsts`,
`nosniff`, `csp`, `referrer-policy`, `mixed-content`, `https-redirect` (`security`),
`robots-sitemap`, `canonical-target`, `hreflang-noindex` (indexability contradictions),
`schema-required` (structured data), and `image-alt`, `heading-order`, `link-text`,
`lang` (on-page).
`compression`, `render-blocking`, `hsts`, `nosniff`, `csp`, `referrer-policy`,
`mixed-content`, `https-redirect`, `robots-sitemap`, `canonical-target`,
`hreflang-noindex`, `schema-required`, `image-alt`, `heading-order`, `link-text`, `lang`.
The last 14 are new in v0.4.0 (`hsts` through `lang`) — six from the new `security` layer,
three indexability-contradiction gates, one structured-data gate, and four on-page
accessibility gates. See [[Security-Layer]] and [[Contradictions]].

Of those 42, **21 can actually produce an error-severity finding and trip `--fail-on`; 21
cannot** — 16 are warning-only, 4 (the `security` response-header gates) are info-only and
can never fail a build under any circumstance, and one (`link-text`) mixes warning and
info. Full breakdown: [[CI-Recipes#which-gates-can-actually-fail-a-build-with---fail-on]].

**`crawl-hygiene` was removed from this enum in v0.2.1** — it validated successfully but
matched no finding a plain `omnirank audit` run could ever produce, since the check that
would emit it needs an explicit removed-URL list no config field supplies. `sitemap-health`
is not inert: `hygiene.check_sitemap()` is wired in as of v0.2.1, distinguishing a
redirecting sitemap entry (warning) from a dead one (error).

Only **21 of the 42** can actually produce an error-severity finding and gate a build —
see [[CI-Recipes#which-gates-can-actually-fail-a-build-with---fail-on]] for the full
breakdown, including the four v0.4.0 security gates that are `info`-severity and can
never fail a build regardless of what you list.

## `smm`

| Field | Type | Description |
Expand Down
143 changes: 143 additions & 0 deletions docs/wiki/Contradictions.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,143 @@
# Contradictions

Five gates, new in v0.4.0, catch a site disagreeing with itself: a sitemap URL its own
`robots.txt` forbids crawling, a canonical pointing at a noindexed, missing or redirecting
page, and an `hreflang` alternate that is itself noindexed. Every finding is 100%
precision — both halves come from the site's own declarations, never a source OmniRank
has to trust.

Verified against `scripts/py/omnirank/gates/contradictions.py`,
`scripts/py/omnirank/robots.py`, and `tests/test_gates_contradictions.py`. This is the
theme release is named for: **"OmniRank finds where your site contradicts itself."**

## Why is this a distinct category from ordinary SEO gates?

Because a contradiction gate never has to trust an outside source to know it found a real
problem. `docs/research/2026-08-04-competitive-gap-analysis.md` §1 identifies this as the
one class of defect OmniRank can credibly claim 100% precision on: a page cannot be both
submitted for crawling and forbidden to crawl, cannot be both the canonical target and
noindexed, cannot be both an `hreflang` alternate and unindexable — these aren't
judgement calls about what Google *should* do, they're the site's own two statements
disagreeing with each other. A `seo.title.long` finding requires trusting that 60
characters is the right cutoff; a `seo.robots-sitemap.disallowed` finding requires
trusting nothing beyond what the site itself published in two files.

## The five gates

| Gate | Finding id | What it detects | Severity |
|---|---|---|---|
| `robots-sitemap` | `seo.robots-sitemap.disallowed` | A URL listed in `sitemap.xml` is also `Disallow`-ed to `Googlebot` in `robots.txt` | **error** |
| `canonical-target` | `seo.canonical-target.noindexed` | A page's canonical points at a URL carrying a `noindex` directive | **error** |
| `canonical-target` | `seo.canonical-target.not-found` | A canonical target returns `404` or `410` | **error** |
| `canonical-target` | `seo.canonical-target.redirects` | A canonical target itself 3xx-redirects | warning |
| `hreflang-noindex` | `seo.hreflang-noindex.alternate` | A page declares an `hreflang` alternate at a page that carries `noindex` | **error** |

Every one of these is `advisory`-tier for fixing — see [[Fix-Tiers-and-Applicability]] —
because each has **two opposite correct fixes** and only the site owner knows which is
intended. A sitemap URL blocked by `robots.txt` could mean "unblock it, we want it
crawled" or "drop it from the sitemap, it shouldn't be discoverable" — OmniRank cannot
tell which, and does not guess. `seo.hreflang-noindex.alternate` attaches to the
*declaring* page, not the noindexed target: naming a page engines may not index as a
locale alternate is discarded by engines wholesale, which is why the page that made the
claim — not the page that can't be indexed — is where the finding lands.

## A real, verified true positive

Auditing `https://danluu.com` with v0.4.0 fires `seo.canonical-target.not-found` on
`https://danluu.com/simple-architectures/`. The page's own markup declares:

```html
<link rel=canonical href=https://danluu.com/simple-architecture>
```

Note the missing trailing `s` — `/simple-architecture`, not `/simple-architectures/`.
That URL returns a genuine `404`:

```
$ curl -sI https://danluu.com/simple-architecture | head -1
HTTP/2 404
```

This is exactly the class of defect this gate group exists for: the page's own `<link
rel=canonical>` names a URL that does not exist, so every ranking signal this page has
earned is nominally being handed to a page search engines will never find. No external
data, no judgement call about SEO best practice — just the page's own declaration checked
against the page it points at.

## Two rules every check in this module follows

Both are load-bearing, and both exist because a false contradiction finding would be
worse than a missed one — the entire value proposition of this gate group is precision.

**1. Never judge a URL OmniRank did not actually see.** A canonical target already in the
crawled set is judged from the `PageData` OmniRank already fetched — no second request. A
target outside the crawled set is probed once, deduplicated across every page pointing at
it, up to `MAX_CANONICAL_PROBES` (25 per audit); anything past that budget is reported as
`budget-exceeded` in `notEvaluated`, never silently skipped. A 5xx or transport failure on
a probe is reported as `page-unreachable`, never as `.not-found` — a transient origin
error is not proof the page is missing, and calling it one would be a guess dressed as a
finding.

**Cross-host sitemap URLs are `notEvaluated`, not judged against the wrong host's
robots.txt.** `RobotFileParser.can_fetch()` discards the host and matches on path alone,
so a sitemap index legitimately listing URLs on a different host (a CDN, a blog
subdomain) would otherwise get those paths judged by an audited host's `robots.txt` that
never governed them at all — a fabricated finding on a correct sitemap. `check_sitemap_vs_robots()`
splits `sitemap_urls` by `robots.txt`'s own netloc first: only same-host URLs are
evaluated against it, and off-host URLs are reported `not-applicable` in `notEvaluated`
instead. This was found and fixed in the final v0.4.0 review, reproduced against two real
local HTTP servers on different hosts before the fix shipped.

**2. A gate that could not run reports why, never silently.** Every function in this
module returns a `NotEvaluated` entry rather than staying quiet when it cannot reach a
verdict — silence would be indistinguishable from a pass, which is the exact failure this
project exists to refuse. `robots-sitemap` returns `no-sitemap`, `page-unreachable`,
`not-applicable`, or `matcher-unsupported`, depending on what stopped it.

## Why does this gate sometimes refuse to answer?

`seo.robots-sitemap.disallowed` needs a robots.txt *matcher*, and the one OmniRank uses —
Python's own `urllib.robotparser` — does not behave identically across the Python
versions this project supports. CPython rewrote the matcher for RFC 9309 compliance in
**Python 3.14**, and backported part of that rewrite to **3.13** — but not all of it: a
3.13.7 interpreter lacks wildcard support that a 3.13.14 interpreter has. On the affected
versions, the matcher can silently miss a path wildcard (`Disallow: /*.pdf$`) or resolve
an overlapping `Allow`/`Disallow` pair by file order instead of RFC 9309's longest-match
rule — and the second failure mode is worse than the first, because it can fabricate a
disallow that RFC 9309 says should be allowed.

**This is deliberately not a Python-version check.** `omnirank/robots.py` runs two small
*behavioural probes* against the live interpreter — does it honour a wildcard, does it
pick the longest match — and only evaluates a site's robots.txt when its actual rules need
a capability the probe confirms this interpreter has. When the rules need a capability the
probe says is missing, the gate reports `matcher-unsupported` in `notEvaluated` rather
than answer with a matcher it knows may be wrong. Probing behaviour instead of checking
`sys.version_info` means a future backport, a distribution-specific patch, or a version
this page hasn't been updated for is picked up automatically and correctly, with no code
change here. Coverage of this one gate is broader on Python 3.14+ than on older
interpreters; correctness is identical on both, because an unevaluated verdict is never
wrong — it is a refusal, not a guess.

## Honest limits

- **A gate firing means a self-contradiction exists, not that fixing it is obvious.**
Every finding here is `advisory`-tier for a reason: OmniRank can prove the two
declarations disagree, but not which one the owner meant.
- **Coverage of `robots-sitemap` genuinely varies by Python version**, for the reason
above — see [[FAQ]] before assuming a clean `robots-sitemap` result means no
contradiction exists on Python 3.11–3.13.
- **Canonical-target probing is bounded at 25 URLs per audit.** A site with more than 25
distinct out-of-crawl canonical targets will see some reported `budget-exceeded` rather
than judged.
- **This module only ever reads declarations the site itself published** — `sitemap.xml`,
`robots.txt`, `<link rel=canonical>`, `noindex` meta, `hreflang`. It has no opinion about
whether those declarations reflect good SEO strategy, only whether they agree with each
other.

## See also

- [[Audit-Skill#indexability-contradictions--new-in-v040]] — the contradiction gates
alongside every other layer and pass
- [[Security-Layer]] — the other new v0.4.0 gate group
- [[Finding-Reference]] — all five contradiction finding ids, generated from the registry
- [[FAQ]] — the robots.txt matcher caveat, asked and answered directly
Loading