Web Attribution and Compliance Scanner
A fast, portable tool for checking web scraping compliance across robots.txt, RSL licenses, TDM policies, and Cloudflare Markdown for Agents. Built with Rust for the OpenAttribution initiative.
PolicyCheck helps you scrape responsibly by checking multiple compliance signals:
- β Robots.txt - What paths you can crawl (REP/RFC 9309)
- π RSL Licenses - Required licensing terms (Responsible Sourcing License)
- π― Content Signals - AI usage preferences (Cloudflare's policy framework)
- π€ TDM Policies - Text & Data Mining permissions (W3C TDMRep)
- π Markdown for Agents - Cloudflare edge markdown delivery detection
- π Privacy Controls - DNT, GPC signals (coming soon)
- π§ Security Contacts - Who to contact about scraping (coming soon)
- π€ AI Bot Analysis - Check 26 known AI crawlers (GPTBot, ClaudeBot, CCBot, etc.)
- π― Content Signals - Detect Cloudflare's AI policy signals (search, ai-input, ai-train)
- π CSV Export - Major AI bots as columns for advertiser analysis
- π Fast - Built with Rust, battle-tested parser (34M+ robots.txt files)
- π¦ Portable - Single binary, no dependencies
- π Comprehensive - User agents, crawl delays, sitemaps, paths, licenses
- π RSL License Detection - Automatically finds Responsible Sourcing Licenses
- π Markdown for Agents - Detect Cloudflare's edge markdown delivery and Content-Signal HTTP headers
- π Multiple Formats - Table, JSON, CSV, or compact text output
- π HTTP API - Run as a service for integration
- π CSV Batch Processing - Analyze thousands of URLs concurrently
- β‘ Concurrent - Parallel URL analysis
Try PolicyCheck instantly at openattribution.org/policycheck
- π No installation required
- π Interactive analysis with visual results
- π₯ Export to CSV for bulk analysis
- π€ See AI bot blocking status at a glance
Perfect for quick checks before integrating the API or CLI.
Requires Rust 1.75+:
cargo install policycheckFor development or the latest unreleased features:
git clone https://github.com/openattribution-org/policycheck.git
cd policycheck
cargo build --release -p policycheckThe binary will be at target/release/policycheck.
Use the core parsing library in your own project (no network I/O, WASM-compatible):
cargo add policycheck-coreuse policycheck_core::PolicyAnalyzer;
let analyzer = PolicyAnalyzer::new("GPTBot".to_string());
let result = analyzer.analyze(
"https://www.nytimes.com",
"User-agent: GPTBot\nDisallow: /\n",
None, // TDM rules
None, // Markdown probe data
);
assert!(!result.is_path_allowed);# Analyze a single URL
policycheck analyze --url https://www.nytimes.com
# Check multiple URLs
policycheck analyze \
--url https://www.nytimes.com \
--url https://github.com \
--url https://techcrunch.com
# Analyze from CSV file (advertiser use case)
policycheck analyze --csv publishers.csv --format csv --output results.csv
# Check for specific user agent
policycheck analyze --url https://www.nytimes.com --user-agent GPTBot
# Output as JSON
policycheck analyze --url https://www.nytimes.com --format json
# Output as CSV with AI bot columns
policycheck analyze --url https://www.nytimes.com --format csv
# Save to file
policycheck analyze --url https://www.nytimes.com --output results.jsonPolicyCheck analyzes 26 known AI crawlers including GPTBot, ClaudeBot, CCBot, and more. Perfect for two key use cases:
Check which AI training bots can access your content:
policycheck analyze --url https://www.nytimes.com --format compactShows comprehensive breakdown of which bots are blocked vs allowed.
Analyze multiple publishers to see which ones block AI search engines (affecting brand visibility):
policycheck analyze --csv publishers.csv --format csv --output analysis.csvExample output:
URL,Status,Path Allowed,RSL Licenses,TDM Reserved,GPTBot,ClaudeBot,Google-Extended,Meta-ExternalAgent,CCBot,Bytespider,OAI-SearchBot,PerplexityBot
https://www.nytimes.com,Success,Yes,0,N/A,Blocked,Blocked,Blocked,Blocked,Blocked,Blocked,Blocked,Blocked
https://github.com,Success,Yes,0,N/A,Allowed,Allowed,Allowed,Allowed,Allowed,Allowed,Allowed,Allowed
https://techcrunch.com,Success,Yes,0,N/A,Blocked,Blocked,Blocked,Allowed,Blocked,Blocked,Allowed,AllowedKey insights:
- NYTimes: Blocks all AI bots (zero AI search visibility)
- GitHub: Allows all AI bots (maximum AI discoverability)
- TechCrunch: Selectively blocks training bots, allows some search bots
Perfect for advertisers evaluating whether publisher placements will appear in ChatGPT, Perplexity, Claude, etc.
PolicyCheck automatically detects RSL license directives from robots.txt files. RSL extends the Robots Exclusion Protocol to enable websites to declare governing license documents for automated crawlers.
RSL introduces a License: directive that can be:
- Global: Outside any User-agent group (applies to all bots)
- Group-scoped: Inside a User-agent group (applies only to that bot)
Precedence rule: Group-scoped licenses override global licenses.
# Global license (applies to all bots unless overridden)
License: https://acme.com/global-license.xml
User-agent: *
Disallow: /private/
Allow: /public/
User-agent: GPTBot
Disallow: /
License: https://acme.com/gptbot-specific-license.xml
In this example:
- Most bots will see the global license
- GPTBot will see only the group-scoped license (global is ignored)
Real-world example: NYTimes blocks AI bots comprehensively:
policycheck analyze --url https://www.nytimes.com --user-agent GPTBot
# Shows: Blocked, with legal notice about prohibited usesPolicyCheck reports three license fields:
active_licenses: The licenses that actually apply (follows RSL precedence rules)global_licenses: Licenses defined outside user-agent groupsgroup_licenses: Licenses defined for the specific user agent
Compact output example:
================================================================================
URL: https://www.nytimes.com
Robots.txt: https://www.nytimes.com/robots.txt
Status: β Success
User Agents:
β’ *
β’ GPTBot
β’ ClaudeBot
β’ (40+ more...)
Path Access (for GPTBot): β Disallowed
AI Bot Analysis:
π« GPTBot: Blocked
π« ClaudeBot: Blocked
π« CCBot: Blocked
β Googlebot: Allowed (with restrictions)
Sitemaps:
β’ https://www.nytimes.com/sitemaps/new/news.xml.gz
β’ (15+ more sitemaps)
================================================================================
JSON output example:
{
"url": "https://github.com",
"robots_url": "https://github.com/robots.txt",
"status": "success",
"user_agents": ["*"],
"ai_bot_analysis": [
{"bot_name": "GPTBot", "company": "OpenAI", "category": "Training", "status": "allowed"},
{"bot_name": "ClaudeBot", "company": "Anthropic", "category": "Training", "status": "allowed"}
],
"global_licenses": [],
"group_licenses": [],
"active_licenses": [],
"crawl_delay": null,
"sitemaps": ["https://github.com/sitemap.xml"],
"is_path_allowed": true
}For more information about RSL, see the RSL Standard.
PolicyCheck automatically detects Content Signals - Cloudflare's framework for expressing AI usage preferences in robots.txt. Adopted by over 3.8 million domains using Cloudflare's managed robots.txt.
Content Signals allow websites to express preferences for how their content can be used after it's been accessed. Three signals are defined:
search- Traditional search indexing and results (not AI-generated summaries)ai-input- Inputting content into AI models (RAG, grounding, generative AI search)ai-train- Training or fine-tuning AI models
User-agent: *
Content-Signal: search=yes, ai-train=no, ai-input=yes
Allow: /
Values can be yes (permitted) or no (not permitted). Omitting a signal means no preference is expressed.
Compact format:
Content Signals:
β search: yes
β ai-train: no
β ai-input: yes
CSV format includes columns: CS-Search, CS-AI-Input, CS-AI-Train
JSON format:
{
"content_signal_search": "yes",
"content_signal_ai_input": "yes",
"content_signal_ai_train": "no"
}policycheck analyze --url https://blog.cloudflare.com --format compactCloudflare's blog permits all AI usage:
search=yes- Allowed in search indexesai-input=yes- Allowed for AI search/RAGai-train=yes- Allowed for model training
For more information, see Cloudflare's Content Signals announcement.
PolicyCheck detects whether a site supports Cloudflare's Markdown for Agents feature. When enabled, Cloudflare converts HTML to clean markdown at the edge when an AI system sends Accept: text/markdown.
- Whether the site returns
Content-Type: text/markdown(supported) - Token count from the
x-markdown-tokensresponse header - Per-response licence signals from the
Content-SignalHTTP header
PolicyCheck sends a GET request with Accept: text/markdown to the target URL. If Cloudflare's feature is enabled, the response comes back as markdown with additional headers. This probe runs concurrently with robots.txt and TDM fetches.
Note: This is different from robots.txt Content Signals. Robots.txt signals are declared in a text file. Markdown for Agents Content-Signal is an HTTP response header returned with the converted content.
Compact format:
Markdown for Agents:
β Supported: Yes
π Token Count: 6960
HTTP Content Signals:
search: yes
ai-input: yes
ai-train: yes
JSON format:
{
"markdown_agents": {
"supported": true,
"token_count": 6960,
"http_content_signal_search": "yes",
"http_content_signal_ai_input": "yes",
"http_content_signal_ai_train": "yes"
}
}CSV format includes columns: Markdown, Markdown Tokens, MD-CS-Search, MD-CS-AI-Input, MD-CS-AI-Train
policycheck analyze --url https://blog.cloudflare.com --format compactCloudflare's blog supports Markdown for Agents with all signals permitted. Available on Cloudflare Pro+ plans (~20% of the web is behind Cloudflare).
Perfect for bulk analysis with AI bot columns:
policycheck analyze --csv publishers.csv --format csv --output analysis.csvCreates a spreadsheet with major AI bots as columns - ideal for Excel/Google Sheets analysis:
URL,Status,Path Allowed,RSL Licenses,TDM Reserved,GPTBot,ClaudeBot,Google-Extended,Meta-ExternalAgent,CCBot,Bytespider,OAI-SearchBot,PerplexityBot
https://www.nytimes.com,Success,Yes,0,N/A,Blocked,Blocked,Blocked,Blocked,Blocked,Blocked,Blocked,Blocked
https://github.com,Success,Yes,0,N/A,Allowed,Allowed,Allowed,Allowed,Allowed,Allowed,Allowed,AllowedPerfect for quick checks:
policycheck analyze --url https://github.com --format tableShows summary information in a clean ASCII table.
Detailed, human-readable output with full AI bot breakdown:
policycheck analyze --url https://www.nytimes.com --format compactShows all details including blocked/allowed AI bots, paths, sitemaps, and licenses.
For programmatic use:
policycheck analyze --url https://www.nytimes.com --format json > results.jsonIncludes ai_bot_analysis array with per-bot status - perfect for integration with other tools.
The PolicyCheck API is available at https://policycheck-d7wv0g.fly.dev
No authentication required for public use. Rate limits may apply.
Start the HTTP API server locally:
policycheck serve --port 3000 --host 0.0.0.0Features:
- β
CORS enabled (all origins, or set
ALLOWED_ORIGINSenv var) - β JSON request/response
- β Concurrent request handling
- β 10s timeout per URL
- β 1MB request body limit
- β Max 100 URLs per request
Health check endpoint.
Response:
{
"status": "healthy",
"service": "policycheck",
"version": "0.2.1"
}Analyze robots.txt and RSL licenses for given URLs.
Request:
{
"urls": ["https://www.nytimes.com", "https://github.com"],
"user_agent": "GPTBot"
}Success Response:
{
"total": 2,
"successful": 2,
"failed": 0,
"results": [
{
"url": "https://www.nytimes.com",
"robots_url": "https://www.nytimes.com/robots.txt",
"status": "success",
"user_agents": ["*", "GPTBot", "ClaudeBot", "..."],
"crawl_delay": null,
"sitemaps": ["https://www.nytimes.com/sitemaps/new/news.xml.gz"],
"allowed_paths": [],
"disallowed_paths": ["/"],
"is_path_allowed": false,
"global_licenses": [],
"group_licenses": [],
"active_licenses": [],
"ai_bot_analysis": [
{"bot_name": "GPTBot", "company": "OpenAI", "category": "Training", "status": "blocked"},
{"bot_name": "ClaudeBot", "company": "Anthropic", "category": "Training", "status": "blocked"}
],
"error": null
}
]
}Error Response (with failures):
{
"total": 2,
"successful": 1,
"failed": 1,
"results": [
{
"url": "https://invalid-domain-xyz.com",
"robots_url": "https://invalid-domain-xyz.com/robots.txt",
"status": "fetch_error",
"error": "Failed to fetch robots.txt",
"user_agents": [],
"ai_bot_analysis": []
},
{
"url": "https://github.com",
"status": "success",
"error": null
}
]
}Using production API:
curl -X POST https://policycheck-d7wv0g.fly.dev/analyze \
-H "Content-Type: application/json" \
-d '{
"urls": ["https://www.nytimes.com"],
"user_agent": "GPTBot"
}'Using local server:
curl -X POST http://localhost:3000/analyze \
-H "Content-Type: application/json" \
-d '{
"urls": ["https://www.nytimes.com"],
"user_agent": "GPTBot"
}'Create a CSV file with URLs to check:
url
https://acme.com
https://example.org
https://test.ioOr with identifiers for tracking:
source_id,url
acme,https://acme.com
example,https://example.org
test,https://test.ioAnalyze all URLs:
policycheck analyze --csv partners.csv --format compact > results.txtPolicyCheck will automatically:
- Find the URL column (looks for headers containing "url", "link", "website", etc.)
- Default to the first column if no URL header is found
- Add
https://prefix if missing - Skip empty rows
- Process all URLs in parallel
Note: Only the URL column is used for analysis. Additional columns (like source_id) can be present for your own tracking but are ignored by PolicyCheck.
import requests
def check_ai_bot_access(urls, user_agent="GPTBot"):
response = requests.post(
"http://localhost:3000/analyze",
json={"urls": urls, "user_agent": user_agent}
)
return response.json()
# Advertiser use case: check which publishers block AI bots
publishers = [
"https://www.nytimes.com",
"https://github.com",
"https://techcrunch.com"
]
result = check_ai_bot_access(publishers)
for site in result['results']:
print(f"\n{site['url']}")
print(f" GPTBot access: {'β Blocked' if not site['is_path_allowed'] else 'β
Allowed'}")
# Check specific AI bots
for bot in site['ai_bot_analysis']:
if bot['bot_name'] in ['GPTBot', 'OAI-SearchBot', 'PerplexityBot']:
status = 'β' if bot['status'] == 'blocked' else 'β
'
print(f" {status} {bot['bot_name']}")const axios = require('axios');
async function checkAIBotAccess(urls, userAgent = 'GPTBot') {
const response = await axios.post('http://localhost:3000/analyze', {
urls,
user_agent: userAgent
});
return response.data;
}
// Advertiser use case: analyze publisher AI visibility
const publishers = [
'https://www.nytimes.com',
'https://github.com',
'https://techcrunch.com'
];
const result = await checkAIBotAccess(publishers);
console.log(`Analyzed ${result.total} publishers`);
result.results.forEach(site => {
const blocked = site.ai_bot_analysis.filter(b => b.status === 'blocked').length;
const allowed = site.ai_bot_analysis.filter(b => b.status === 'allowed').length;
console.log(`${site.url}: ${blocked} blocked, ${allowed} allowed`);
});package main
import (
"bytes"
"encoding/json"
"net/http"
)
type AnalyzeRequest struct {
URLs []string `json:"urls"`
UserAgent string `json:"user_agent"`
}
func checkCompliance(urls []string, userAgent string) (*AnalyzeResponse, error) {
reqBody := AnalyzeRequest{URLs: urls, UserAgent: userAgent}
jsonData, _ := json.Marshal(reqBody)
resp, err := http.Post(
"http://localhost:3000/analyze",
"application/json",
bytes.NewBuffer(jsonData),
)
if err != nil {
return nil, err
}
defer resp.Body.Close()
var result AnalyzeResponse
json.NewDecoder(resp.Body).Decode(&result)
return &result, nil
}Pull and run the official image from GitHub Container Registry:
# Latest version
docker pull ghcr.io/openattribution-org/policycheck:latest
docker run -p 3000:3000 ghcr.io/openattribution-org/policycheck:latest
# Specific version
docker pull ghcr.io/openattribution-org/policycheck:0.2.1
docker run -p 3000:3000 ghcr.io/openattribution-org/policycheck:0.2.1Multi-platform images available for linux/amd64 and linux/arm64.
FROM rust:1.85-bookworm as builder
WORKDIR /app
COPY Cargo.toml Cargo.lock ./
COPY crates ./crates
RUN cargo build --release -p policycheck
FROM debian:bookworm-slim
RUN apt-get update && apt-get install -y ca-certificates && rm -rf /var/lib/apt/lists/*
COPY --from=builder /app/target/release/policycheck /usr/local/bin/policycheck
EXPOSE 3000
CMD ["policycheck", "serve", "--host", "0.0.0.0", "--port", "3000"]Build and run:
docker build -t policycheck .
docker run -p 3000:3000 policycheckapiVersion: apps/v1
kind: Deployment
metadata:
name: policycheck
spec:
replicas: 3
selector:
matchLabels:
app: policycheck
template:
metadata:
labels:
app: policycheck
spec:
containers:
- name: policycheck
image: ghcr.io/openattribution-org/policycheck:latest
ports:
- containerPort: 3000
livenessProbe:
httpGet:
path: /health
port: 3000
---
apiVersion: v1
kind: Service
metadata:
name: policycheck-service
spec:
selector:
app: policycheck
ports:
- protocol: TCP
port: 80
targetPort: 3000
type: LoadBalancerFor local development or production deployments using podman-compose:
# podman-compose.yml
services:
policycheck:
image: ghcr.io/openattribution-org/policycheck:latest
ports:
- "3000:3000"
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:3000/health"]
interval: 30s
timeout: 10s
retries: 3Run with:
podman-compose up -d- Robots.txt parsing (REP/RFC 9309)
- RSL license detection
- User agent matching
- Crawl delay detection
- Sitemap discovery
- Path permission checking
- CSV batch processing
- HTTP API server
- Multiple output formats
- TDM (Text & Data Mining) policy detection (
/.well-known/tdmrep.json) - Content Signals (Cloudflare AI policy framework)
- Markdown for Agents detection (Cloudflare edge markdown delivery)
- Security contact discovery (
/.well-known/security.txt) - Privacy control detection (DNT, GPC)
- AI plugin manifest detection (
/.well-known/ai-plugin.json) - OpenID configuration for gated content
- Caching layer for repeated checks
- GitHub Action for PR compliance checks
- Pre-commit hook for URL validation
Analyze robots.txt and RSL licenses from URLs.
Options:
-u, --url <URL>- URL to analyze (can be repeated)-c, --csv <PATH>- CSV file containing URLs-a, --user-agent <AGENT>- User agent to check (default: "*")-f, --format <FORMAT>- Output format: table, json, compact (default: table)-o, --output <PATH>- Save output to file
Start HTTP API server.
Options:
-p, --port <PORT>- Port to listen on (default: 3000)--host <HOST>- Host to bind to (default: 127.0.0.1)
PolicyCheck is designed for speed:
- Concurrent analysis: Multiple URLs analyzed in parallel
- Optimized builds: Release builds use LTO and aggressive optimization
- Battle-tested parser: Based on
texting_robots, tested against 34M+ real-world files - Low memory footprint: Efficient parsing with minimal allocations
Typical performance:
- Single URL analysis: ~50-200ms (network dependent)
- 100 URLs analyzed concurrently: ~2-5 seconds
The Paradox: robots.txt exists for bots to check before crawling, but some sites block datacenter IPs, preventing policy checkers from accessing robots.txt.
Why this happens:
- Sites like Medium block cloud provider IP ranges to prevent scraping
- PolicyCheck runs from cloud infrastructure (Fly.io)
- Appears as "generic scraper" rather than "compliance checker"
How legitimate crawlers solve this:
- IP whitelisting - Googlebot, GPTBot, ClaudeBot use published IP ranges that sites whitelist
- Reverse DNS verification - Sites verify bot identity via DNS lookups
- User agent + IP combo - Both must match expected patterns
Impact on PolicyCheck:
- β Works: Most sites (GitHub, Cloudflare, NYTimes, etc.)
- β Blocked: Some sites that aggressively block datacenter IPs (e.g., Medium)
- π‘ Workaround: Test locally with
cargo runor use sites that don't block datacenter IPs
Why this matters: If compliance checkers are blocked, publishers can't verify their own policies are working correctly. This is a gap in the current web crawling ecosystem.
- Input validation: URLs are validated before processing
- Size limits: robots.txt files limited to 500KB (Google's recommendation), request bodies limited to 1MB
- Rate limiting: Max 100 URLs per request
- Timeouts: HTTP requests timeout after 10 seconds
- No arbitrary code execution: Pure parsing, no eval or dynamic code
- CORS enabled: API server has CORS enabled by default
PolicyCheck implements the following standards:
- β RFC 9309: Robots Exclusion Protocol (REP)
- β RSL Standard: Responsible Sourcing License
- β Content Signals: Cloudflare's AI Policy Framework (CC0 License)
- β W3C TDMRep: Text and Data Mining Reservation Protocol
- β Markdown for Agents: Cloudflare edge markdown delivery and Content-Signal HTTP headers
- π§ RFC 9116: security.txt (planned)
- π§ RFC 8615: Well-Known URIs (planned)
PolicyCheck is part of the OpenAttribution initiative. Contributions welcome!
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Make your changes
- Add tests if applicable
- Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
This project is licensed under the MIT License - see LICENSE for details.
Third-party software notices are in NOTICE.
- texting_robots (MIT OR Apache-2.0) - Robust robots.txt parsing by @Smerity
- reqwest (MIT OR Apache-2.0) - HTTP client
- clap (MIT OR Apache-2.0) - CLI argument parsing
- axum (MIT) - HTTP server framework
- See NOTICE for complete attribution list
PolicyCheck is built for the OpenAttribution initiative, which aims to make web attribution transparent, accessible, and machine-readable.
Mission: Enable responsible AI development through clear content licensing and attribution standards.
- π Report issues: GitHub Issues
- π¬ Discussions: GitHub Discussions
- π§ Contact: openattribution.org
- π Website: OpenAttribution.org
Built with β€οΈ by the OpenAttribution community.
Special thanks to:
- @Smerity for texting_robots
- The Rust community for excellent tooling
- Everyone contributing to open web standards
Made with Rust π¦ | Part of OpenAttribution π | MIT Licensed π