What your market actually says about you — and whether your own community is lying to you about it.
Beta. It runs end to end and passes its --selftest, but flags and output formats may still change. Issues and feedback welcome.
Reddit is where buyers go for the opinion that isn't a paid review, which makes it the highest-signal free source of voice-of-customer data there is — and the one most likely to be quoted back at you by an LLM. Pulling it is easy. Pulling it without fooling yourself is the hard part, and that's what this does.
Three things separate a real read from a misleading one:
- The venue gap. A brand's own subreddit almost always reads better than neutral ground. The question is how much better — and whether that gap is big enough that a shopper who cross-checks will find a different story. This measures it, with the controls that make the number defensible.
- Peer-baselined complaints. Comparing a brand's complaint rate against the whole corpus is wrong: brand-mentioning documents are systematically longer, so every brand scores "above average" and the comparison says nothing. The baseline here is the other brands.
- Noise that isn't optional. AutoModerator removal notices were 28% of comment volume in one subreddit tested. Removed and deleted bodies are another large slice. Leave them in and your sentiment is measuring moderation policy.
No API key, no OAuth, no dependencies. Reads from Arctic Shift, the maintained public Reddit archive — which also means it works from hosts and CI runners that reddit.com itself blocks with a 403.
Voice of Customer r/SyntheticGear
==========================================================================
150 posts · 289 comments · 2026-01-01 → 2026-06-28
sentiment 51% positive · 7% negative · ratio 7.26
Most-asked (58 of 150 posts are questions — 39%)
22c 12p Should I return my Acme Audio?
21c 41p Anyone using Zenith for long flights?
21c 24p Acme Audio vs Northwind — which holds up better?
Share of voice
Acme Audio 51 27.3% ██████████████████
Northwind 48 25.7% █████████████████·
Kestrel 44 23.5% ████████████████··
Complaint fingerprint (100 = same as peer brands; >130 distinctive)
baseline: 150 brand-mentioning documents
Zenith n=44 19 pos / 5 neg sharpest: customer service 320
customer service 320 █████████████████████▶
returns & refunds 146 ████████████████······
Northwind n=48 14 pos / 8 neg sharpest: defective on arrival 278
defective on arrival 278 █████████████████████▶
quality / durability 238 █████████████████████▶
No package to install and nothing to build. You need Python 3.8+ and the single reddit_voc.py file.
# Option A: clone the repo
git clone https://github.com/seoprocheck/reddit-voc.git
cd reddit-voc
# Option B: grab just the one file
curl -O https://raw.githubusercontent.com/seoprocheck/reddit-voc/main/reddit_voc.pyConfirm it runs (no network needed, exercises the analysis internals):
python3 reddit_voc.py --selftest # expect: 15/15 passedThen try it on the bundled fixture before hitting the live archive:
python3 reddit_voc.py --fixture fixtures/sample-corpus.json \
--brands "Acme Audio,Zenith,Northwind,Kestrel" --min-docs 25python3 reddit_voc.py SEO --days 30 # what a community talks about
python3 reddit_voc.py SEO --days 90 --query "core update" # narrow to a topic
python3 reddit_voc.py headphones --brands "sony,bose,sennheiser,anker"
python3 reddit_voc.py headphones --brands "sony,bose" --themes --json
python3 reddit_voc.py headphones --brands "sony,bose" --verbatims # real quotes, not just numbersA fingerprint tells you that a brand over-indexes on "battery life"; --verbatims shows you the actual comments behind it, top-voted, split into what people praise and what they complain about, plus one real example per complaint axis. The sentiment tag is decided on the same sentence that gets quoted, so the text and its label can never disagree. It is still a lexical read — sarcasm and negation slip through — so treat every quote as a thread worth opening, not a line to paste into a slide. --quotes N sets how many per brand per tone (default 3).
Measure a brand community against neutral ground:
python3 reddit_voc.py acmeaudio --brand acme --vs headphonesPull once, analyse many times — the archive is free and run by volunteers, so don't re-fetch what you already have:
python3 reddit_voc.py headphones --days 180 --save headphones.json
python3 reddit_voc.py --fixture headphones.json --brands "sony,bose" --min-docs 50Try it instantly with the bundled fixture — fictional brands, generated text, no real Reddit content:
python3 reddit_voc.py --fixture fixtures/sample-corpus.json \
--brands "Acme Audio,Zenith,Northwind,Kestrel" --min-docs 25Verify the analysis internals — word-boundary matching, ratio smoothing, peer baselining, and the invariant that a corpus compared against itself must return a gap of exactly 1.0:
python3 reddit_voc.py --selftestThe default complaint axes are generic — price, durability, service, returns, warranty, delivery, setup, sizing. Swap them for your category with a JSON file of {label: regex}:
{
"battery life": "battery (life|drain)|dies after|won't hold a charge",
"comfort": "clamp(ing)? force|hurts my ears|too tight|sore after",
"bluetooth drop": "drops? (out|connection)|keeps disconnecting|pairing"
}python3 reddit_voc.py headphones --brands "sony,bose" --issues audio-issues.jsonSentiment here is lexical, not classified. Documents are matched against positive and negative marker patterns. That is robust enough for comparison — brand A versus brand B, own sub versus neutral — and not robust enough to quote as "49% of customers are happy". It cannot read sarcasm or negation. Use it to rank, not to score.
The ratio is smoothed. (positive+1)/(negative+1), not positive/negative. Twenty-two positives and zero negatives is not a ratio of 22; it's a ratio you don't have the data to state.
Gaps get a robustness check. Whether the brand's own side is filtered to documents naming the brand, and whether the ratio is smoothed, are both arbitrary calls that move the number. The tool re-runs the gap under all four combinations and tells you the range. If the variants disagree, it says so and tells you to compare brands by rank instead of by multiple — a finding that only holds under one arbitrary choice is not a finding.
Small samples are flagged, not hidden. Under 50 documents on either side and you get a warning; brands under --min-docs mentions are listed as skipped rather than silently dropped.
What you cannot see. Removed and deleted content is gone. Moderation strips complaints as well as spam, in unknown proportion — one subreddit tested deletes roughly three quarters of submissions. Treat every rate here as a rate among surviving posts.
Check the subreddit is what you think it is. Brand subreddits are far rarer than people assume, and brand names collide. r/Corona is a city in California, not the drink; r/Java is a programming language, not a coffee. Most major consumer brands have no subreddit at all, and the general category community is where the conversation actually happens.
Arctic Shift rate-limits with HTTP 422 and the message "slow down a bit" rather than a 429; the client backs off and retries automatically. Requests are spaced by default. Full-text --query works on posts only, which is an upstream constraint.
MIT © SEO Pro Check · built by @seoprocheck.