Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

reddit-voc

What your market actually says about you — and whether your own community is lying to you about it.

Python 3.8+ Zero dependencies No API key License: MIT Status: beta

Beta. It runs end to end and passes its --selftest, but flags and output formats may still change. Issues and feedback welcome.


Reddit is where buyers go for the opinion that isn't a paid review, which makes it the highest-signal free source of voice-of-customer data there is — and the one most likely to be quoted back at you by an LLM. Pulling it is easy. Pulling it without fooling yourself is the hard part, and that's what this does.

Three things separate a real read from a misleading one:

  • The venue gap. A brand's own subreddit almost always reads better than neutral ground. The question is how much better — and whether that gap is big enough that a shopper who cross-checks will find a different story. This measures it, with the controls that make the number defensible.
  • Peer-baselined complaints. Comparing a brand's complaint rate against the whole corpus is wrong: brand-mentioning documents are systematically longer, so every brand scores "above average" and the comparison says nothing. The baseline here is the other brands.
  • Noise that isn't optional. AutoModerator removal notices were 28% of comment volume in one subreddit tested. Removed and deleted bodies are another large slice. Leave them in and your sentiment is measuring moderation policy.

No API key, no OAuth, no dependencies. Reads from Arctic Shift, the maintained public Reddit archive — which also means it works from hosts and CI runners that reddit.com itself blocks with a 403.

Voice of Customer  r/SyntheticGear
==========================================================================
  150 posts · 289 comments · 2026-01-01 → 2026-06-28
  sentiment  51% positive · 7% negative · ratio 7.26

  Most-asked  (58 of 150 posts are questions — 39%)
      22c   12p  Should I return my Acme Audio?
      21c   41p  Anyone using Zenith for long flights?
      21c   24p  Acme Audio vs Northwind — which holds up better?

  Share of voice
    Acme Audio                51   27.3%  ██████████████████
    Northwind                 48   25.7%  █████████████████·
    Kestrel                   44   23.5%  ████████████████··

  Complaint fingerprint  (100 = same as peer brands; >130 distinctive)
    baseline: 150 brand-mentioning documents
    Zenith               n=44     19 pos / 5   neg    sharpest: customer service 320
        customer service         320  █████████████████████▶
        returns & refunds        146  ████████████████······
    Northwind            n=48     14 pos / 8   neg    sharpest: defective on arrival 278
        defective on arrival     278  █████████████████████▶
        quality / durability     238  █████████████████████▶

Install

No package to install and nothing to build. You need Python 3.8+ and the single reddit_voc.py file.

# Option A: clone the repo
git clone https://github.com/seoprocheck/reddit-voc.git
cd reddit-voc

# Option B: grab just the one file
curl -O https://raw.githubusercontent.com/seoprocheck/reddit-voc/main/reddit_voc.py

Confirm it runs (no network needed, exercises the analysis internals):

python3 reddit_voc.py --selftest        # expect: 15/15 passed

Then try it on the bundled fixture before hitting the live archive:

python3 reddit_voc.py --fixture fixtures/sample-corpus.json \
  --brands "Acme Audio,Zenith,Northwind,Kestrel" --min-docs 25

Usage

python3 reddit_voc.py SEO --days 30                       # what a community talks about
python3 reddit_voc.py SEO --days 90 --query "core update" # narrow to a topic
python3 reddit_voc.py headphones --brands "sony,bose,sennheiser,anker"
python3 reddit_voc.py headphones --brands "sony,bose" --themes --json
python3 reddit_voc.py headphones --brands "sony,bose" --verbatims       # real quotes, not just numbers

A fingerprint tells you that a brand over-indexes on "battery life"; --verbatims shows you the actual comments behind it, top-voted, split into what people praise and what they complain about, plus one real example per complaint axis. The sentiment tag is decided on the same sentence that gets quoted, so the text and its label can never disagree. It is still a lexical read — sarcasm and negation slip through — so treat every quote as a thread worth opening, not a line to paste into a slide. --quotes N sets how many per brand per tone (default 3).

Measure a brand community against neutral ground:

python3 reddit_voc.py acmeaudio --brand acme --vs headphones

Pull once, analyse many times — the archive is free and run by volunteers, so don't re-fetch what you already have:

python3 reddit_voc.py headphones --days 180 --save headphones.json
python3 reddit_voc.py --fixture headphones.json --brands "sony,bose" --min-docs 50

Try it instantly with the bundled fixture — fictional brands, generated text, no real Reddit content:

python3 reddit_voc.py --fixture fixtures/sample-corpus.json \
  --brands "Acme Audio,Zenith,Northwind,Kestrel" --min-docs 25

Verify the analysis internals — word-boundary matching, ratio smoothing, peer baselining, and the invariant that a corpus compared against itself must return a gap of exactly 1.0:

python3 reddit_voc.py --selftest

Your vertical, your complaints

The default complaint axes are generic — price, durability, service, returns, warranty, delivery, setup, sizing. Swap them for your category with a JSON file of {label: regex}:

{
  "battery life":   "battery (life|drain)|dies after|won't hold a charge",
  "comfort":        "clamp(ing)? force|hurts my ears|too tight|sore after",
  "bluetooth drop": "drops? (out|connection)|keeps disconnecting|pairing"
}
python3 reddit_voc.py headphones --brands "sony,bose" --issues audio-issues.json

Reading the output honestly

Sentiment here is lexical, not classified. Documents are matched against positive and negative marker patterns. That is robust enough for comparison — brand A versus brand B, own sub versus neutral — and not robust enough to quote as "49% of customers are happy". It cannot read sarcasm or negation. Use it to rank, not to score.

The ratio is smoothed. (positive+1)/(negative+1), not positive/negative. Twenty-two positives and zero negatives is not a ratio of 22; it's a ratio you don't have the data to state.

Gaps get a robustness check. Whether the brand's own side is filtered to documents naming the brand, and whether the ratio is smoothed, are both arbitrary calls that move the number. The tool re-runs the gap under all four combinations and tells you the range. If the variants disagree, it says so and tells you to compare brands by rank instead of by multiple — a finding that only holds under one arbitrary choice is not a finding.

Small samples are flagged, not hidden. Under 50 documents on either side and you get a warning; brands under --min-docs mentions are listed as skipped rather than silently dropped.

What you cannot see. Removed and deleted content is gone. Moderation strips complaints as well as spam, in unknown proportion — one subreddit tested deletes roughly three quarters of submissions. Treat every rate here as a rate among surviving posts.

Check the subreddit is what you think it is. Brand subreddits are far rarer than people assume, and brand names collide. r/Corona is a city in California, not the drink; r/Java is a programming language, not a coffee. Most major consumer brands have no subreddit at all, and the general category community is where the conversation actually happens.

Notes

Arctic Shift rate-limits with HTTP 422 and the message "slow down a bit" rather than a 429; the client backs off and retries automatically. Requests are spaced by default. Full-text --query works on posts only, which is an upstream constraint.

License

MIT © SEO Pro Check · built by @seoprocheck.

About

Reddit voice-of-customer mining — brand share-of-voice, peer-baselined complaint fingerprints, and the brand-sub vs neutral-ground sentiment gap. No API key. Zero deps.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages