An open-source AI outcome analysis system that finds fake winners, real winners, and conversion leaks, even when your tracking is messy.
Built for OpenClaw and Claude Cowork style agent workflows.
Ads + Pages + Outcome Signals → Real Winners → Fake Winners → Leaks → Next Moves
Outcome Kit reads your ad data, landing page data, and whatever business outcome signal you actually have, leads, bookings, signups, activations, purchases, revenue.
Then it tells you:
- What’s actually working — the angles, creatives, and pages driving outcomes
- What only looks like it’s working — high CTR, cheap leads, weak downstream quality
- Where the leak is — ad, page, or follow-through
- What to do next — scale, cut, rewrite, rebuild, or test
The result: a daily operator brief that gives you decision-grade truth, not dashboard theater.
Most teams don’t have a data problem.
They have a fake winner problem.
The ad with the best CTR often brings the worst buyers. The cheapest lead often books the fewest calls. The campaign report that looks great in-platform often hides the weakest downstream outcomes.
And most businesses do not have perfect attribution. No clean CRM. Broken UTMs. Partial tracking. Calendly bookings in one place, ad spend in another, landing page data somewhere else.
That’s normal.
Outcome Kit exists because you still need to make decisions anyway.
This system doesn’t wait for perfect infrastructure. It uses the signals you do have and turns them into a clear read on:
- real winners
- fake winners
- leaks
- underfed opportunities
So you can stop trusting vanity metrics and start acting on what’s actually producing outcomes.
7:08am. Telegram from the agent:
OUTCOME KIT — DAILY BRIEF
Primary outcome: Demo Bookings
Secondary signal: Show Rate
Confidence: MEDIUM
REAL WINNER
• "Stop guessing" angle
- 18% lower cost per booking
- 24% higher show rate
- strongest blended outcome score
FAKE WINNER
• "Save time" creative family
- highest CTR in Meta
- weakest booking rate from page visit
- likely curiosity traffic
LEAK
• Founder-led psychology ads → generic SaaS page
- ads pull strong clicks
- page converts 34% below account average
- message/page mismatch
UNDERFED WINNER
• Competitive teardown angle
- low spend
- highest booking efficiency
- only 9% of budget
NEXT 3 MOVES
1. Shift 15% budget from "Save time" to teardown angle
2. Build a new page for founder-led psychology traffic
3. Generate 4 new creatives from the "Stop guessing" family
That’s the product.
Not another dashboard. Not 14 tabs open across Meta, GA4, and your booking tool. Not "it depends."
Just:
- what’s real
- what’s fake
- what’s leaking
- what to do next
Outcome Kit is for operators with enough data to be confused, but not enough infrastructure to be certain.
That usually means:
- SaaS founders
- agencies
- coaches / consultants
- lean in-house marketing teams
- service businesses
- e-commerce brands with partial post-click tracking
If you’re spending money on traffic and trying to figure out which message actually produces the outcome you care about, this is for you.
You choose the business outcome that matters most.
- leads
- bookings
- qualified leads
- signups
- trial starts
- activated users
- purchases
- revenue
- show rate
- activation rate
- purchase rate
- close rate
- AOV
- LTV proxy
- qualification score
The more downstream data you connect, the smarter the agent gets.
But you do not need a perfect CRM to get value from this.
Every run classifies your funnel into four buckets:
Strong upstream response and strong downstream outcome quality.
Looks good in-platform, weak on actual outcomes.
The interest is real, but something after the click is breaking.
Low volume, strong efficiency, deserves more budget or more testing.
This framework is the core of the kit.
INGEST → MAP → SCORE → DIAGNOSE → RECOMMEND → LEARN
Pull data from:
- Meta Ads
- Google Ads (optional)
- landing pages / site analytics
- booking or signup tools
- CRM or revenue systems if available
Normalize everything into a shared structure:
- angle
- asset
- page
- audience
- source
- outcome event
Evaluate each angle, creative family, and page against:
- cost per outcome
- conversion efficiency
- quality signal
- data confidence
Find:
- real winners
- fake winners
- leaks
- underfed opportunities
Turn that into action:
- scale
- cut
- rewrite
- rebuild
- clone
- test
Store patterns over time:
- which hooks drive low-intent traffic
- which pages break conversion
- which creative styles produce better outcomes
- which angles are over-credited upstream
This is not a snapshot. It compounds.
Most tools show you numbers.
Outcome Kit makes decisions.
Most tools assume:
- pristine UTMs
- perfect CRM hygiene
- fully connected revenue data
Outcome Kit assumes reality:
- messy tracking
- incomplete downstream data
- disconnected tools
- still needing to make better calls anyway
It doesn’t pretend to know more than it knows. Every recommendation is confidence-aware.
That means the agent can say:
- High confidence: this angle is the strongest by booked-call efficiency
- Medium confidence: this creative family is likely bringing low-intent traffic
- Low confidence: revenue linkage is weak, recommendation withheld
That honesty is the point.
| Skill | What It Does |
|---|---|
meta-source-reader |
Pulls Meta campaign, ad set, and ad performance data |
page-signal-reader |
Reads landing page sessions, CVR, and engagement data |
outcome-event-reader |
Pulls leads, bookings, signups, purchases, or other outcome events |
metadata-loader |
Loads angle tags, creative tags, audience tags, and page mappings |
angle-mapper |
Groups ads, pages, and outcomes into message families |
outcome-scorer |
Builds blended scoring around your chosen outcome |
fake-winner-detector |
Finds assets that look good upstream but fail downstream |
leak-diagnoser |
Identifies whether the break is in the ad, page, or follow-through |
opportunity-finder |
Finds low-volume, high-quality angles worth scaling |
decision-writer |
Converts analysis into concrete operator moves |
brief-sender |
Delivers daily or weekly briefs to Telegram, Slack, or WhatsApp |
pattern-memory |
Stores recurring patterns and learns over time |
Each skill can run standalone or as part of the full pipeline.
Most marketing reports are built around:
- campaigns
- ad sets
- ads
- channels
That’s not enough.
Outcome Kit is built around angles.
Because “Ad 27” is not the strategic question.
The real question is:
- Is the time-savings angle working?
- Is founder-led proof driving better bookings?
- Is the competitive teardown message producing stronger buyers?
- Is the page failing the message?
That’s the level that actually matters.
- Find which ad angle produces the best cost per booking
- Spot high-CTR ads that bring weak show rates
- Identify page-message mismatch on demo pages
- Compare cheap lead angles vs qualified lead angles
- Find the creative family that drives the best booked calls
- Detect landing pages suppressing form completion
- Find which campaigns bring signups that actually activate
- Catch curiosity hooks that inflate top-of-funnel volume
- Recommend more creatives around the best activation-producing message
- Find which creative angles drive actual purchases, not just clicks
- Compare page variants by conversion efficiency
- Spot underfed winners worth scaling
OUTCOME KIT — DAILY BRIEF
PRIMARY OUTCOME: Activated Signups
SECONDARY SIGNAL: Purchase Conversion
CONFIDENCE: HIGH
REAL WINNER
• "Stop guessing" angle
- lowest cost per activated signup
- strongest downstream purchase rate
- stable across 14-day and 30-day windows
FAKE WINNER
• "Save hours" angle
- best CTR
- weakest activation rate
- high curiosity, low buyer intent
LEAK
• Founder-led ads → generic product page
- strong clickthrough
- weak signup CVR
- message-page mismatch
UNDERFED WINNER
• Competitive teardown angle
- low spend
- high activation efficiency
- deserves more budget
NEXT 3 MOVES
1. Shift 20% spend from "Save hours" to teardown angle
2. Create a new page matching founder-led proof messaging
3. Generate 3 new creatives from the "Stop guessing" family
This kit is intentionally focused.
- Meta as the primary paid source
- landing page / site signal ingestion
- one outcome source
- angle tagging
- confidence-aware scoring
- daily or weekly operator brief
- persistent learning / memory
- full multi-touch attribution
- perfect revenue reconciliation
- automatic budget changes
- enterprise dashboards
- every ad platform on earth
The point of V1 is not “complete marketing measurement.” It’s better decisions, fast.
This repo now supports both:
- OpenClaw for full agent runtime, memory, cron, and chat workflows
- Claude Cowork for teams who prefer Claude's desktop agent with sub-agent coordination
The underlying pipeline is the same. The difference is the agent shell around it.
# Clone the repo
git clone https://github.com/TheMattBerman/outcome-kit.git
cd outcome-kit
cp .env.example .env
cp config.example.json config.json
# Add your keys, data sources, and primary outcomenpm run doctornpm run run:samplenpm run runOpenClaw is the best fit if you want:
- Telegram / chat-based operation
- cron jobs and recurring briefs
- memory and longer-lived agent behavior
- tool routing across a bigger agent system
npm install -g openclaw
openclaw startThen message it naturally:
- “Analyze which ad angles are driving booked calls”
- “Find fake winners in my Meta funnel”
- “What’s leaking between the ads and the page?”
- “Which angles deserve more budget?”
- “Give me the daily outcome brief”
Claude Cowork is a strong fit if you want:
- desktop-native agent workflows with sub-agent coordination
- agent instructions that live in
CLAUDE.mdand.claude/ - explicit skill, rule, and hook structure via plugins
This repo includes:
CLAUDE.md.claude/agents/.claude/skills/.claude/rules/.claude/hooks/
That means Claude Cowork or Claude Code can understand:
- pipeline order
- what each stage does
- which artifacts matter
- what good diagnosis looks like
- how confidence should constrain recommendations
The repo now includes:
- pipeline-order hook checks so later stages do not run on missing or invalid artifacts
- confidence-aware diagnosis rules
- per-angle attribution quality surfaced in briefs
- validate-first suppression when evidence or attribution quality is too weak to justify action
If you use Claude Cowork, this gives the agents much better rails than a generic repo prompt.
{
"primary_outcome": "bookings",
"secondary_signal": "show_rate",
"sources": {
"meta_ads": true,
"google_ads": false,
"page_signals": true,
"crm": false,
"booking_tool": true
},
"confidence_rules": {
"min_sample_size": 3,
"allow_medium_confidence_recommendations": true,
"withhold_low_confidence_actions": true
},
"source_enrichment": {
"rules": [
{
"match": "competitor breakdown",
"angle": "competitive-teardown",
"page": "/teardown"
}
]
},
"angles": [
{
"name": "time-savings",
"audience": "heads-of-growth",
"creative_family": "founder-direct",
"page": "/time",
"match_tokens": ["save time", "time savings", "hours back"]
}
]
}outcome-kit/
├── README.md
├── SETUP.md
├── SPEC.md
├── AGENTS.md
├── SOUL.md
├── .env.example
├── config.example.json
├── package.json
├── scripts/
│ ├── lib.js
│ ├── run.js
│ ├── meta-source-reader.js
│ ├── page-signal-reader.js
│ ├── outcome-event-reader.js
│ ├── map-angles.js
│ ├── score.js
│ ├── diagnose.js
│ ├── recommendations.js
│ ├── write-memory.js
│ ├── brief.js
│ └── validate-config.js
├── CLAUDE.md
├── .claude/
│ ├── agents/
│ │ ├── data-reader.md
│ │ ├── diagnostician.md
│ │ └── brief-writer.md
│ ├── skills/
│ │ ├── meta-source-reader/
│ │ ├── page-signal-reader/
│ │ ├── outcome-event-reader/
│ │ ├── angle-mapper/
│ │ ├── outcome-scorer/
│ │ ├── fake-winner-detector/
│ │ ├── leak-diagnoser/
│ │ ├── opportunity-finder/
│ │ ├── decision-writer/
│ │ ├── brief-sender/
│ │ ├── pattern-memory/
│ │ └── run-analysis/
│ ├── rules/
│ │ ├── data-quality.md
│ │ ├── scoring.md
│ │ └── diagnosis.md
│ └── hooks/
│ └── check-pipeline-order.sh
└── workspace/
└── outcomes/
By default, AI and dashboards both tend to do dumb shit here.
Outcome Kit is built specifically to avoid:
- declaring winners from CTR alone
- treating cheap leads as high-quality outcomes
- hiding broken tracking instead of surfacing it
- overclaiming certainty when confidence is low
- flattening different message angles into one blended average
- reporting numbers without decisions
- optimizing for platform metrics that don’t match business outcomes
If the data is messy, the agent says so. If confidence is low, it says so. If the platform winner is actually garbage, it says that too.
Built by Matt Berman.
- 🐦 Twitter/X: @themattberman
- 📰 Newsletter: Big Players
- 🏢 Agency: Emerald Digital
This is for operators who need truth before they need another chart.
Star the repo if this helps. It tells me to keep building.