Skip to content

Repository files navigation

Open Loops

Finds the things nobody is chasing — and tells an executive's assistant which is closest to falling over.

An executive assistant's real job isn't scheduling — it's being the safety net. Someone promised something three weeks ago in a thread nobody reopened, and there is no system anywhere that knows about it. Open Loops reads Slack and the calendar, extracts the promises, works out which ones closed on their own, and ranks what's left by how close it is to falling over.

Use it with Claude or with ChatGPT through the Codex app. Both run the same local detector and deliver the results to Slack.

That is the whole product: a skill. There is no app and nothing to log into. It reads your channels and your calendar, posts a ranked list to your own Slack DM each evening, and you correct it by replying to the message. The output is text, and it looks like What comes out below.


What it reads, and where that goes

Anyone pointing this at a workspace they do not personally own is right to ask, and it should be answerable without reading the source.

The detector runs locally. There are no network calls in src/ or slack-run.js. The one file that makes any is tools/report.js, and only if you opt in to diagnostic reports after your first digest: a fixed set of fields — a random ID, versions, the date, and where a run failed or which kind of item you corrected — never message text, names, paths or error messages. "diagnostics": false in the config stops it and discards anything unsent. Claude or Codex retrieves messages through your connected Slack and calendar tools and posts the digest to your own DM. Those connector responses are also processed by the host assistant; local detection does not mean the data never leaves your machine.

Midday alerts are off unless you opt in. Saying yes adds a second scheduled task; the daily digest task is not touched. At 12:00 and 15:00 on weekdays, a check re-reads the same channels and calendar as the digest and posts one short alert to your own DM, only if a commitment became due or overdue or a new urgent one appeared. It uses the digest's own status fields, alerts once per level, never writes the ledger, and leaves the evening digest and its reply numbering as they are. It sends no notification of its own: Slack may stay silent for a message you post yourself, and the Claude app may show a routine "task completed" notice after every check. See SKILL.md, "Offering midday alerts".

You choose what it reads. channels.include is an allowlist: name two channels and it reads two channels. Direct messages are never included unless you add them by name, because they are the most sensitive thing in a workspace and the least likely to hold a commitment anyone is tracking.

What lands on disk is ledger.json, in your working directory. By default it keeps the sentence each commitment came from, so the digest can say what cleared. "storeText": false keeps the tracking — keys, dates, verdicts, accuracy — and removes stored row text, learned phrases and the rerun snapshot's text on the next real run. It does not erase raw input captures or separate backups. Every open item is re-detected from live messages on each run and still quotes its sentence in full. The one thing lost is cleared since the last run, which reads from the ledger and falls back to (text not kept).

No model decides anything. On the Slack path a model fetches and posts; what counts as a commitment is regexes and plain comparisons. Nothing is sent anywhere to be classified.

What it trusts. Everyone who can post in a channel you read. Message text is parsed from what the connector returns, and text alone cannot tell a line that looks like a message header inside somebody's message from a real one, so a workspace member could write a message that reads as a message from someone else, including from you. A read whose timestamps repeat or run out of order is flagged in the digest (ORDER), which catches a careless forgery and not a careful one. Point it at workspaces where that member is not an adversary, and keep the channel allowlist short. Structured message records would remove this; the Claude connector returns text only.

Stopping and removing it are different things:

  • Stop it: delete or pause the scheduled task. The ledger stays, so picking it up later resumes rather than restarts.
  • Remove the skill: npx skills remove open-loops for Claude, or delete the open-loops folder under ~/.codex/skills for Codex.
  • Delete what is on your machine: the working directory (~/open-loops-data), which holds ledger.json, each run's input and output files, and the code checkout.
  • What deleting does not reach: the digests already posted to your Slack DM (delete them there), the assistant's conversation history for past runs, and whatever the connector providers keep.

Problems and questions: open an issue at github.com/kaarizhussain/open-loops/issues. This is a working prototype for one workspace, not a supported product.


What it finds

Seven ways a commitment slips, all detected from message text — none of it tagged by hand.

Five of them fire on nothing having happened. That is the whole point, and it is worth putting first: an assistant already knows about the loud problems. What they need is the thread that has been silent for three weeks, and silence is the one thing a summariser cannot surface, because there is nothing there to summarise.

Signal What fires it
Unanswered An inbound request, and no reply from your side absence
No reply yet You asked a question; silence absence
No follow-up sent An external meeting ended and nothing went out after absence
Meeting unprepped External attendees, no agenda attached absence · two sources
Agreed, not booked A call was agreed to, but no calendar hold exists for it absence · two sources
You promised An outbound commitment with no matching delivery since a sentence exists
They promised An inbound commitment and nothing has arrived a sentence exists

The last two are the ones anything can do. Point a good summariser at the same mailbox and it will find sentences that sound like promises too. The five above it cannot: they are all statements about what is missing, and two of them exist only in the gap between a mailbox and a calendar — the seam no single-product assistant reaches across.

What comes out

From 25 messages and 5 meetings, the brief — the message itself:

OPEN LOOPS — for 2026-08-06 · Thu
4 overdue · 2 due today · 2 due by Fri · 15 open

TODAY — highest priority
 1  1d late    Answer Program — "Can you confirm your final talk title and a short bio by…"
               Conference · RevOps Summit — speaker confirmation · Sat
 2  3d late    Chase j.mercer — "We'll send the countersigned copy back Monday."
               Partner · Solstice partner agreement · 27 Jul
 3  today      Send an agenda — Larkspur Retail — renewal decision
               m.osei@larkspurretail.com · Key account · Larkspur Retail — renewal decision · Thu

NEEDS TO GO OUT TODAY
  [LATE] Larkspur Retail — renewal decision — today, no agenda (m.osei@larkspurretail.com)
  Vector Freight — Q3 QBR — tomorrow, no agenda (lena.borg@vectorfreight.com)

ALSO OPEN — most pressing first
 4  today      Needs the executive: "Yes — I'll send it Thursday EOD."
 5  Fri        No agenda attached — Vector Freight — Q3 QBR
 6  Fri        You owe Rachel: "I'll book the venue by end of week — let me look at…"
 7  Sat        Needs the executive: "I'll review and get back to you before the deadline."
 8  asked Fri  Lena asked: "Could you send an agenda ahead of the QBR so our…"
    + 7 more, 3 closed and today's spot check in the thread ↓

Reply  3 7 not real · k 1 4 already knew · miss b answers the spot check

Two messages, because what the detector needs to remember and what a person needs to read are different lists. The brief says what to do: three items with their evidence, a few one-liners, and the way down. The details, posted as a reply in its thread, say why: every item grouped by who acts next — what only the executive can do, what someone else owes you, what you can close out yourself — with the sentence it came from, what closed and what closed it, what was read, any warnings, and the spot check.

TODAY is chosen, not sorted by one score: a stated deadline that has just arrived, then what lands in the next three days, then someone waiting on your answer, then everything else overdue — which still leads the one-liners, so an old promise is never buried. It opens with what to do, not with how many messages it read; a digest that leads with its own statistics is a system reporting on itself.

Every item quotes the sentence it came from and says who said it, where and when — and when a deadline was borrowed from another message, whose it was. What closed itself says what closed it. That is what a reader checks the list against, in seconds.

The same fifteen items every morning is a list nobody reads by Thursday — not because it is inaccurate, but because it is identical. So what changed rides on the counts line, on a day with old and new items mixed each item says NEW or how long it has sat, and anything that dropped off is reported once as cleared.

Seeing it without installing it

The text above is the product. This is not — it is a browser page that runs the same detector over an invented mailbox, so the reasoning can be poked at without connecting anything. Nobody who installs the skill sees this screen; it exists because a ranked list is easier to argue with when you can click a row and read why it fired.

The demo page — the same detector over an invented mailbox

▶ Try it — a synthetic CRO's Thursday. Click any row for the sentence that triggered it, the rule that fired, and a drafted chase note. Every name in it is fictional.

How it works

messages + events
   → commitment extraction     cue phrases, per sentence
   → deadline resolution       "Thursday EOD" → 2026-08-06, relative to send date
   → closure matching          did a later message deliver it, and was it that one?
   → silence + meeting checks  unanswered asks, unbooked calls, missing agendas
   → risk ranking              overdue days, proximity, signal type, relationship
   → the ledger                what is new, what cleared, what you already rejected

The detector is deterministic and dependency-free. It takes a normalized shape, so the source is an adapter concern, not an engine concern:

detectLoops(messages, events, { exec: 'dana@northstar.io', today: '2026-08-06' })

The Slack runner supports both Claude and Codex, accepting verbatim connector text or structured Slack records. Both use the same detector, ledger and digest renderer. Pointing it at something else below describes the normalized input needed for another source.

Deterministic is load-bearing, not a preference. No model decides what counts as a commitment: the same messages, settings and evaluation date produce the same detections. Classification happens in local code. On the Slack path the host assistant fetches and posts; it never judges.

Who you support

The list of people you support is optional, and empty means nobody — you are reading your own work, which is the common case for anyone trying this on themselves.

principals: []                          // nobody. Two piles: chase them, and yours.
principals: [{ label: 'Dana' }]         // one. Her pile is "Needs Dana".
principals: [{ label: 'Dana',   address: 'dana@northstar.io' },
             { label: 'Marcus', address: 'marcus@northstar.io' }]

That is what the detector's options and a run input call the list. In openloops.config.json it is "supporting" — "supporting": [{ "label": "Dana" }] — and a principals key there does nothing.

Your own "I'll…" is always yours. A principal's pile holds what was promised in their name — "Dana will send the signed copy" — and with several, the name in the promise decides whose, and goes on each line. Anyone else's promise is theirs to chase. Leaving it out entirely keeps the original behaviour: one unnamed executive reading their own inbox, where "I" is the executive.

How you argue with it

Marking something wrong has to cost about as much as ignoring it, or it does not happen for fourteen days running — which is exactly how long it takes to learn anything. So the digest numbers its items and you reply to it. Nothing to open, nothing to log into.

3 7            those two are not real commitments
k 1 4          those are real, but I already knew
miss b         the spot check found something it walked past

Three replies, three different numbers, and they measure different things:

Precision — of what it showed you, how much was real. Whether it can be trusted.

Novelty — of what was real, how much you did not already know. Whether it is worth reading. Someone with a good memory could get a flawless digest every morning and gain nothing from it, and precision alone would call that a success.

Recall — how much it walked straight past. Nothing else in the loop can see this: a miss produces nothing to reject, so corrections could only ever teach it to be quieter, never more thorough. Each digest therefore samples its own silence — a handful of messages it found nothing in — and asks whether it should have. The two worst bugs found in this detector were both false negatives, invisible to every other number here.

                       ── ILLUSTRATIVE. Not this project's results. ──
Tracked 63 items.
   9 wrong        → 86% held up
  41 already known
  13 genuinely new → 21% told you something
Recall so far: about 84% — 3 misses found in 45 messages spot-checked.

Invented numbers, showing the shape of the report. The real ones are in Evaluation and they are thinner: five messages spot-checked, not forty-five. A reviewer read this block as a result and congratulated the project on it, which is a fair warning about how a sample renders inside a document that is otherwise trying hard to state what it does not know.

Past four rejections of the same phrase, with none of them kept, it stops asking and mutes it — announcing the change, listing it in every digest afterwards, and undoing it on one line. It learns the kind, not just the instance, because otherwise the same bad pattern arrives fresh every morning forever.

Nothing it learns is hidden. A detector that rewrites its own rules invisibly is one nobody can predict, and predictability is most of the reason this is regexes instead of a model.

The parts that were actually hard

Every real bug in this thing was found by running it against real messages, never by reading the code. Five separate times, on five different days. A closure rule that let the word signed mark a future promise as already delivered. Deadlines borrowed from unrelated messages further up a channel. An app footer poisoning the topic match. A question buried by the sender's own later message, so it vanished instead of ageing. And the biggest: every rule assumed a channel was one conversation with one person. The first run with three people in it showed whoever posted last "answering" every open question, and every promise pinned on whoever had spoken first.

None were visible in review, and every one looked obvious afterwards.

→ What was actually hard — the full list, and what each broke.

Quick start

ChatGPT integration via the Codex app

This integration runs in the Codex desktop app, using its Slack connection and local Node.js execution. It is not a hosted ChatGPT web app or a custom GPT. You need Git, Node.js, and a Slack connection in Codex with channel history, thread reads, user profiles and message posting. No OpenAI API key or additional Node dependencies are required.

1. Clone the repository and install the skill.

git clone https://github.com/kaarizhussain/open-loops.git
cd open-loops
node tools/install-codex.js

If you already have a checkout, run the installer there. It installs the skill in your personal Codex skills directory and records the checkout's location, so keep that folder in place. Restart Codex if the skill does not appear.

To pick up a new version, pull the checkout and run node tools/install-codex.js --update. It replaces the installed copy as a whole, so the instructions and the runner stay the same version, and keeps the old copy beside it as open-loops.bak-<time>.

2. Connect Slack in Codex. Claude's Slack authorization does not carry over. Google Calendar is optional and needs its own connection for meeting coverage.

3. Ask Codex to set it up.

Use $open-loops to set up my Slack digest.

Codex shows your work channels and asks which to track, then asks whether you track your own work or support someone else. You can include that information in your first request. Setup saves your choices in openloops.config.json and explains where fetched text is processed and saved before reading message history.

Codex runs the first digest, posts the brief to your self-DM with details in its thread, and echoes the verified brief in chat. Reply in Slack with 3 7 to reject items or k 1 4 for items you already knew. If a spot check is included, miss b answers it. Calendar connection is optional and does not block the first Slack digest.

4. Schedule it after the first successful run. Ask Codex to run it daily at your preferred local time. Installation alone does not create a schedule. Local scheduled runs need the computer awake and the app running.

The optional diagnostics question comes last, after the first digest and your schedule choice. Midday checks currently have a Claude workflow only; Codex does not offer or enable them.

Alongside Claude: use separate data directories and ledgers while comparing the two hosts. Stop the old schedule before migrating to a shared ledger; two schedulers must not write to it concurrently.

Tested in a Slack sandbox: channel and thread reads, EA/executive ownership, digest delivery, completion matching, and a numbered correction processed in a simulated next-day run. Live calendar access and unattended Codex scheduling have not yet been validated.

See Codex setup for prerequisites, data formats, and running alongside Claude.

Claude setup

npx skills add kaarizhussain/open-loops

Ask Claude to set up Open Loops with its connected Slack tools. See the Claude Slack workflow for the setup and daily run instructions.

Run the code and demo

No dependencies, no install, no API keys.

git clone https://github.com/kaarizhussain/open-loops.git
cd open-loops
npm test              # every suite
node build.js         # rebuilds index.html (downloads fonts on first run)

index.html is committed, so you can also just open it in a browser. npm test fails if it has drifted behind src/, because a demo that silently ships an old detector is worse than no demo — that happened, and it is the page most people see.

The screenshot at the top is the same page, captured headless. It goes stale the same way and nothing checks it, so re-run this when the layout changes:

chrome --headless --hide-scrollbars --force-device-scale-factor=2 \
  --window-size=1440,900 --screenshot=docs/screenshot.png \
  "file://$PWD/index.html#present"

#present suppresses the guided tour, which otherwise covers the app it is touring.

The tests live in test/ and are the documentation for how each part is meant to fail: test.js for the detector against the demo fixture, test_slack.js and test_store.js for the adapters, test_ledger.js for what the digest remembers between runs, test_digest.js for the rendering, test_slack_run.js for the whole Slack path end to end, test_channel.js for a channel with several people in it, and test_replay.js and test_replay_seed.js for real connector output, sanitized and replayed.

Pointing it at something else

The engine never touches a mail or chat API. Write an adapter producing these two shapes and nothing in src/loops.js changes:

message = { id, threadId, subject, from, to: [], date: 'YYYY-MM-DDTHH:MM', body, attach: bool }
event   = { id, title, start: 'YYYY-MM-DDTHH:MM', attendees: [], agenda: bool, series }

Direction is derived from from against the reader's own address, so an adapter never labels inbound or outbound. Two exist: src/slack.js, which ships, and tools/enron.js, a benchmark harness of about a hundred lines.

→ Writing an adapter — the field-by-field mapping, and the privacy limits that matter more than the shapes do.

Evaluation

What it ran on Size Labelled? What it establishes
Demo fixture (src/fixture.js) 25 messages, 5 events yes, by assertion that a change has not broken known behaviour
A live Slack workspace 17 messages, one member 1 rejection, 1 spot check that the whole path runs unattended
The same workspace, three people (test/replay/2026-09-14/) 31 messages, two test accounts yes, 13 scripted items that attribution holds with more than one person — 10 of 13 right before the fix, 13 after
Enron corpus (tools/benchmark.js) 3,725 emails, 16 mailboxes no how often it fires — 35.8 items per 100
Top-of-digest, hand-graded 79 items, two labellers yes that the task is well-posed — kappa 0.76

Read the third row carefully. 35.8 per 100 is a firing rate, not an accuracy. That corpus has no ground truth, so nothing in it says how many of those were real. It is an honest answer to "how noisy is this" and no answer at all to "how right is it".

Precision on real correspondence is still unmeasured. Nobody has used this for a fortnight and marked what it got wrong. Until somebody has, every claim here is about the code and none of it is about results. The machinery to measure that is built and has five data points in it.

Three of the seven signals have never fired on anything real. Two need a calendar the Slack path only just got. No follow-up sent needs a recipient list that chat does not carry, and now says so rather than firing blind.

Everything in src/fixture.js is invented. Dana Whitfield, Northstar Systems and every counterparty in it are fictional, written to exercise each detector including the cases that must not fire.

→ Evaluation in full — the corpus work, the seven hypotheses that died against it, and why a benchmark can show a filter is consistent but never that its threshold is right.

License

MIT — see LICENSE. Embedded fonts are SIL OFL 1.1; see fonts/NOTICE.md.

About

Finds the commitments nobody is chasing in Slack and the calendar, and tells an assistant which is closest to falling over.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages