Skip to content

Scan for uncovered placeholder name families (#43) - #52

Merged
PiratesIRC merged 3 commits into
mainfrom
feat/issue-43-placeholder-scan
Sep 6, 2026
Merged

PiratesIRC merged 3 commits into
mainfrom
feat/issue-43-placeholder-scan

Conversation

@PiratesIRC

Copy link
Copy Markdown
Owner

Implements #43, a scan that finds the numbered stream-name families that no Placeholder Name Pattern covers.

What it does

A new action, Scan for Placeholder Patterns. It groups every stream name into a family by replacing digit runs with a slot, so MAX 100 and MAX 101 become the one family MAX #, and reports the families that no configured pattern covers, each with an anchored regular expression to paste. It reads one database column that matching already loads, opens no provider connection, changes no setting and writes nothing to the database. The full readout goes to /config/stream-mapparr/placeholder-name-scan.txt, since a notification shows only about 280 characters.

The grouping lives in a new file, Stream-Mapparr/placeholder_scan.py, which imports nothing but the standard library, so it is tested directly. The action is the wrapper that loads the streams and writes the readout.

What was measured rather than assumed

Everything below was measured against 25,068 live stream names and the live guide database.

A digit followed by K is left alone, because 4K and 8K are resolution tags. With that rule off, 15 further templates covering 418 streams group by their resolution tag instead of by a slot number.

A family needs three different numbers in its slot, not three streams. A separate stream-count threshold was written and removed, because three distinct numbers already implies three streams, so it could never refuse anything the first rule admitted.

The readout is plain ASCII with any other character written as a backslash-u escape, the same rule as the CSV export preamble. Provider names here carry such characters. Python regular expressions accept that form, so every suggested pattern still matches the name it came from.

The ranking was reversed after review

The first version ranked families by how many of their streams carried an EPG identifier, and suggested a pattern only for those that did. Four independent reviews and a database measurement showed that was wrong in both directions.

A stream with no EPG identifier returns from the resolver immediately, so a pattern over such a family cannot change matching at all. A stream that resolves has its matching name replaced by the programme currently airing. The old ranking therefore promoted the families where a pattern changes behaviour and demoted the ones where it is inert.

Measured: 3,700 streams carry an EPG identifier, only 218 of those identifiers match a guide row, and 28 streams could resolve a programme at the moment of measurement. Of the 16 families the old ranking promoted, one held a stream whose identifier matched a guide row and none could resolve a programme. That one was a numbered channel lineup, the case a pattern harms. The largest family on the installation, at 273 streams, was reduced to one line with no pattern.

Families are now ranked by size, which is what the reporter of the issue did by hand. Every uncovered family gets a pattern. The EPG count stays as a note on each family, and a family where at least half its streams carry an identifier is marked with a caution that a pattern there replaces a working name. On live data that caution fires on 13 families, every one a recognisable channel lineup, and they moved from the top of the list to ranks 31 to 60.

What this does not solve

Size alone does not separate event placeholder families from informative numbered channel families either. The six largest uncovered families include three of each kind. The readout says so in its own text rather than implying the ranking settled it.

Verification

1,510 tests pass. Twenty-nine mutations were applied one at a time and every one was caught by a named test, with a comment-only control leaving the suite green. That set includes restoring the old ranking, which now fails a test that says why it was removed. The publish audit is clean with all seven deny rules proven to fire.

🤖 Generated with Claude Code

https://claude.ai/code/session_017ZBufwmvsXqzkq2h2VkAwB

PiratesIRC and others added 3 commits September 6, 2026 10:51
Adds the action Scan for Placeholder Patterns. It groups every stream name
into a numbered family by replacing its numbers with a slot, and reports the
families that no configured Placeholder Name Pattern covers, each with an
anchored regular expression to paste.

Requested as issue #43. The setting only ever helped with the naming schemes
the operator already thought to write down, and nothing in the interface told
apart "this installation has no placeholder families" from "the patterns you
wrote match none of them".

The grouping lives in a new stdlib-only module, Stream-Mapparr/placeholder_scan.py,
with no Django import, so it is tested directly. The action is the thin wrapper
that loads the streams and writes the readout to
/config/stream-mapparr/placeholder-name-scan.txt.

Measured against 25,068 live stream names, and three decisions came from that
rather than from assumption:

  A digit immediately followed by K is left alone, because 4K and 8K are
  resolution tags. With the rule off, 15 further templates covering 418 streams
  group by their resolution tag instead of by a slot number.

  A family needs three different numbers in its slot, not three streams. A
  separate stream-count threshold was written and removed: three distinct
  numbers already implies three streams, so it could never refuse anything the
  first rule admitted, and no test could show it doing work.

  Families whose streams carry EPG data are reported first and separately. On
  this installation 131 families are uncovered and only 16 hold a stream with an
  EPG identifier, so one undifferentiated list would report a problem eight
  times larger than the one worth acting on.

The readout is plain ASCII with any other character written as a backslash-u
escape, the same rule as the CSV export preamble. Provider names here really do
carry such characters, and an earlier version of the ASCII test used only
synthetic names, so it passed while the live report broke the rule. Python
regular expressions accept that form; all 132 suggestions generated from live
data were checked against their own example name.

Every guard was proven by mutation rather than by reading: each was broken in
turn and a named test confirmed to fail, with a comment-only control leaving the
suite green. That found one vacuous test, which sized its own input from the
threshold it was checking and so moved with it.

Reports only. No setting is changed and nothing is written to the database.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ZBufwmvsXqzkq2h2VkAwB
Two reviews of the first version of the placeholder family scan found four
defects. Each was reproduced before it was changed, and each is now covered by
a named test that fails when the fix is reverted.

A character above the basic multilingual plane, an emoji or a flag, was escaped
with the lower-case backslash-u form. That form carries a MINIMUM of four hex
digits, not exactly four, so an emoji produced five and Python read only the
first four. The pattern compiled, matched nothing, and the family kept
reporting as uncovered with nothing saying why. The existing round-trip test
used a character inside the basic plane, which is the one class the old form
already handled. Characters above the plane now use the upper-case eight-digit
form.

A literal backslash was left unescaped in the template while a literal hash was
escaped, so a backslash followed by a digit produced the same two characters as
an escaped hash. Two different names grouped into one family and the suggested
pattern matched neither.

The scan ran inside the request with no yield, no input cap and no time budget.
The pattern safety gate deliberately admits patterns that can backtrack
polynomially, on the stated promise that the runtime bounds the input, and that
promise was kept only on the regex pre-processing path. The walk now hands the
worker back every 500 names, skips a name over 500 characters and stops at a
five second budget, using the same PluginConfig limits as that path. A stopped
walk is reported as partial rather than passed off as finished: a family the
walk never reached is missing, not covered.

A failure to write the readout returned a plain success. The readout exists
only in that file, so the notification now says the file is missing, and says
it FIRST, because the notification helper drops lines from the end and a test
caught the warning being truncated away.

Also from those reviews: distinct slot numbers are counted within one slot
rather than across slots, so two slots holding two values each no longer count
as four; a resolution tag groups regardless of the case of its K, since
patterns are compiled case-insensitively; a suggestion longer than the setting
will accept is marked as such instead of being offered as though it worked; a
database row that is not a dictionary is skipped rather than raising out of the
walk; and the second listing is capped and says how many it left out.

Verification: 1,506 tests pass. Twenty-four mutations were applied one at a
time and every one was caught by a named test, with a comment-only control
leaving the suite green. One mutation initially escaped, because the test
asserted the argument NAMES appeared in the call and setting them to None keeps
the names; the assertion now pins the values. Re-measured over the same 25,068
live stream names: 132 families, 131 uncovered, 16 carrying an EPG identifier,
0.26 seconds, no name skipped, budget not tripped, output plain ASCII, and all
132 suggestions match their own example name.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ZBufwmvsXqzkq2h2VkAwB
The first version ranked uncovered families by how many of their streams
carried an EPG identifier, and split the readout in two, suggesting a pattern
only for families that carried one. Four independent reviews and a measurement
of the live database agree that was wrong in both directions.

WHICH DIRECTION THE RISK RUNS, read from the code. In
_resolve_current_epg_title_for_stream a stream with no EPG identifier returns
immediately, so the matching name is never reassigned and adding a pattern for
such a family cannot change matching at all. A stream that does resolve has its
matching name replaced by the programme currently airing, and on the channel
side the channel name is replaced. So the old ranking promoted the families
where a pattern changes behaviour and demoted the families where it cannot.

MEASURED against the live installation rather than argued. 3,700 streams carry
an EPG identifier. Only 218 of those identifiers match a guide row, and 28
streams could resolve a programme at the moment of measurement. Of the 16
families the old ranking promoted, exactly one held a stream whose identifier
matched a guide row, and none could resolve a programme. That one is a numbered
channel lineup, which is the case a pattern harms. The largest family on the
installation, at 273 streams, was reduced to a single line with no suggested
pattern. Separately, the 16 families carrying identifiers and the 9 families
whose names contain an event word do not overlap at all.

WHAT CHANGED. Families are ranked by size, largest first, which is what the
reporter of the issue did by hand. Every uncovered family gets a suggested
pattern. The EPG identifier count stays as a note on each family, reported as
how many of its streams carry one out of how many. A family where at least half
of them do is marked with a caution saying a pattern there replaces a working
name with whatever is airing. The notification leads with the largest family
and carries the same caution when it applies.

On the live data the caution fires on 13 of the 131 uncovered families, every
one a recognisable channel lineup such as ITV, BeIN Sports or TNT Sport, and
those families moved from ranks 1 to 16 down to ranks 31 to 60. The three
largest families now lead the readout, each with a pattern to paste, and all
three carry no EPG identifier at all.

Verification: 1,510 tests pass. Twenty-nine mutations were applied one at a
time and every one was caught by a named test, with a comment-only control
leaving the suite green. That includes putting the old ranking back, which now
fails a test that says why it was removed. Re-measured over the same 25,068
live stream names: output plain ASCII, all 132 suggestions match their own
example name, no name skipped and the time budget not tripped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ZBufwmvsXqzkq2h2VkAwB
@PiratesIRC
PiratesIRC merged commit 66d9ce9 into main Sep 6, 2026
6 checks passed
@PiratesIRC
PiratesIRC deleted the feat/issue-43-placeholder-scan branch September 6, 2026 18:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant