Scan for uncovered placeholder name families (#43) - #52
Merged
Merged
Conversation
Adds the action Scan for Placeholder Patterns. It groups every stream name into a numbered family by replacing its numbers with a slot, and reports the families that no configured Placeholder Name Pattern covers, each with an anchored regular expression to paste. Requested as issue #43. The setting only ever helped with the naming schemes the operator already thought to write down, and nothing in the interface told apart "this installation has no placeholder families" from "the patterns you wrote match none of them". The grouping lives in a new stdlib-only module, Stream-Mapparr/placeholder_scan.py, with no Django import, so it is tested directly. The action is the thin wrapper that loads the streams and writes the readout to /config/stream-mapparr/placeholder-name-scan.txt. Measured against 25,068 live stream names, and three decisions came from that rather than from assumption: A digit immediately followed by K is left alone, because 4K and 8K are resolution tags. With the rule off, 15 further templates covering 418 streams group by their resolution tag instead of by a slot number. A family needs three different numbers in its slot, not three streams. A separate stream-count threshold was written and removed: three distinct numbers already implies three streams, so it could never refuse anything the first rule admitted, and no test could show it doing work. Families whose streams carry EPG data are reported first and separately. On this installation 131 families are uncovered and only 16 hold a stream with an EPG identifier, so one undifferentiated list would report a problem eight times larger than the one worth acting on. The readout is plain ASCII with any other character written as a backslash-u escape, the same rule as the CSV export preamble. Provider names here really do carry such characters, and an earlier version of the ASCII test used only synthetic names, so it passed while the live report broke the rule. Python regular expressions accept that form; all 132 suggestions generated from live data were checked against their own example name. Every guard was proven by mutation rather than by reading: each was broken in turn and a named test confirmed to fail, with a comment-only control leaving the suite green. That found one vacuous test, which sized its own input from the threshold it was checking and so moved with it. Reports only. No setting is changed and nothing is written to the database. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017ZBufwmvsXqzkq2h2VkAwB
Two reviews of the first version of the placeholder family scan found four defects. Each was reproduced before it was changed, and each is now covered by a named test that fails when the fix is reverted. A character above the basic multilingual plane, an emoji or a flag, was escaped with the lower-case backslash-u form. That form carries a MINIMUM of four hex digits, not exactly four, so an emoji produced five and Python read only the first four. The pattern compiled, matched nothing, and the family kept reporting as uncovered with nothing saying why. The existing round-trip test used a character inside the basic plane, which is the one class the old form already handled. Characters above the plane now use the upper-case eight-digit form. A literal backslash was left unescaped in the template while a literal hash was escaped, so a backslash followed by a digit produced the same two characters as an escaped hash. Two different names grouped into one family and the suggested pattern matched neither. The scan ran inside the request with no yield, no input cap and no time budget. The pattern safety gate deliberately admits patterns that can backtrack polynomially, on the stated promise that the runtime bounds the input, and that promise was kept only on the regex pre-processing path. The walk now hands the worker back every 500 names, skips a name over 500 characters and stops at a five second budget, using the same PluginConfig limits as that path. A stopped walk is reported as partial rather than passed off as finished: a family the walk never reached is missing, not covered. A failure to write the readout returned a plain success. The readout exists only in that file, so the notification now says the file is missing, and says it FIRST, because the notification helper drops lines from the end and a test caught the warning being truncated away. Also from those reviews: distinct slot numbers are counted within one slot rather than across slots, so two slots holding two values each no longer count as four; a resolution tag groups regardless of the case of its K, since patterns are compiled case-insensitively; a suggestion longer than the setting will accept is marked as such instead of being offered as though it worked; a database row that is not a dictionary is skipped rather than raising out of the walk; and the second listing is capped and says how many it left out. Verification: 1,506 tests pass. Twenty-four mutations were applied one at a time and every one was caught by a named test, with a comment-only control leaving the suite green. One mutation initially escaped, because the test asserted the argument NAMES appeared in the call and setting them to None keeps the names; the assertion now pins the values. Re-measured over the same 25,068 live stream names: 132 families, 131 uncovered, 16 carrying an EPG identifier, 0.26 seconds, no name skipped, budget not tripped, output plain ASCII, and all 132 suggestions match their own example name. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017ZBufwmvsXqzkq2h2VkAwB
The first version ranked uncovered families by how many of their streams carried an EPG identifier, and split the readout in two, suggesting a pattern only for families that carried one. Four independent reviews and a measurement of the live database agree that was wrong in both directions. WHICH DIRECTION THE RISK RUNS, read from the code. In _resolve_current_epg_title_for_stream a stream with no EPG identifier returns immediately, so the matching name is never reassigned and adding a pattern for such a family cannot change matching at all. A stream that does resolve has its matching name replaced by the programme currently airing, and on the channel side the channel name is replaced. So the old ranking promoted the families where a pattern changes behaviour and demoted the families where it cannot. MEASURED against the live installation rather than argued. 3,700 streams carry an EPG identifier. Only 218 of those identifiers match a guide row, and 28 streams could resolve a programme at the moment of measurement. Of the 16 families the old ranking promoted, exactly one held a stream whose identifier matched a guide row, and none could resolve a programme. That one is a numbered channel lineup, which is the case a pattern harms. The largest family on the installation, at 273 streams, was reduced to a single line with no suggested pattern. Separately, the 16 families carrying identifiers and the 9 families whose names contain an event word do not overlap at all. WHAT CHANGED. Families are ranked by size, largest first, which is what the reporter of the issue did by hand. Every uncovered family gets a suggested pattern. The EPG identifier count stays as a note on each family, reported as how many of its streams carry one out of how many. A family where at least half of them do is marked with a caution saying a pattern there replaces a working name with whatever is airing. The notification leads with the largest family and carries the same caution when it applies. On the live data the caution fires on 13 of the 131 uncovered families, every one a recognisable channel lineup such as ITV, BeIN Sports or TNT Sport, and those families moved from ranks 1 to 16 down to ranks 31 to 60. The three largest families now lead the readout, each with a pattern to paste, and all three carry no EPG identifier at all. Verification: 1,510 tests pass. Twenty-nine mutations were applied one at a time and every one was caught by a named test, with a comment-only control leaving the suite green. That includes putting the old ranking back, which now fails a test that says why it was removed. Re-measured over the same 25,068 live stream names: output plain ASCII, all 132 suggestions match their own example name, no name skipped and the time budget not tripped. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017ZBufwmvsXqzkq2h2VkAwB
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements #43, a scan that finds the numbered stream-name families that no Placeholder Name Pattern covers.
What it does
A new action, Scan for Placeholder Patterns. It groups every stream name into a family by replacing digit runs with a slot, so
MAX 100andMAX 101become the one familyMAX #, and reports the families that no configured pattern covers, each with an anchored regular expression to paste. It reads one database column that matching already loads, opens no provider connection, changes no setting and writes nothing to the database. The full readout goes to/config/stream-mapparr/placeholder-name-scan.txt, since a notification shows only about 280 characters.The grouping lives in a new file,
Stream-Mapparr/placeholder_scan.py, which imports nothing but the standard library, so it is tested directly. The action is the wrapper that loads the streams and writes the readout.What was measured rather than assumed
Everything below was measured against 25,068 live stream names and the live guide database.
A digit followed by K is left alone, because
4Kand8Kare resolution tags. With that rule off, 15 further templates covering 418 streams group by their resolution tag instead of by a slot number.A family needs three different numbers in its slot, not three streams. A separate stream-count threshold was written and removed, because three distinct numbers already implies three streams, so it could never refuse anything the first rule admitted.
The readout is plain ASCII with any other character written as a backslash-u escape, the same rule as the CSV export preamble. Provider names here carry such characters. Python regular expressions accept that form, so every suggested pattern still matches the name it came from.
The ranking was reversed after review
The first version ranked families by how many of their streams carried an EPG identifier, and suggested a pattern only for those that did. Four independent reviews and a database measurement showed that was wrong in both directions.
A stream with no EPG identifier returns from the resolver immediately, so a pattern over such a family cannot change matching at all. A stream that resolves has its matching name replaced by the programme currently airing. The old ranking therefore promoted the families where a pattern changes behaviour and demoted the ones where it is inert.
Measured: 3,700 streams carry an EPG identifier, only 218 of those identifiers match a guide row, and 28 streams could resolve a programme at the moment of measurement. Of the 16 families the old ranking promoted, one held a stream whose identifier matched a guide row and none could resolve a programme. That one was a numbered channel lineup, the case a pattern harms. The largest family on the installation, at 273 streams, was reduced to one line with no pattern.
Families are now ranked by size, which is what the reporter of the issue did by hand. Every uncovered family gets a pattern. The EPG count stays as a note on each family, and a family where at least half its streams carry an identifier is marked with a caution that a pattern there replaces a working name. On live data that caution fires on 13 families, every one a recognisable channel lineup, and they moved from the top of the list to ranks 31 to 60.
What this does not solve
Size alone does not separate event placeholder families from informative numbered channel families either. The six largest uncovered families include three of each kind. The readout says so in its own text rather than implying the ranking settled it.
Verification
1,510 tests pass. Twenty-nine mutations were applied one at a time and every one was caught by a named test, with a comment-only control leaving the suite green. That set includes restoring the old ranking, which now fails a test that says why it was removed. The publish audit is clean with all seven deny rules proven to fire.
🤖 Generated with Claude Code
https://claude.ai/code/session_017ZBufwmvsXqzkq2h2VkAwB