Skip to content

Detect uncovered EPG placeholder-name patterns automatically #43

Description

@JHeat

Is your feature request related to a problem? Please describe.

I'm new to Stream-Mapparr and to this kind of stream-matching setup in general, and I hit a problem that I think will bite other new users too: epg_placeholder_name_patterns only helps with the exact naming schemes you think to add — there's no way to find out you're missing one.

I recently enabled EPG-aware placeholder matching and it looked complete: I had patterns for PPV EVENT #, LIVE EVENT #, PPV # |, and HBO #, all matching fine. It wasn't until I manually pulled every stream name via the API and grouped them by a digit-stripped template that I discovered MAX # (128 streams — my single largest category, bigger than the other four combined) and Triller TV | Event # (14 streams) were sitting there the whole time, silently never eligible for EPG resolution, because nothing matched their exact name shape. There was no error, no warning, no indicator in the UI — those channels just quietly never got current EPG titles resolved, and I'd have had no way to know unless I happened to go digging.

Describe the solution you'd like

An action (or a section of an existing one, like Validate Settings) that scans all current stream names, groups them by a digit-stripped template (e.g. MAX 100 → MAX #), and reports which templates look like numbered-placeholder families (recurring, sequential numbering, a decent-sized count) but don't match any configured epg_placeholder_name_patterns entry — something like:

⚠️ 2 likely placeholder pattern(s) not covered by your current settings:

  • MAX # (128 streams, e.g. "MAX 100") — add ^MAX \d+$?
  • Triller TV | Event # (14 streams, e.g. "Triller TV | Event 8") — add ^Triller TV \| Event \d+$?

It doesn't need to auto-add anything — just surfacing the candidates (with a ready-to-paste regex) would have saved me a manual API audit, and would help anyone else onboarding onto this feature see the full picture of their own stream list instead of only what they already knew to look for.

Describe alternatives you've considered

  • Manually auditing stream names via the Dispatcharr API and grouping by template, which is what I actually did — works, but it's not something a typical user should have to know how to do, and it's easy to think you're done when you're not (nothing tells you otherwise).
  • Documenting common provider placeholder formats in the README as a checklist — helps a bit, but every provider's naming scheme is different, so a static list will always be incomplete. A scan of the user's own actual stream list is the only thing that's guaranteed to be complete for their setup.

Additional context

This came out of the same EPG-aware placeholder matching work as PR #41/#42 (closed #41 → reopened as #42 after a rebase, both AI-assisted and disclosed as such) — not a one-off observation from a single session. Over the last couple of days on that PR: the initial feature landed with 4 placeholder patterns (PPV EVENT #, LIVE EVENT #, PPV # |, HBO #); a follow-up fix fixed the matcher stripping a schedule-suffix naming convention (Channel Name | Weekday @ Time) that was silently defeating matches; and only then, while re-validating that fix against live data, did the manual stream-name audit turn up two more entire placeholder families (MAX # — 128 streams, the single largest category — and Triller TV | Event # — 14 streams) that had been sitting uncovered the whole time with zero indication anything was missing. That progression — ship the feature, fix a matching bug, then discover a whole coverage gap by accident — is exactly why I think this needs to be a built-in scan rather than something a user stumbles into by getting lucky enough to go digging.

Built with AI-assisted tooling (Claude Code), same as the rest of this work — flagging that up front the same way as my other recent issues/PRs here. Happy to share the actual grouping approach I used (digit-stripped template + count) if it's useful as a starting point.

Activity

  1. PiratesIRC commented on Aug 8, 2026

    @PiratesIRC
    Owner

    Sorry for the slow reply. Pull request #42 is merged and on main now, and I have it deployed and running here as 1.26.2202106.

    I think this is the right idea, and the argument for it is the strongest part of the report: a setting that only helps with the naming schemes you already thought of gives you no way to find out you are missing one, and nothing in the UI distinguishes "no placeholders here" from "your patterns do not match anything". That is the same failure shape I try to avoid elsewhere in this plugin, where a check that cannot tell reports success.

    Before replying I tested your grouping approach against my own installation, 25,323 stream names, to see whether it holds up on data that is not yours. It does. Grouping by a digit-stripped template and keeping templates with at least 3 streams and at least 3 distinct small numbers gives 107 candidates covering 2,898 streams, which is 11.4% of the list. PPV EVENT # shows up with 180 streams, along with families I would never have thought to write a pattern for, including :MAX US # at 296 and :Paramount+ # at 253. So the core claim holds: the scan finds things a hand-written pattern list does not.

    Two things came out of that test that I think are worth building in from the start rather than discovering later.

    Digits inside resolution tags get stripped too. 4K and 8K become #K, which produces templates like US: CINEMANIA HOLLYWOOD # #K and, worse, merges names that are not the same family. My largest single false grouping was 293 streams collapsed under NO EVENT STREAMING NOW - | #K EXCLUSIVE | US: DAZN P..., where the template is really being driven by the resolution tag rather than by a slot number. Leaving a digit alone when it is immediately followed by K and preceded by a boundary would fix most of it.

    Not every numbered family is an event placeholder. The scan also surfaces UK: BBC RED BUTTON # at 96, US: HULU ORIGINALS # at 76 and UK: KARAOKE # at 32. Those are numbered channel families whose names are perfectly informative, and adding a placeholder pattern for them would make matching worse rather than better. A report that lists all 107 without ranking would be long and mostly not actionable.

    For that second point, the signal I would reach for is whether the streams in a template actually carry EPG data at all, since a placeholder can only ever be resolved if it does. That is already available on the stream rows the plugin loads. Ranking by "how many of these streams have EPG data and are not already covered by an existing pattern" would put the actionable families at the top and push the merely-numbered ones down.

    Two things I cannot judge from here, in the interest of being straight about it. My own placeholder-named streams carry no EPG identifiers at all, so the EPG-aware matching feature resolves nothing on this installation and your live testing remains the only end-to-end evidence it works. And I have not written any of this, so the refinements above are untested suggestions rather than a design I have proved out.

    Leaving this open as accepted. I am not going to promise a date. If you want to take it on, I would look at a pull request, and the grouping approach you described is the right starting point.

  2. JHeat commented on Aug 11, 2026

    @JHeat
    ContributorAuthor

    Much appreciated. I'll put some human eyes on this :)

  3. PiratesIRC commented on Aug 12, 2026

    @PiratesIRC
    Owner

    Following up on the status note above: 1.26.2241602 is now released, so the EPG-aware placeholder matching work is no longer main-only. It will reach the Dispatcharr plugin browser once the Hub listing updates, which is a separate step and can lag by a day.

    Nothing changes about this issue. It stays open and accepted, and the two things worth designing around are still the ones from the earlier reply: leaving a digit alone when it is immediately followed by K, so resolution tags do not collapse unrelated names into one family, and ranking candidates by how many of their streams actually carry EPG data, so numbered channel families that are perfectly well named do not crowd out the real placeholders.

  4. added a commit that references this issue on Sep 6, 2026
  5. PiratesIRC commented on Sep 6, 2026

    @PiratesIRC
    Owner

    This is built and merged to main as of today. Pull request #52.

    The action is called Scan for Placeholder Patterns. It does what you described: every stream name is grouped into a family by replacing its digit runs with a slot, so MAX 100 and MAX 101 become the one family MAX #, and the families that no configured pattern covers are reported, each with an anchored regular expression to paste. It reads one database column that matching already loads, opens no provider connection, changes no setting and writes nothing. The full readout goes to a file, because a notification here shows only about 280 characters.

    Both refinements from the earlier reply are in. A digit followed by K is left alone, so resolution tags do not drive the grouping. With that rule off, 15 further templates covering 418 streams group by their resolution tag instead of by a slot number on my data.

    The ranking went the other way from what I proposed, and I want to be straight about why, because the measurement contradicted me.

    I suggested ranking by how many streams in a family carry EPG data. I built that first, then measured it. On my installation 3,700 streams carry an EPG identifier, but only 218 of those identifiers match a guide row, and 28 streams could resolve a programme at the moment I measured. Of the 16 families that ranking promoted, one held a stream whose identifier matched a guide row and none could resolve a programme. That one was a numbered channel lineup, which is the case where a pattern makes things worse. Meanwhile the largest family on my box, at 273 streams, was pushed to a single line with no suggested pattern.

    Reading the code explains the direction. A stream with no EPG identifier returns from the resolver immediately, so a pattern over that family cannot change matching at all. A stream that resolves has its matching name replaced by the programme currently airing. So ranking by EPG presence promoted the families where a pattern changes behaviour and demoted the families where it is inert. That is backwards for a report meant to suggest patterns.

    So families are now ranked by size, which is what you did by hand, and every uncovered family gets a suggested pattern. The EPG count survives as a note on each family, saying how many of its streams carry an identifier out of how many, and a family where at least half do is marked with a caution that a pattern there replaces a name that is already matching. On my data that caution fires on 13 families, every one a recognisable channel lineup, and they moved from the top of the list down to ranks 31 to 60.

    One thing it does not solve, and the readout says so rather than implying otherwise. Size alone does not separate event placeholder families from numbered channel families whose names are already informative. My six largest uncovered families include three of each kind. I did not find a single measured signal that separates them, so the report ranks by size, marks what it can, and leaves the judgement to the reader.

    This is on main and not yet in a release, so it will reach the plugin browser at the next one. If you try it against your own list, I would be interested in whether the largest families it reports are the ones you would have picked out by hand.

  6. PiratesIRC commented on Sep 6, 2026

    @PiratesIRC
    Owner

    Released as 1.26.2491549: https://github.com/PiratesIRC/Stream-Mapparr/releases/tag/1.26.2491549

    The action is Scan for Placeholder Patterns. The Hub listing update is submitted as Dispatcharr/Plugins#283 and is waiting on a maintainer merge, so the plugin browser will keep offering the previous version until that lands. A manual install from the release works now.

    Closing this as done. If the scan misses a family on your list, or ranks one oddly, please open a new issue with the family name and roughly how many streams are in it, since that is the input I cannot see from here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions