Skip to content

feat: add pinned repository optimizer - #23

Merged
quangshuynh merged 3 commits into
mainfrom
feat/pinned-repository-optimizer
Sep 21, 2026
Merged

quangshuynh merged 3 commits into
mainfrom
feat/pinned-repository-optimizer

Conversation

@quangshuynh

Copy link
Copy Markdown
Owner

What this adds

GitHub profiles pin up to six repositories, and GitProfileLens had no answer to which six. Existing pin advice ranked individual repositories by score, which is the ordering the score already gives.

The Pinned repository optimizer answers a set question instead: which combination of repositories forms the strongest portfolio set, and why. It reads finished audits and writes to none of them.

This is explicitly not sort by score descending; take 6.

Product distinction

Repository score Portfolio candidacy Pinned optimizer
Question How well does this repository present itself? Is this a good repository to feature? Which combination forms the strongest set?
Scope One repository One repository The whole set
Computed by scoreRepository classifyPortfolioCandidate optimizePinnedSet

Pin limit

Six, from GitHub's own documentation: "Select up to six repositories and gists, combined." The product already encoded the same invariant in pinnedItems(first: 6), in pin advice, and in the corpus fixture builder. pinned-optimizer.js now carries the constant with the documentation quoted beside it, and the optimizer recommends up to six, never padding to reach it.

Eligibility, decided before any ranking

  1. Not classified De-emphasize.
  2. Something verified a visitor can read: a description scoring ≥ 70, or a README verified present/comprehensive.
  3. No open high finding on a Worth polishing repository.
  4. Public.

A numeric score never qualifies a repository. Archive status, fork status, and score are deliberately not eligibility rules — they are handled by ranking.

Selection: rules, not a hidden score

No internal utility value, no weighted sum, nothing a user could mistake for a second score. One total lexicographic order:

  1. candidacy tier → 2. originality (original > unknown > fork) → 3. active > archived → 4. redundancy → 5. presentation score → 6. maintenance → 7. verified metadata → 8. repository name

Rule 8 makes the order total, so the result never depends on GitHub's response order. Because redundancy sits at rule 4, breadth can never promote a lower tier, a fork over an original, or an archive over an active repository.

Diversity signals actually used

Primary language and topic overlap. Nothing else. GitProfileLens has no reliable project or domain categories, so it does not invent any — no "backend"/"mobile"/"DevOps" inference from README prose. Topic overlap needs two shared topics, because one shared topic such as python says very little.

Breadth is a credit earned by verified distinctness. An unreported language or absent topics earns no claim and takes no penalty — an unknown signal ranks exactly with a repeated one.

Homepage presence is recorded as set completeness, not as redundancy: two repositories both linking a demo is not a story overlap.

No circularity

Current pin state is read in one place, readCurrentPins, and used for one purpose: the comparison. Asserted behaviorally — for every corpus profile, re-running with every repository pinned, and again with none pinned, must produce the identical recommended set.

F2 tension

The optimizer does not read scorePortfolioFocus or any profile-level score. Feeding a score that rewards archived/forked repositories into a selection built from candidacy that penalizes them would compound the disagreement, not resolve it. F2 is unchanged and remains open on its own terms.

Corpus findings

Two problems surfaced from inspecting all 13 profiles before freezing the baseline, and both were fixed in the rules rather than snapshotted:

  • An unreported language originally outranked a known-duplicate language, which dropped a flagship repository from oss-maintainer in favor of a lower-scoring repository with no reported language. Unknown was acting as positive evidence. Redundancy is now a credit earned only by verified distinctness.
  • "Polish first" was emitted for a repository with nothing to polish (Worth polishing solely because of a fork relationship), producing a broken sentence. That action now requires gaps the audit actually named.

fork-dominated receives six forks, because it has no eligible original repositories at all. Each entry states the authorship limit and none claims the owner did no work. Recommending nothing there would have implied forks are worthless.

UI location

A sixth top-level tab, measured rather than assumed. The tab strip is overflow-x: auto by design and already overflows at phone widths with five tabs (514px of content in 364px at 390px wide), so a sixth extends a handled case rather than introducing one; it fits at ≥768px. Nesting it under Overview would have buried it beneath the score layout, six category explainers, five recommendation cards, and the methodology block, and would have blurred the product distinction the codebase works to maintain.

Opening the tab issues no additional GitHub request — it runs over the audit already in memory.

Regression guarantees

Unchanged: repository presentation scores, category scores, portfolio candidate classifications, fork semantics, Network behavior, Network Markdown, repository Full/Compact Markdown, GitHub authentication permissions.

npm run eval reports no corpus outcome changed. The only existing expectation this interval changes is the five-tab keyboard test, which becomes a six-tab test.

Markdown export is deliberately untouched and documented as a possible later enhancement.

Verification

Command Result
npm run check pass
npm test 260/260 pass
npm run eval No corpus outcome changed
npm run eval:pins No corpus recommendation changed
npm run test:browser 29/29 pass, 0 skipped
git diff --check clean

GitHub profiles pin up to six repositories, and the product had no answer
to which six. Pin advice ranked individual repositories by score, which is
the same ordering the score already gives.

The optimizer answers a set question instead: which combination of
repositories forms the strongest portfolio set. It reads finished audits
and writes to none of them, so no repository score, category score,
finding, or candidacy label moves when it runs.

Eligibility is decided from candidacy and fact before any ranking, so a
high score alone never qualifies a repository. Selection is a fixed
lexicographic order rather than a weighted utility value: candidacy, then
confirmed original work over unreported fork status over a
GitHub-identified fork, then active over archived, then redundancy against
the set so far, then presentation score, maintenance, verified metadata,
and finally repository name. The name tie-break makes the order total, so
the result never depends on the order GitHub returned repositories in.

Redundancy uses only primary language and topic overlap, because
GitProfileLens has no reliable project or domain categories and does not
invent any. Breadth is a credit earned by verified distinctness: an
unreported language earns no claim and takes no penalty, matching how
candidacy already treats unknown evidence.

Current pin state is read in one place and used for one purpose, the
comparison. Using it as evidence would make a recommendation
self-reinforcing once acted on.

The interface adds a sixth tab. The tab strip is already horizontally
scrollable and already overflows at phone widths with five tabs, so a
sixth extends a handled case rather than introducing one, and the panel
would have buried Overview had it been nested there.
Adds three layers of coverage plus a corpus report.

tests/pinned-optimizer.test.js pins the rules against constructed profiles
built to trigger one rule at a time, including every required
falsification case: redundant top scores yielding to a lower-scoring
candidate with observable breadth, a high-scoring fork that does not
displace a confirmed original, a high-scoring archive that does not
displace an active project, a de-emphasized repository that is not
selected however well it scores, a Worth polishing repository taking the
final slot when its gaps are minor, a private Strong candidate recognized
without being called pinnable, unknown fork status claimed as neither
original nor forked, an already-optimal set, a short set, and stable
tie-breaking across repeated and reversed runs.

tests/scoring/pinned-sets.test.js pins the same rules against the
evaluation corpus, where the optimizer has to choose between repositories
that are all plausible. It also asserts the two properties that keep the
feature honest: pin state changes the comparison and never the
recommendation, and running the optimizer moves no score, category score,
finding, or candidacy on any corpus profile.

evaluation/pinned-report.js is a separate report from npm run eval on
purpose. A recommendation is not a scoring outcome, so mixing the two
would make a selection change read as a scoring regression and would
invite recording a new scoring baseline to settle a recommendation
question. Every corpus recommendation was inspected before the baseline
was recorded; two findings came out of that pass and were fixed in the
rules rather than snapshotted.

Browser coverage renders the recommended set, the current pins, the
suggested changes, and the rationale, and follows the link from a
recommendation to that repository's own audit card. Three further browser
tests cover a short set, unverified pin metadata, and strong private work
in the authenticated audit.

The existing five-tab keyboard test becomes a six-tab test, which is the
only existing expectation this interval changes.
Records the three-way distinction the product now makes: the repository
score measures how one repository presents itself, candidacy measures
whether one repository is worth featuring, and the optimizer measures
which combination forms the strongest set.

Documents the pin limit and its evidence, the eligibility rules and why
each exists, the full selection order, what the diversity signals are and
what they deliberately are not, fork and archive and private handling,
missing-evidence handling, the current-versus-recommended comparison, and
the limitations.

Records two things explicitly. The optimizer does not read
scorePortfolioFocus or any profile-level score, so the F2 tension is not
one of its inputs and F2 is unchanged. And GitProfileLens recommends a
portfolio set from observable repository presentation and metadata; it
does not measure engineering ability and does not determine how much code
a profile owner authored within a fork.
@vercel

vercel Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
gitprofilelens Ready Ready Preview Sep 21, 2026 5:09pm UTC

@quangshuynh
quangshuynh merged commit fae3a0b into main Sep 21, 2026
4 checks passed
@quangshuynh
quangshuynh deleted the feat/pinned-repository-optimizer branch September 21, 2026 17:21

This branch was successfully deployed

1 active deployment
Preview — 35012053 Deployed Sep 21, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant