Skip to content

Make field matching meaningful and add an eligibility evaluation harness - #145

Open
ma7moudalysalem wants to merge 1 commit into
mainfrom
feat/field-matching-and-eligibility-evaluation
Open

ma7moudalysalem wants to merge 1 commit into
mainfrom
feat/field-matching-and-eligibility-evaluation

Conversation

@ma7moudalysalem

Copy link
Copy Markdown
Owner

What this changes

The recommender and the eligibility checker both compared a student's field of study against a listing by exact string equality, across two vocabularies that never met. Listings are written from a fifteen-item canonical list; the seeder wrote fine-grained specialisms and the student profile takes free text. A student interested in Software Engineering could never match a listing open to Engineering.

Scoring

The match score is now the share of the attainable weight a listing earns across the criteria it actually constrains, instead of a fixed point sum. A criterion the listing leaves open drops out of both the numerator and the denominator, so an award open to every discipline is no longer punished for that openness — previously Gates Cambridge and Fulbright scored zero on the 40-point field term for being open to everyone.

Field comparison is graded: an exact match counts in full, a shared significant word counts as a half. The explanation now names the criteria that matched rather than restating the score band, and says plainly when nothing matched.

Eligibility

The field criterion could never report "not met" — its final branch returned partial, so a law student was told a computer science award was partially met. Unrelated fields now report not met, and "partially met" is reserved for a genuinely adjacent field.

Data

The seeder writes FieldsOfStudyJson on generated listings and states student preferences in the same canonical vocabulary.

before after
open listings carrying a field 0 of 336 733 of 763
(student, listing) pairs sharing a field 0.0% 12.2%

The curated external dataset carries fieldsOfStudy: five entries restrict discipline, nineteen are explicitly open to any field.

Evaluation harness

server/tools/ScholarPath.Eval runs the delivered checker over a seeded database with a fixed seed and reports per-criterion outcomes, so a figure can be reproduced:

dotnet run --project server/tools/ScholarPath.Eval -- --students 60 --listings 60 --seed 20260918

On 3,600 pairs:

field criterion before after
not met 0% 90.4%
partially met 91.9% 5.2%
met 8.1% 4.4%
unrelated field reported as partially met 93.5% 0%

Two claims made true in code

The SignalR Redis backplane is now wired when Redis is configured (AddStackExchangeRedis was never called), and WebKit and Edge are added to the Playwright projects.

Verification

  • dotnet build — 0 warnings, 0 errors
  • dotnet test — 951 passing, 18 new (10 scoring, 8 field matching). The scoring path had no test coverage at all before this.

Note for the reviewer

Program.cs in the working tree also carries an unrelated, unfinished "Live System Activity" feature whose files are untracked. That work is deliberately not included here — only the SignalR hunk is.

…ation harness

The recommender and the eligibility checker both compared a student's field of
study against a listing by exact string equality, across two vocabularies that
never met: listings are written from a fifteen-item canonical list, while the
seeder wrote fine-grained specialisms and the student profile takes free text.
A student interested in "Software Engineering" could never match a listing open
to "Engineering".

Scoring
- The match score is now the share of the attainable weight a listing earns
  across the criteria it actually constrains, rather than a fixed point sum.
  A criterion the listing leaves open drops out of both sides of the ratio, so
  an award open to every discipline is no longer punished for that openness.
- Field comparison is graded: an exact match counts in full, a shared
  significant word counts as a half.
- The explanation now names the criteria that matched instead of restating the
  score band, and says plainly when nothing matched.

Eligibility
- The field criterion could never report "not met" — its final branch returned
  "partially met", so a law student was told a computer science award was
  partially met. Unrelated fields now report "not met" and "partially met" is
  reserved for a genuinely adjacent field.

Data
- The seeder writes FieldsOfStudyJson on generated listings and states student
  preferences in the same canonical vocabulary. Measured on a seeded database,
  listings carrying a field went from 0 of 336 to 733 of 763, and the share of
  (student, listing) pairs sharing a field from 0.0% to 12.2%.
- The curated external dataset carries fieldsOfStudy; five entries restrict
  discipline and nineteen are explicitly open to any field.

Evaluation
- server/tools/ScholarPath.Eval runs the delivered checker over a seeded
  database with a fixed seed and reports per-criterion outcomes. On 3,600 pairs,
  fields with nothing in common reported as "partially met" fell from 93.5% to
  zero.

Also enables the SignalR Redis backplane when Redis is configured, and adds
WebKit and Edge to the Playwright projects, so both claims hold in fact.

Tests: 951 passing (18 new).
Copilot AI lite review requested due to automatic review settings September 18, 2026 12:04

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants