Right now the summary rules are tuned by manually comparing the manager's hand-written summary against the model output, day by day (see docs/SUMMARY_RULES.md).
It would be great to turn that into an offline eval script:
Discussion on scoring metrics welcome.
Right now the summary rules are tuned by manually comparing the manager's hand-written summary against the model output, day by day (see
docs/SUMMARY_RULES.md).It would be great to turn that into an offline eval script:
roll-call text -> expected summarysamples as the gold standard;Discussion on scoring metrics welcome.