Skip to content

Final paper and small tokenizer fix - #8

Merged
jsture merged 3 commits into
mainfrom
paper-n-tokenizer
Jun 19, 2026
Merged

Final paper and small tokenizer fix#8
jsture merged 3 commits into
mainfrom
paper-n-tokenizer

Conversation

@jsture

@jsture jsture commented Jun 18, 2026

Copy link
Copy Markdown
Contributor

No description provided.

@jsture jsture self-assigned this Jun 18, 2026
jsture and others added 2 commits June 18, 2026 16:45
- Skip malformed sequences during training instead of aborting the run,
  and map malformed input to <unk> at encode time so it stays detectable
  via unk_rate rather than raising mid-batch.
- Debit both constituents when a merge is applied so *_freq.json reflects
  accurate post-merge counts; primitive keys are retained for coverage.
- Drop the per-merge full copy of pair_counts (never read) and reuse
  untouched sequences by reference during rebuild. Output is unchanged;
  merge order stays pinned for reproducibility.
- Document that encoding is greedy longest-match, not a replay of learned
  merge order, and that both phases share it.
- Remove dead two-letter metal branches from SMILES_RE (always bracketed
  in canonical SMILES); verified identical splits on the shipped vocab.
- Correct the misleading "unbalanced parens = defect" comment: boundary-
  crossing SMILES tokens are inherent to APE and match the paper.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@jsture
jsture merged commit c97cef2 into main Jun 19, 2026
1 of 2 checks passed
@jsture
jsture deleted the paper-n-tokenizer branch June 19, 2026 07:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant