Repository navigation
Add --max-model-len & add Laya multi-lingual and typed-decisions - #13
Merged
Merged
Conversation
alvarobartt
force-pushed
the
add-missing-laya-variants
branch
from
September 24, 2026 20:28
0485758 to
9296aa1
Compare
alvarobartt
added this pull request to stack #15
September 25, 2026 07:48
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
This PR adds the
--max-model-lenflag which defaults to themax_lenvalue set in in therl_agent_config.jsonfile for Laya, whereas usually the encoder context length is defined inmax_position_embeddingsin theconfig.json(in this case placed underencoder/config.json).Additionally, this PR adds the missing multi-lingual and typed-decisions Laya variants other than the English-only one; which is linked with the
--max-model-lenaddition, since for the multi-lingual version, despite being trained with sequences up to 1024 setting--max-model-len 8192is recommended to leverage a larger context length, as the encoder supports sequences up to 8192 (as set in itsconfig.jsonunderencoder/config.json).Finally, it also includes some slight changes in the GitHub Actions to make sure that the cache is reused efficiently and that's only triggered when applicable.