Skip to content

There is an issue with running the simple bench #67

Description

@sqz1169846252

In the 143rd sample group of the GPT-OSS-120B dataset, the results obtained by Fastokens and HF tokenizers are inconsistent, as Fastokens employs multithreading during processing, leading to discrepancies compared to the single-threaded text segmentation results. Currently, this appears to be caused by a bug in Fastokens' handling of the O200k tokenization scheme.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions