Skip to content

Long unbroken tokens create empty and oversized knowledge chunks #322

Description

@Calmingstorm

Failure scenario

  1. Ingest a non-empty document consisting of one unbroken word longer than CHUNK_SIZE, for example 1,501 x characters.
  2. KnowledgeStore._chunk_text() enters the oversized-paragraph word loop with an empty current_chunk.
  3. The first word does not fit, so the overflow branch appends current_chunk.strip() before assigning the word.

The result is two chunks: an empty chunk followed by the entire 1,501-character word. Ingest reports and stores both chunks, so the source has a phantom empty chunk and total_chunks = 2; the actual content chunk also remains larger than the advertised hard chunk size. Repeating the pattern with another oversized word inserts another empty chunk.

Site

  • src/knowledge/store.py:881-891 appends the empty accumulator when the first word itself exceeds CHUNK_SIZE, then carries the oversized word forward without a character-level fallback.

Expected result

Never emit empty chunks, and split a token that exceeds CHUNK_SIZE at a bounded character boundary so every stored chunk is non-empty and respects the chunk-size contract.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions