Skip to content

Identifying the maximum concurrent user count #63

Description

@avikeid2007

Prompt: Max User Count Estimation

You are tasked with identifying the maximum concurrent user count that can be supported by a given AI model deployment.
Use the following parameters to guide your calculation:

  • VRAM Capacity (GB): Total available GPU memory.
  • Model Size (GB): Memory footprint of the loaded model.
  • Context Window (tokens): Maximum sequence length per user request.

Requirements:

  1. Calculate how many users can be served simultaneously without exceeding VRAM.
  2. Factor in memory overhead for:
    • Model weights
    • Key-value cache per user (scales with context window)
    • System/runtime buffers
  3. Provide a formula or step-by-step reasoning.
  4. Output:
    • Estimated Max Concurrent Users
    • Breakdown of VRAM usage per user
    • Assumptions made (e.g., precision type FP16/INT8, average prompt length).

Example Input:

  • VRAM: 24 GB
  • Model Size: 12 GB
  • Context Window: 4,096 tokens

Example Output:

  • Max Concurrent Users: 8
  • VRAM per User: 1.5 GB (KV cache + buffers)
  • Assumptions: FP16 precision, average prompt length 2,048 tokens

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions