Prompt: Max User Count Estimation
You are tasked with identifying the maximum concurrent user count that can be supported by a given AI model deployment.
Use the following parameters to guide your calculation:
- VRAM Capacity (GB): Total available GPU memory.
- Model Size (GB): Memory footprint of the loaded model.
- Context Window (tokens): Maximum sequence length per user request.
Requirements:
- Calculate how many users can be served simultaneously without exceeding VRAM.
- Factor in memory overhead for:
- Model weights
- Key-value cache per user (scales with context window)
- System/runtime buffers
- Provide a formula or step-by-step reasoning.
- Output:
- Estimated Max Concurrent Users
- Breakdown of VRAM usage per user
- Assumptions made (e.g., precision type FP16/INT8, average prompt length).
Example Input:
- VRAM: 24 GB
- Model Size: 12 GB
- Context Window: 4,096 tokens
Example Output:
- Max Concurrent Users: 8
- VRAM per User: 1.5 GB (KV cache + buffers)
- Assumptions: FP16 precision, average prompt length 2,048 tokens
Prompt: Max User Count Estimation
You are tasked with identifying the maximum concurrent user count that can be supported by a given AI model deployment.
Use the following parameters to guide your calculation:
Requirements:
Example Input:
Example Output: