Non-technical guide for small teams running a local AI chat server on Windows with NVIDIA GPUs.
ExLlamaSharp turns a Windows PC with an NVIDIA GPU into a private ChatGPT-style server for your office. People open a browser, pick a model, and chat. Developers can also call the same OpenAI-compatible API.
Default address after install: http://localhost:14563 (or the server’s LAN IP if networking was enabled in the wizard).
- Check the machine meets the basics (NVIDIA GPU with ~8 GB+ VRAM, Windows 10/11, enough free disk). IT can run
packaging/Check-Requirements.ps1. - Install with the ZIP package (
Install.ps1as Administrator) or ask IT to deploy the published server. - Open the browser to
http://localhost:14563.
If the page does not load, wait a minute for the Windows service to start, then open Diagnostics (/diagnostics) or ask IT to check the service status.
On first launch you will be guided through:
- Create admin — username and password for the web UI.
- Model folder — where downloaded models are stored (default under Program Data).
- Network — listen only on this PC, or allow other PCs on the LAN.
- Download a starter model — pick a recommended model and wait for the job to finish.
- Done — you land on the dashboard.
You can change these later under Settings.
- Open Models.
- Choose Pull / download from the library (Hugging Face style repo id) or Import a folder you already have.
- Wait until the job shows completed (Jobs list / progress).
- Click Load so the GPU starts serving that model.
Aliases let you give a short name (for example office-assistant) that the Chat page and API use.
Large models need more VRAM and disk. If load fails, try a smaller quantized model or free GPU memory.
- Open Chat.
- Select the loaded model (or alias).
- Type a message and send.
Tips:
- System prompts and conversation history live in the UI; clear a thread when starting a new topic.
- If replies stop or error, check Diagnostics — often “No model loaded”.
- For API use from apps, create a key under API Keys and see the in-app API guide.
Admins can:
- Create users and assign roles.
- Issue API keys with quotas / scopes for apps and integrations.
- Enable teams / tenants (when multi-tenancy is on) so groups stay isolated.
- Review audit activity and set moderation rules if needed.
Non-admins usually only need Chat (and maybe a personal API key if your policy allows it).
| Need | Where |
|---|---|
| Is the server healthy? | Diagnostics → Run health check |
| GPU / version | About |
| Change port or CORS | Settings → Network |
| Backup config & keys | Settings / Backup (admin) |
- Open Diagnostics and note which component is red/yellow (
database,engine,inference,disk). - See troubleshooting.md for the same common issues listed on that page.
- Admins: admin-guide.md.