Measure what a local model actually does on the hardware you actually have.
VirtualV LLM is a hardware-aware evaluation suite for GGUF/llama.cpp and 1Cat-vLLM workloads on heterogeneous GPU pools. It keeps quality, throughput, hardware topology and provenance together, while failures remain explicit instead of becoming invented scores.
The reference host combines an RTX A4000, two Tesla V100-SXM2 32 GB cards over NVLink and an RTX 4000 Ada. Its measurements illustrate the method; they are not universal performance claims.
- Interactive benchmark dashboard
- Curated machine-readable evidence
- Business standard and operator manual
- Current model and runtime roadmap
- Business PDF and weight-free v1 backup
The completed migration from numerai-signals is documented in Git history.
The active runners, scheduled dashboard build and systemd unit now use this
repository. Model weights, secrets, private evaluation packs and raw logs are
excluded from Git; raw evidence is published separately as a checksummed
release artifact.
git clone https://github.com/virtuanalytica/virtualv_llm.git
cd virtualv_llm
cp .env.example .env
python3 scripts/benchmarks/bootstrap_public_data.py
python3 -m py_compile scripts/benchmarks/*.py scripts/reporting/*.py
python3 scripts/benchmarks/build_dual_v100_html.pyBenchmark an OpenAI-compatible local endpoint:
python3 scripts/benchmarks/well_known_suite.py MODEL_ID \
--external-url http://127.0.0.1:PORT \
--external-model SERVED_MODEL \
--physical-gpus 1,2 \
--topology '2x Tesla V100-SXM2-32GB NVLink' \
--engine 'llama.cpp' \
--out reports/well_known_suite_local.jsonThe full run protocol, failure handling and evidence requirements are in the
operator standard. Python 3.12+ is
recommended. A compatible local inference runtime and lm-eval are required
for actual GPU runs.
- Curated JSON, HTML, documentation and PDFs are tracked.
- Raw logs are immutable release artifacts with SHA-256 manifests.
- No result is comparable without matching protocol, engine, context and hardware metadata.
- Long GPU jobs are resumable, exclusive and independently managed by systemd.
Licensed under Apache-2.0. Third-party models, datasets and runtimes retain
their own licenses; consult NOTICE before redistributing them.
Maintainers can assemble the raw evidence without copying it into Git:
python3 scripts/reporting/build_release_evidence.py \
--source /path/to/reports/lm_eval_runs