| title | RAG Leaderboard v2.1 |
|---|---|
| emoji | 🏁 |
| colorFrom | blue |
| colorTo | indigo |
| sdk | docker |
| pinned | false |
Leaderboard for evaluating RAG (Retrieval-Augmented Generation) systems.
- Download the public question set from
data/questions/questions_public.jsonl - Run your RAG pipeline and generate answers
- Upload a JSONL file with your answers — one JSON object per line:
{"id": "0", "answer": "Your answer here"}
{"id": "1", "answer": "Another answer"}- Each answer is graded by Grok (LLM-as-judge) on a 0 or 1 scale:
1— correct (semantically equivalent to gold answer)0— wrong or empty
| Variable | Description |
|---|---|
XAI_API_KEY |
Your xAI API key (required for judging) |
HF_TOKEN |
HuggingFace token (for gold answers dataset + leaderboard upload) |
GOLD_DATASET_ID |
HF dataset with gold answers (default: datakomarov/RAG-data-v2) |
GOLD_FILENAME |
Filename in the dataset (default: answers_gold.jsonl) |
THIS_SPACE_ID |
This Space's repo ID, e.g. datakomarov/RAG-LB-v2 |
EVAL_MODEL |
Grok model to use (default: grok-4-1-fast-reasoning) |
EVAL_CONCURRENCY |
Parallel judge calls (default: 5) |
Store your gold answers in a private HF dataset:
{"id": "19-1", "question": "Какую модель использовал Николай Кобало?", "answer": "Модель SEIR...", "context": "Опциональный контекст из корпуса..."}
{"id": "14-3", "question": "Как тимлид может поддерживать мотивацию?", "answer": "Декомпозировать задачи..."}Поля question и context опциональны, но рекомендуются — судья использует их при оценке.