Evidence-based Energy Estimation for AI Workloads using Retrieval Augmented Generation
WattBot 2025 challenges participants to build RAG systems that extract credible environmental impact estimates from academic literature. Our system must provide:
- π― Concise, citation-backed answers
- π Document IDs and supporting evidence
- β Explicit handling of unanswerable questions
- 32 scholarly articles (2019-2025) on AI's environmental impact
- Topics: Energy consumption, carbon emissions, water usage, sustainability
- Question Types: Numeric values, categorical terms, True/False
| Component | Weight | Criteria |
|---|---|---|
answer_value |
75% | Numeric accuracy (Β±0.1% tolerance) or exact categorical match |
ref_id |
15% | Jaccard overlap with ground truth citations |
is_NA |
10% | Proper handling of unanswerable questions |
| Metric | Score |
|---|---|
| Final WattBot Score | 0.886 |
| Public Leaderboard Rank | 1 |
| Private Leaderboard Rank | 1 |
Our HERO (Hierarchical Evidence Retrieval & Optimization) pipeline consists of four key components:
| Feature | HERO Approach | Traditional RAG |
|---|---|---|
| Document Structure | Preserves hierarchy (DocβSectionβLeaf) | Flat chunking |
| Retrieval Strategy | Hybrid (BM25 + Dense) with reranking | Single-method |
| Storage | SQLite single-file | External vector DBs |
| Visual Processing | OCR + Table extraction | Text-only |
| Citation Handling | Multi-document evidence aggregation | Single-source |
PDF Documents β Hierarchical Parsing β Structured Chunks
- Parser: PyMuPDF for text extraction
- Structure Preservation: Document β Section β Paragraph hierarchy
- Visual Processing: OCR for tables/figures containing critical metrics
- Chunking Strategy: ~150-180 words per chunk with overlap
Query β [BM25 + Dense Embeddings] β Top-K Chunks
- BM25: Keyword-based retrieval for precise term matching
- Dense Retrieval: Sentence-Transformers for semantic search
- Reranking: Weighted combination (Ξ±=0.6 for BM25, Ξ²=0.4 for Dense)
- Storage: SQLite for single-file reproducibility
Retrieved Context + Question β Gemini 2.5 Pro β Structured Answer
- Model: Google Gemini 2.5 Pro (gemini-2.0-flash-exp)
- Prompt Engineering: Few-shot examples + structured output format
- Rate Limiting: Exponential backoff for API stability
- Output: JSON with answer, citations, and supporting evidence
Raw Answers β Validation β Unit Normalization β Final Submission
- Numeric validation (Β±0.1% tolerance)
- Citation format checking
- "Unable to answer" fallback handling
Based on our experimental results:
π‘ Lesson: Do not treat PDFs as flat text strings. Preserving document hierarchy (Doc β Section β Leaf) allows precise targeting.
Implementation:
- Tracked section headers and metadata
- Maintained parent-child relationships in chunks
- Enabled context-aware retrieval
π‘ No Vector DB: Using SQLite as a single-file datastore reduces complexity and ensures reproducibility.
Benefits:
- β No external dependencies
- β Easy version control
- β Portable across environments
- β Resilient to API rate limits with backoff logic
π‘ Lesson: Critical energy metrics are often hidden in charts. We must process images (via Vision Models/OCR), not just filter them out.
Strategy:
- Extracted figures/tables as separate chunks
- Used OCR for table data
- Linked visual content to text context
.
βββ notebooks/
β βββ v1_gemini_2_5pro.ipynb # Initial pipeline implementation
β βββ Gemini_2_5_pro_ζ°pipeline_0_821.ipynb # Optimized version (0.886 score)
βββ docs/
β βββ WattBot2025_AttentionPlease_Technical_Report.pdf
β βββ WattBot2025_AttentionPlease_Presentation.pdf
βββ data/
β βββ metadata.csv # Document index
β βββ train_QA.csv # Training Q&A pairs
β βββ test_Q.csv # Test questions
βββ README.md
pip install pymupdf rank-bm25 sentence-transformers google-generativeai pandas# Set Google API Key
export GOOGLE_API_KEY="your-api-key-here"# Open the notebook
jupyter notebook notebooks/Gemini_2_5_pro_ζ°pipeline_0_821.ipynb
# Or run as script (if converted)
python src/main.py --input data/test_Q.csv --output outputs/submission.csvβ
Document Parsing: 32/32 papers processed
π Total chunks: 1778
π Hybrid Search: BM25 + Dense retrieval ready
π€ Generating answers...
Progress: 100% |ββββββββββββββββββββ| 250/250
πΎ Submission saved: outputs/submission.csv
| Metric | Baseline | HERO (Ours) | Ξ |
|---|---|---|---|
| Recall@K (Coverage) | |||
| Recall@1 | 79.49% | 80.49% | +1.00% |
| Recall@3 | 89.74% | 89.80% | +0.06% |
| Recall@5 | 92.31% | 92.68% | +0.37% |
| Recall@10 | 92.31% | 95.12% | +2.81% |
| nDCG@K (Ranking Quality) | |||
| nDCG@1 | 0.7949 | 0.8049 | +0.0100 |
| nDCG@3 | 0.8497 | 0.8502 | +0.0005 |
| nDCG@5 | 0.8564 | 0.8584 | +0.0020 |
| nDCG@10 | 0.8617 | 0.8665 | +0.0048 |
| Overall Accuracy | |||
| MRR | 0.8526 | 0.8560 | +0.0034 |
| Type | Count | Accuracy |
|---|---|---|
| Numeric | ~40% | [TBD] |
| Categorical | ~35% | [TBD] |
| True/False | ~15% | [TBD] |
| Unanswerable | ~10% | [TBD] |
| Component | Technology |
|---|---|
| Document Parsing | PyMuPDF, regex |
| Keyword Retrieval | BM25 (rank-bm25) |
| Dense Retrieval | Sentence-Transformers (all-MiniLM-L6-v2) |
| Vector Storage | SQLite |
| LLM | Google Gemini 2.5 Pro |
| Development | Python 3.10+, Jupyter Notebook |
- Hybrid retrieval outperformed single-method approaches by 12%
- Hierarchical chunking improved citation accuracy by preserving context
- Few-shot prompting reduced hallucinations in numeric answers
- Exponential backoff handled API rate limits gracefully
- Pure dense retrieval missed exact term matches (e.g., "BERT-base")
- Large chunks (>300 words) diluted relevant information
- Zero-shot prompting generated inconsistent citation formats
| Configuration | WattBot Score | Ξ |
|---|---|---|
| BM25 only | 0.64 | -0.18 |
| Dense only | 0.71 | -0.11 |
| HERO (Hybrid) | 0.82 | baseline |
| HERO + Visual processing | 0.85 | +0.03 |
- Competition Dataset: Endemann, C., Paul, D. J., & Zhao, A. (2025). WattBot 2025. Kaggle.
- Retrieval Methods: Robertson, S., & Zaragoza, H. (2009). The Probabilistic Relevance Framework: BM25 and Beyond.
- RAG Survey: Gao, Y., et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey.
Team Attention Please
- Shao-Hua Wuβ‘ - Document Processing & Retrieval System & Project Lead
- Xie-Pei Juβ‘ - Document Processing & Retrieval System & Hybrid Search Implementation & Answer Generation & Prompt Engineering & Optimization
- Bo-Hao Chenβ‘ - Answer Generation & Prompt Engineering & Optimization
- Yi-Chen Hsiaoβ - Report Generation
- Yi-Yang Xueβ - Retrieval System & Hybrid Search Implementation
β‘ Equal contribution
If you find our approach useful, please cite:
@misc{wattbot2025_hero,
title={HERO: Hierarchical Evidence Retrieval \& Optimization for WattBot 2025},
author={Wu, Shao-Hua and Ju, Xie-Pei and Chen, Bo-Hao and Hsiao, Yi-Chen and Xue, Yi-Yang},
year={2025},
howpublished={\url{https://github.com/your-repo/wattbot2025-hero}}
}This project is released under the MIT License. See LICENSE for details.
- Competition organizers at ML+X, University of Wisconsin-Madison
- Google for Gemini API access
- Open-source community for tools and libraries