Items identified during the architectural audit that are not yet implemented. Organized by priority.
Model defaults (last updated 2026-04-23):
CHAT_MODELdefault changed toqwen/qwen3-6b-plus(wasgoogle/gemini-2.5-flash-lite)SPARK_MODELadded — defaults togoogle/gemini-2.5-pro; used only during active Spark sessions (Spark analysis requires correlated multi-signal reasoning that exceeds flash-lite capability)
| Tool | Source | Purpose |
|---|---|---|
check_server_status |
mcsrvstat.us API | Let the LLM check if a user's server is online when diagnosing |
get_paper_config_docs |
RAG + PaperMC docs | Dedicated config key lookup |
check_plugin_compatibility |
Modrinth version API | Check specific MC version support |
get_java_version_info |
Static mapping | MC version → recommended Java version |
Current problems:
_truncate_safe(result_text, limit=6000)per tool ×MAX_TOOL_ROUNDS=4= up to 24K chars of tool results- History replays 3 full Q&A pairs verbatim
max_tokens=1024truncates complex answers
Recommendations:
# Track cumulative context size
MAX_CONTEXT_CHARS = 20000
# Summarize older history
if history and len(history) > 1:
summary = f"Resumo: {history[-2]['question']} → {history[-2]['answer'][:200]}"
messages.append({'role': 'user', 'content': summary})
# Only keep last exchange verbatim
messages.append({'role': 'user', 'content': history[-1]['question']})
messages.append({'role': 'assistant', 'content': history[-1]['answer']})
# Increase max_tokens
max_tokens=2048Current behavior:
- If the LLM calls
search_docsmultiple times (in different tool rounds), the same document chunks may appear in both results - The assistant repeats information redundantly in its final answer
Action items:
- Verify — check
docs_rag.pyagentic loop to confirm if duplicatesearch_docscalls occur in practice - If confirmed — track
all_sourcesacross rounds and filter duplicates before sending to LLM - Optional enhancement — add a memo/context field to the LLM stating "avoid repeating these sources you just retrieved"
Note: This is lower priority than 1.3/1.4 and depends on runtime observation.
Rationale:
- Most common admin questions involve choosing between Survival, Creative, Adventure, and Spectator
- Resource usage, difficulty balancing, and version compatibility vary per gamemode
- Currently the LLM has no structured gamemode reference
Implementation:
- Create a new RAG vector store or static context chunk:
gamemode_reference.md- Survival: description, typical configs, RAM usage, common settings changes
- Creative: building/testing specific requirements, performance tips
- Adventure: map design constraints, custom rules
- Spectator: spectator-only server considerations
- Include per-gamemode recommendations:
- Most common MC versions (e.g., Survival = latest; Creative = stable)
- PaperMC config tuning (e.g.,
mob-spawning.per-player,difficulty) - Plugin compatibility by gamemode
- Embed during initial
/reindexso the LLM can cite it naturally
Spark Profiler — ✅ Implemented. See cogs/spark.py + cogs/spark_parser.py. Passive URL detection, /spark command, agentic detail sections, and dedicated SPARK_MODEL are all live.
Legacy Timings — Remaining work:
- Add
/timingscommand accepting Paper timings report URL (timings.aikar.co) - Parse the timings JSON: server info, tick breakdown, plugin timings, world stats
- Feed structured data to the LLM for diagnosis (similar to Spark flow)
- Still commonly used on older Paper servers that haven't migrated to Spark
Multi-step discord.ui.Modal:
- Collect: RAM, player count, MC version, server type (Paper/Purpur/Fabric)
- LLM generates optimized config snippets from RAG context
- Outputs
server.properties,paper-global.yml,purpur.ymlchanges - Converts passive documentation into active, personalized guidance
/compat <plugin> <version>:
- Query Modrinth/Hangar version-specific API endpoints
- Cross-reference with known incompatibilities
- Expose as an LLM tool (
check_compatibility) for natural-language queries - Addresses the most repetitive question type in support channels
Standalone command exposing the get_java_version_info tool (§1.3) as a user-facing slash command.
- Input: MC version (e.g.,
1.21.4) - Output: recommended Java version with explanation
- Static mapping, no external API needed:
- 1.21+ → Java 21
- 1.17–1.20.x → Java 17
- ≤1.16 → Java 8
- Also useful as a quick-reference embed in support channels
When users ask "which is better, X or Y?", fetch metadata from both plugins and present a side-by-side comparison embed:
- Downloads, last updated date, supported MC versions, description
- Uses existing
plugin_apis.pyinfrastructure (Modrinth, Hangar, SpigotMC) - Addresses a common question pattern in support channels
Input: server type (Paper/Purpur/Spigot) + MC version. Output: actionable checklist of recommended config optimizations.
- Paper:
chunk-loading,entity-activation-range,mob-spawning.per-player,view-distance,simulation-distance - Purpur: additional Purpur-specific tweaks (
purpur.yml) - Each item includes the config key, recommended value, and a one-line explanation
- More actionable than requiring users to know what to ask in
/ask
Lookup any server.properties key and return:
- Description of what it does
- Default value and valid range
- Common tuning advice
- Backed by RAG over PaperMC docs or a static knowledge base
- Related to §1.3
get_paper_config_docstool but as a user-facing command
Admin-only command showing bot health status:
- API connectivity: OpenRouter, GitHub, mclo.gs, Modrinth (latency + status)
- Vector store: age, chunk count, last reindex timestamp
- Conversation cache: active conversations, TTL stats
- Error counts since startup (by category)
- Useful for operators to verify bot health after deployment or config changes
Items identified and scoped but not yet implemented. Ordered by priority.
Problem: All GitHub API calls during indexing (tree discovery, commit SHAs) are unauthenticated. GitHub's unauthenticated rate limit is 60 req/hour. A single full reindex hits the tree API once per source (4 calls) plus one commit-SHA fetch per source (4 more). Exhausting the rate limit causes silent indexing failures.
Value: Prevents production indexing failures. Also unlocks private repos as future doc sources.
Implementation:
- Add to
config.py:GITHUB_TOKEN: Final[str | None] = os.getenv('GITHUB_TOKEN')
- In
DocsRAG.cog_load, ifGITHUB_TOKENis set, passAuthorization: Bearer <token>as a default header on theaiohttp.ClientSession. - Document in
SETUP.md— token only needspublic_repo(read) scope.
Problem: Only mclo.gs and pastebin.com are passively detected. hastebin.com, paste.gg, and GitHub Gist raw URLs are common in Minecraft admin communities and are silently ignored.
Implementation: Extend _parse_link() in cogs/log_analyzer.py:
hastebin.com/(\w+)→ raw:hastebin.com/raw/\1paste.gg/p/\w+/(\w+)→ raw:paste.gg/p/<...>/files/\1/rawgist.github.com/\w+/(\w+)→ raw:gist.githubusercontent.com/\w+/\1/raw
During the Spark integration design (see SPARK_INTEGRATION.md §14), two tools were explicitly excluded from the Spark agent's tool set:
web_search — Removed from Spark sessions.
Uncontrolled web results introduce noise and hallucination risk in a structured diagnostic workflow. The indexed PaperMC + Spark RAG docs provide grounded, version-specific recommendations without external noise. web_search remains available in /ask and /chat commands where open-ended Q&A benefits from real-time information.
get_known_issues — Not implemented.
This tool was considered to surface known plugin performance bugs. Without web_search as a data source, it has no reliable runtime data. Plugin-specific issues should be surfaced by search_docs(source="PaperMC") and by directing users to the plugin's own issue tracker.
Problem: There is no feedback signal to detect bad AI responses. Users receive incorrect answers with no recourse beyond retrying, and operators have no visibility into failure modes.
Value: Surfaces systematic failures (hallucinated config keys, wrong versions) and builds user trust through visible accountability.
Implementation:
- Add a
discord.ui.Viewwith 👍 / 👎 buttons to/askand/analyzeresponses. - On interaction, log
(user_id, question, answer, rating, timestamp)to a configurable feedback channel (Discord webhook or dedicated#bot-feedbackchannel). - No database required; structured logging is sufficient for triage.
Problem: Crash reports shared as file attachments (.txt, .log) are not auto-detected. Users must manually run /analyze after uploading a crash report, unlike mclo.gs/pastebin links which trigger automatic pattern matching.
Implementation:
- Extend
on_messageincogs/log_analyzer.pyto check file attachments for---- Minecraft Crash Report ---- - Auto-trigger analysis (same flow as mclo.gs link detection)
- Add
👀reaction and offer "Analyze with AI" button - Crash reports are one of the most frequently shared artifacts in support channels
Problem: All bot features are globally active. Server admins cannot disable passive log detection in off-topic channels or restrict /ask to specific support channels.
Implementation:
- Admin command:
/config <feature> <channel> enable|disable - Features:
log_detection,spark_detection,passive_errors,ask,analyze - Store configuration in
data/channel_config.json - Check config in each cog's
on_messagelistener and command checks - Default: all features enabled in all channels (backward compatible)
validate_config() in config.py only asserts BOT_TOKEN. If OPENROUTER_API_KEY is missing, all AI commands silently respond with "
Fix: Add a warning log in validate_config() (or in setup() of AI cogs) if the key is absent:
if not OPENROUTER_API_KEY:
logger.critical("OPENROUTER_API_KEY is not set — all AI commands will be disabled")Only RateLimitError is caught explicitly on LLM calls. asyncio.TimeoutError, aiohttp.ServerConnectionError, and other transient failures return an error immediately with no retry. Under brief API instability the bot fails every request.
Fix: Wrap LLM calls in a simple retry helper with exponential backoff (2–3 attempts, base delay 1s):
async def _llm_call_with_retry(self, **kwargs):
for attempt in range(3):
try:
return await self.client.chat.completions.create(**kwargs)
except (RateLimitError, openai.APIConnectionError, openai.APITimeoutError) as e:
if attempt == 2:
raise
await asyncio.sleep(2 ** attempt)/ask and /chat set the embed title to the raw user question (f'❓ {question}'). Discord embed titles are capped at 256 characters and truncate silently, which looks broken for long questions.
Fix: Truncate the question before embedding it in the title:
title_text = question if len(question) <= 240 else question[:237] + '…'
title = f'❓ {title_text}' if i == 0 else ''main.py already configures logging.basicConfig() with a proper format, but some cogs may still use print() for output. Additionally, log messages lack structured context (guild_id, user_id, command name), making production debugging harder.
Fixes:
- Audit all cogs for stray
print()calls → replace withlogger.info/warning/error - Add structured context to log messages where relevant:
logger.info("Processing /ask from user=%s guild=%s", interaction.user.id, interaction.guild_id)
- Add
LOG_LEVELenv var toconfig.py(default:INFO), apply inmain.py:logging.basicConfig(level=getattr(logging, LOG_LEVEL, logging.INFO), ...)
The current 18 patterns in responses/errors.py cover common cases. The following patterns are missing and would improve passive detection coverage:
| Pattern | Category | Notes |
|---|---|---|
OutOfMemoryError: Metaspace |
Memory | Different from heap OOM — indicates too many plugins loaded |
ConcurrentModificationException |
World | Usually from async world access or corrupted chunks |
java.lang.StackOverflowError |
Performance | Recursive chunk loading or plugin logic loops |
ModResolutionException |
Forge/Fabric | Mod dependency resolution failures |
IncompatibleModSetException |
Forge/Fabric | Mod version conflicts |
ExceptionInInitializerError |
Plugin | Plugin static initialization failures |
BackendConnectionException |
Proxy | BungeeCord/Velocity backend connection failures |
io.netty.handler.codec.DecoderException |
Network | Packet decoding errors, often from version mismatch |
Implementation: Add entries to _raw_patterns in responses/errors.py with pt-BR response templates. Follow existing ordering convention: specific patterns before generic ones within each category.
Problem: The only test file (tests/test_spark_parser.py, 954 lines) duplicates the entire Spark parser and has no assertions, fixtures, or test runner. There is no pytest in requirements.txt.
Implementation:
- Add
pytestandpytest-asynciotorequirements.txt - Refactor
tests/test_spark_parser.pyinto proper pytest tests with assertions - Add tests for:
cogs/utils.py—truncate_safeandsplit_responseedge cases (empty string, exact boundary, unicode)responses/errors.py— pattern matching against known log snippetscogs/plugin_apis.py— URL construction and response parsing with mock datacogs/spark_parser.py—build_summary()andbuild_detail()with fixture JSON
- Add a
conftest.pywith shared fixtures (mock Spark JSON, sample logs)
Problem: No containerization. Deployment requires manual Python setup, which is a barrier for non-Python operators.
Implementation:
- Create
Dockerfile:FROM python:3.11-slim WORKDIR /app COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt COPY . . CMD ["python", "main.py"]
- Create
docker-compose.ymlwith volume mount fordata/(vector store persistence) and.envpassthrough - Document in
SETUP.md
Add a GitHub Actions workflow:
- Lint —
ruff check .orflake8 - Test —
pytest(once 6.1 is implemented) - Build — verify
pip install -r requirements.txtsucceeds - Trigger on push to
mainand on PRs - Optional: automated Docker image build and push to GHCR