AI-powered Discord bot for the Miners' Refuge community. Helps Minecraft server admins troubleshoot logs, diagnose performance with Spark, search documentation, and find plugins — all from Discord.
Built with discord.py + RAG (Retrieval-Augmented Generation) over multiple documentation sources.
- Features
- Slash Commands
- Passive Features
- AI & RAG System
- Documentation Sources
- Project Structure
- Installation
- Configuration
- Running the Bot
- Docker
- Tests
| Category | Capability |
|---|---|
| Q&A (RAG) | Answers questions grounded in indexed docs (Miners' Refuge, PaperMC, PurpurMC, Spark) |
| General Chat | Open-ended assistant with web search, server/history context awareness |
| Log Analysis | AI analysis of mclo.gs/pastebin links, file uploads, and screenshots |
| Error Detection | Instant passive replies for 20+ known Minecraft error patterns |
| Spark Profiler | Full report parsing, bottleneck diagnosis, platform-aware recommendations |
| Plugin Search | Concurrent search across Modrinth, Hangar, SpigotMC |
| Server Tools | JVM flags, server status checks, docs changelog |
| Channel Summaries | /resumo and natural-language "o que perdi?" — map-reduce summaries of a channel or thread over any period, with jump links and permission checks |
| Persistent Memory | The bot remembers: atomic facts about the server and its members, auto-injected into every conversation, with inline LLM writes, pinning, dedupe, full audit history and admin control |
All AI responses are in Brazilian Portuguese and formatted for Discord embeds.
Ask any Minecraft server administration question. Uses an agentic RAG loop — the LLM automatically calls search_docs, search_plugins, memory_search, search_history, and context tools to ground its answer in real documentation.
- Model:
CHAT_MODEL(default:qwen/qwen3.6-plus) - Tools available:
search_docs,search_plugins,memory_search,memory_write,memory_about,get_channel_history,get_guild_info,search_history,get_message_context,find_user,sql_history,add_reaction - Features: Image/screenshot analysis, 5-minute ephemeral prompt+report cache, paginated multi-embed output with source links, reply to continue conversation (30 min TTL, 200 conversations), cooldown per user
- Follow-up: Reply to the bot's response to continue the conversation with full history
General-purpose assistant (same agentic loop as /ask but with web search).
- Tools:
web_search,web_extract(via Tavily), plus the same context/history tools as/ask, the scheduler tools andsummarize_channel(see/resumo) - Reactions:
add_reactionlets the model react with a Unicode emoji or a:nome:custom emoji of the server, only to messages in the current channel (the triggering message by default) and at most 3 per answer.REACTION_TOOL_ENABLED=falseremoves the tool - Web search: Supports
search_depth(basic/advanced),time_range, domain filtering - Reply follow-up and @mention mode: Mention the bot (
@QuillBot <question>) to chat without a slash command — same rate-limit and conversation handling as/chat - Set
CHAT_MENTION_ENABLED=falseto disable mention mode
Summarize what happened in a channel or thread while you were away.
canal— text/voice channel or thread (default: the current one). Archived threads work tooperiodo— default since your last message in that channel (your own messages from the last 10 minutes don't count, so "voltei!" +/resumostill covers the gap); also30m/6h/2d,hoje,ontemor a date like24/09/2026 08:00(Brasília time). Autocompletes common presetsfoco— optional topic to focus on ("backup", "o lag do servidor")- Output: ephemeral embed with Tópicos, Decisões e soluções and Pendências, each item linked (↗) to the message it came from, plus a private 🔔 Mencionaram você list (messages that mentioned or replied to you). 📢 Publicar no canal posts it publicly (without your mentions); replying to the published summary continues it as a
/chatconversation - How it works (
cogs/summary.py):- Messages are read straight from Discord, newest first, until the period start,
SUMMARY_MAX_MESSAGESorSUMMARY_MAX_DAYS. Bot answers are included, since in support channels they're often the fix - Ranges that fit
SUMMARY_SEGMENT_CHARStake one LLM call. Longer ones are split at conversation gaps (≥20 min of silence), each part is summarized in parallel, and the notes are merged - The model cites messages as
[msg_id=…]; code turns them into jump links and drops any ID that wasn't in the fetched range - Mentions are rewritten as plain names, so summaries never ping anyone
- Messages are read straight from Discord, newest first, until the period start,
- Permissions: only channels that you can view and read the history of (private threads also require membership or Manage Threads), in the current server
- Natural language: the same engine is the
summarize_channeltool in/chat, @mention, reply follow-ups and scheduled tasks. For example@QuillBot o que perdi?,resume o que rolou no #suporte desde ontemor a daily scheduled "resumo do #geral". The chat model turns "hoje"/"desde ontem" into timestamps using its clock context and receives a finished summary, not raw messages, so the 6000-char tool-result cap and conversation replay stay small - Set
SUMMARY_ENABLED=falseto remove the command and the tool
AI-powered log/crash-report analysis.
- Inputs (at least one required):
log_link— mclo.gs or pastebin.com URL (parsed to raw API endpoint)log_file—.logor.txtattachment (auto-uploaded to mclo.gs for sharing, streaming read up to 5 MB)image— screenshot of an error/log (vision model)
- Processing: Sanitizes non-printable chars, extracts relevant lines for large logs (keeps 60 header lines + error/warning/exception lines, truncates to 12k chars), sends to LLM with a structured diagnostic prompt
- Output: Sections — Resumo, Erros Encontrados, Avisos, Recomendações (only non-empty sections shown)
- Button fallback: When a pasted log has no regex pattern match, a "🔍 Analisar com IA" button is offered to trigger the same analysis
Analyze a Spark Profiler report with deep AI diagnosis.
- Input: Full
https://spark.lucko.me/<code>URL or bare report code - Parser (
spark_parser.py): Fetches fromspark-json-service.lucko.me, extracts server identity, JVM flags, TPS/MSPT, GC, CPU hotspots (chain-collapsed call tree withself_pct), world/entity stats, configs, game rules, plugins - LLM flow: Injects a compact summary (~1.7k chars) as synthetic prior exchange, follows an 8-step diagnostic protocol (TPS → MSPT → lag spike detection → waitForNextTick → GC → hotspots → config → grounded recommendation), uses
SPARK_MODEL(default:google/gemini-2.5-pro) - Spark tools:
get_spark_detail(sections:hotspots,jvm,profiler,plugins,world,game_rules,configs:<file>) andget_config_key(single-key lookup) - Platform awareness: Detects Paper/Purpur/Spigot/Forge/Fabric/Vanilla and never recommends configs that don't exist for that platform
- Passive detection: Any
spark.lucko.melink in chat triggers a "🔥 Analisar com IA" button
Re-index documentation vectors.
- Without
source: full reindex of allDOC_SOURCES - With
source: reindex a single source by label (e.g.PaperMC) - Shows
⏳ Indexando...guard — blocks/ask//docsqueries during indexing - Per-source commit SHA tracking for granular updates
The bot maintains long-term memories: atomic one-sentence facts about the server (guild-wide) or about individual users (scoped, privacy-guarded). Relevant memories are automatically injected into every /ask, /chat and @mention turn — the LLM doesn't need to search first.
| Subcommand | Who | Description |
|---|---|---|
/memory list [kind] |
everyone | Paginated list (users see guild memories + their own; admins see all) |
/memory show <id> |
everyone | Memory details (own or guild-wide only); admins also see history IDs |
/memory edit <id> |
Admin | Edit content / importance |
/memory delete <id> |
Admin | Archive (always reversible — no hard deletes) |
/memory restore <id> |
Admin | Restore any version by history ID |
/memory pin / /memory unpin <id> |
Admin | Toggle pinned (always injected; admins bypass the pin limit) |
/memory whoami |
everyone | See everything the bot remembers about you |
/memory forget-me |
everyone | Hard-delete all memories about you (irreversible) |
/memory forget <user> |
Admin | Hard-delete all memories about a user |
/memory export |
Admin | Full JSON export |
How it works:
- LLM tools —
memory_search(hybrid lexical→semantic recall with importance/recency re-ranking),memory_write(create/update/forget, one mutation path, dedupe-guarded),memory_about(compact per-user profile) - Auto-injection — every turn embeds the current question and selects the top memories (cosine + importance + participant boost); pinned memories always ride along (importance-ranked, capped). User memories only participate when their subject is in the conversation (speaker, @mention, reply target)
- Pinning — bot can pin too (
pinned=trueinmemory_write, capped atMEMORY_PIN_LIMIT/guild, audited viapinned_by); admins pin/unpin without the cap - Dedupe — creates are refused when semantically similar (cosine ≥
MEMORY_DEDUPE_THRESHOLD) to an existing memory in the same scope; the LLM is told toupdateinstead - Single mutation path — every write funnels through
MemoryStore, which applies the change and records a before/after snapshot intomemory_historyin one SQLite transaction - Rate guard — bot writes capped at
MEMORY_BOT_WRITE_LIMIT/hour per guild;MEMORY_MAX_PER_SUBJECTactive memories per scope - Transparency channel — with
MEMORY_LOG_CHANNEL_IDset, every mutation posts an embed with a ↩️ Reverter button (admin-only, same guild) - Privacy —
/memory whoami+/memory forget-meare self-service; purges hard-delete memories and their history - One-time migration — existing
lore.dbentries import automatically on first start (curated → pinned, pending skipped;lore.dbis kept on disk untouched) - Tool errors degrade gracefully — memory failures return error strings to the LLM instead of aborting
/ask//chat - Set
MEMORY_ENABLED=falseto remove the tools and prompt guidance entirely
| Command | Description | Input |
|---|---|---|
/docs [query] |
Search indexed docs (RAG) or show docs link. Supports all sources + source label prefix. | query (optional) |
/plugin <name> |
Search Modrinth, Hangar, SpigotMC concurrently. Paginated view (one source per page). Autocomplete via Modrinth. | name (required) |
/status <ip> |
Check Minecraft server status via mcsrvstat.us. Shows MOTD, player count, version. |
ip (hostname or IP, optional port) |
/changelog |
Last 5 commits to MinersRefuge/docs (SHA, message, author, date) via GitHub API. |
— |
/flags <ram> |
Generate Aikar's JVM flags for a given RAM amount in MB. Auto-tunes G1GC thresholds at >= 12 GB. Validates 512–65536 MB range. | ram (MB, integer) |
/plov |
Embed: PLOV info needed to choose hosting — Plano/Players, Localização, Orçamento, Versão. | — |
/plano |
Embed: info needed to recommend a plan — Version, Players, Mods/Plugins, Gamemode. | — |
/health |
(Admin) Bot diagnostics: uptime, concurrent API pings (OpenRouter, Tavily, GitHub, mclo.gs, Modrinth) with latency, vector store stats, web search status, conversation cache usage, WebSocket latency. | — |
/sync |
(Admin) Re-sync slash commands with Discord. | — |
/help |
Auto-generated list of all registered slash commands + passive feature summary. | — |
The bot monitors every non-bot message via on_message listeners (load order: log_analyzer → history_rag → commands → spark → docs_rag).
- Triggers on:
https://mclo.gs/<id>andhttps://pastebin.com/<id> - Behavior: Reacts with 👀, fetches raw content (streaming, 5 MB cap), checks against known error patterns
- If matched: Replies instantly with the pattern's troubleshooting response
- If no match + AI available: Replies with "🤔 Não reconheci nenhum erro" + "🔍 Analisar com IA" button
- Triggers on:
.log/.txtattachments - Behavior: Reads file (streaming), checks error patterns, uploads to
mclo.gs(POST /1/log), replies with the mclo.gs link (or a fallback message). If large (> 5 MB), notes truncation. Offers AI analysis button when no pattern matches.
25 pre-compiled regex patterns — first match wins (specific before generic):
| Category | Examples |
|---|---|
| Plugin | Ambiguous plugin name, missing dependency, Error occurred while enabling, Could not pass event |
| Startup / JAR | Unable to access jarfile, Current Java is X but we require Y, Unsupported Java, JNI error |
| EULA | You need to agree to the EULA |
| Memory | OutOfMemoryError, Can't keep up! Is the server overloaded? |
| Port / Network | FAILED TO BIND TO PORT, Perhaps a server is already running on that port |
| World / Data | Failed to load chunk, Region file is truncated, Session lock is no longer valid |
| Permissions | <player> was denied the command |
| Version | Outdated server, Outdated client |
| Crash | This crash report has been saved to, ---- Minecraft Crash Report ---- |
| Generic | The received string length is longer than maximum allowed, Connection throttled |
Regex https://spark.lucko.me/<code> → replies "🔥 Relatório Spark detectado!" with analysis button.
Reply to any /ask, /chat, or /spark bot response to continue the thread. The bot replays up to 16 history turns (for /ask//spark) with the Spark report carried forward. Per-user cooldown enforced. Supports image attachments in follow-ups.
Mention the bot with a question (@QuillBot como otimizar meu servidor?) — equivalent to /chat with the same tool loop and conversation storage.
Background RAG over the entire Discord server's message history:
- Backfill on startup: reads all accessible text channels/threads (oldest-first, configurable limit)
- Live ingestion: Every new message is chunked with a 5-message sliding window for local context, embedded, and persisted per-guild (
data/history/<guild_id>.json + .npy) - Dedup + deletion: Tracks
msg_id → index, handleson_message_delete - Search tool:
search_history/search_docsequivalent available to the LLM in all agentic loops - SQL tool:
sql_history— LLM-written analytical SELECTs over the history DB (exact counts, GROUP BY, mentions, replies, FTS5, REGEXP), sandboxed four ways: single-SELECT validation + guild scoping, read-only URI +PRAGMA query_only,set_authorizer(writes/ATTACH/PRAGMA denied,chunks.embeddingnever emitted), and a progress-handler timeout (HISTORY_SQL_TIMEOUT_SECONDS). Disable withHISTORY_SQL_TOOL_ENABLED=false; rows/output caps viaHISTORY_SQL_MAX_ROWS.
- Provider: Switchable via
EMBEDDING_PROVIDER—openai(remote, default) orlocal(sentence-transformers, e.g.all-MiniLM-L6-v2) - Models:
EMBEDDING_MODEL(defaultqwen/qwen3-embedding-8b),RERANK_MODEL(defaultcohere/rerank-4-fast) - Reranking: Enabled when
RERANK_ENABLED=trueandOPENAI_BASE_URLis OpenRouter — reranks top3×top_kcandidates via the OpenRouter/rerankendpoint, falls back to cosine similarity - Storage:
data/vectors.json(metadata) +data/vectors.npy(float32 binary embeddings) — migrated automatically from old JSON-embedded format - Chunking: Markdown split by headings (
##/#), then by1500-char paragraph windows; title extracted from frontmatter or first#
Shared loop used by /ask, /chat, and Spark analysis:
while tool_calls and rounds < MAX_TOOL_ROUNDS:
→ display status label on deferred interaction
→ execute each tool (deduplicate sources via dedup_key)
→ append tool results (truncated to 6000 chars)
→ final answer (max_tokens=`LLM_MAX_TOKENS`, default 8192)
→ if the answer comes back empty (e.g. a reasoning model spent the whole token budget thinking): one retry without tools + conciseness nudge, then fallback message
→ paginated embeds + source pages
- Parallel tool calls per round when the LLM requests multiple
Each turn captures the exact rendered user message plus the serialized internal loop (assistant tool calls, tool results, final answer). On follow-ups, the most recent CONVERSATIONS_TRAJECTORY_TURNS turns are replayed verbatim — the model literally sees its own prior tool calls and their results instead of a distilled Q/A pair. Older turns (and turns captured before this feature, or whose trajectory was dropped as invalid/oversized) fall back to the compact attributed rendering.
Replay windows advance in CONVERSATIONS_TRAJECTORY_STEP-turn steps (hysteresis), so the replayed prefix only changes at deliberate collapse boundaries — not sliding every turn — which is what keeps the provider prefix cache effective. Explicit cache_control breakpoints mark the system prompt and the last replayed history message for Anthropic/Gemini via OpenRouter; OpenAI-compatible endpoints ignore them and cache automatically. Captured trajectories are sanitized on persist (no reasoning/thinking payloads, tool-call/result pairing enforced).
COOLDOWN_RATE / COOLDOWN_PER (default 1 per 30s) per user on /ask, /chat, /analyze. Follow-up replies share a TTLCache cooldown. Exceeding returns an ephemeral ⏳ Aguarde Xs message.
Every AI call receives a <contexto> block with user (display name, account age, join date, roles), guild (name, member count, channels, roles), channel, and temporal (BRT + UTC) context.
When CHANNEL_CONTEXT_MESSAGES > 0 (default 10), /ask, /chat, @mention and reply follow-ups also receive a <mensagens_recentes_do_canal> block — the latest N channel messages in chronological order, same format as the get_channel_history tool. Image attachments in that context are persisted and sent as vision parts, up to the four-image per-message budget (current-message images take priority, then the newest channel images); filenames remain visible in the text context. Only images posted within CHANNEL_CONTEXT_IMAGE_MAX_AGE_MINUTES (default 60) are sent, and CHANNEL_CONTEXT_IMAGES_ENABLED=false turns channel images off entirely (use it with text-only CHAT_MODELs). Attachments are downloaded once and reused by attachment id. Stickers ([figurinha:Nome]) and GIF/image link embeds (Tenor, Giphy, direct image links; [gif: …]) count as images too; animated GIFs and stickers are sent as one numbered sheet of four evenly spaced frames, and Tenor/Giphy links use a small animated rendition with the still preview as fallback. Lottie stickers are shown by name only. Custom emojis appear as :nome: in the text, and each message's custom emojis are drawn on one labelled sheet that only uses image slots left over after real media. The same media from the triggering message (and the message it replies to) is sent as the user's own images. Reactions are listed on each line before msg_id as [reações: 👍 3 (Ana, Bruno +1)], with up to REACTION_USERS_LIMIT reactor names per reaction. The triggering message is excluded; set 0 to disable. Reply follow-ups use the captured gap block (prior_context, see CONVERSATIONS_GAP_MESSAGES) instead of this recent-channel window whenever a gap was captured, including images from those gap messages.
The LLM message list is ordered so conversation follow-ups reuse a cached prefix: [system (persona + static conversation summary)] → [replayed history turns] → [current message]. All per-request blocks — <contexto> (clock), the semantically-selected memory block, and the recent channel window — ride on the final user message instead of the system prompt, so the system prompt + history stay byte-identical across turns and providers can prefix-cache them (typically 80–95% input-token savings on cached prefixes).
Configured in config.py:DOC_SOURCES. Each entry:
| Key | Required | Description |
|---|---|---|
repo |
Yes | GitHub owner/name |
branch |
Yes | Branch to index |
base_url |
Yes | Docs website base URL |
label |
Yes | Human-readable label shown in results |
summary |
No | Path to mdBook SUMMARY.md for discovery |
path_prefix |
No | Only index files under this prefix |
url_strip_prefix |
No | Strip prefix when building website URLs |
max_files |
No | Cap per source (default 200) |
| Source | Label | Discovery | Base URL |
|---|---|---|---|
MinersRefuge/docs |
Miners' Refuge | SUMMARY.md |
https://docs.minersrefuge.com.br |
PaperMC/docs |
PaperMC | Tree API (src/content/docs/paper/admin/) |
https://docs.papermc.io |
PurpurMC/PurpurDocs |
PurpurMC | Tree API (mkdocs/purpur/) |
https://purpurmc.org/docs/purpur/ |
lucko/spark-docs |
Spark | Tree API (docs/) |
https://spark.lucko.me/docs/ |
Periodic reindex every REINDEX_INTERVAL_HOURS (default 6h) via composite MD5 of all latest commit SHAs.
QuillBot/
├── main.py # Entrypoint — validates config, loads cogs in order, syncs slash commands
├── config.py # Centralized env-var config, DOC_SOURCES, feature flags
├── requirements.txt
├── Dockerfile
├── docker-compose.yml
├── cogs/
│ ├── commands.py # /chat, /docs, /help, /flags, /plov, /plano, /health, /sync + @mention & reply follow-ups
│ ├── docs_rag.py # RAG pipeline — indexing, search, reranking, /ask, /reindex, agentic loop + Spark diagnosis
│ ├── history_rag.py # Server-wide history RAG — backfill, live ingestion, per-guild vector stores
│ ├── log_analyzer.py # Passive log detection, pattern matching, /analyze, file upload to mclo.gs
│ ├── memory.py # Persistent memory — auto-injection, LLM write tools, admin commands, audit history, log channel, lore migration
│ ├── plugins.py # /plugin, /status, /changelog — plugin search & server status
│ ├── summary.py # /resumo + summarize_channel tool — period parsing, fetch, map-reduce summary, permission checks
│ ├── spark.py # /spark command + passive spark.lucko.me detection
│ ├── spark_parser.py # Spark JSON parsing, summary/detail builders, call-tree rendering
│ ├── plugin_apis.py # Shared Modrinth/Hangar/SpigotMC API helpers
│ ├── tavily_tools.py # Tavily web_search / web_extract tool definitions & execution
│ └── utils.py # Shared: truncate_safe, split_response, PaginatedEmbedView, context builders, run_tool_loop
├── responses/
│ └── errors.py # 25 compiled Minecraft error regex patterns + pt-BR responses
├── tests/ # pytest suite, one file per module/feature (see Tests)
└── data/
├── vectors.json/.npy # Docs RAG vector store
├── memory.db # Persistent memory (memories + history)
├── lore.db # Legacy lore encyclopedia (kept as migration source)
└── history/ # Per-guild history stores
Requires Python 3.11+
git clone https://github.com/Zeptiny/QuillBot.git
cd QuillBot
pip install -r requirements.txt- Create a bot at https://discord.com/developers/applications
- Under Bot → Privileged Gateway Intents enable Message Content Intent
- Invite with scopes
bot+applications.commandsand permissions: Send Messages, Embed Links, Read Message History, Add Reactions, Use Application Commands
Copy and fill in your environment file:
cp .env.example .env # if available, otherwise create .env manually| Variable | Description |
|---|---|
BOT_TOKEN |
Discord bot token — bot will not start without it |
OPENROUTER_API_KEY |
API key for OpenRouter (or any OpenAI-compatible provider). Required for all AI commands. Also aliased as OPENAI_API_KEY / LLM_API_KEY |
| Variable | Default | Description |
|---|---|---|
OPENAI_BASE_URL |
https://openrouter.ai/api/v1 |
OpenAI-compatible endpoint. Aliases: LLM_BASE_URL, OPENROUTER_BASE_URL |
CHAT_MODEL |
qwen/qwen3.6-plus |
Used for /ask, /chat, /analyze |
LLM_MAX_TOKENS |
8192 |
Max completion tokens for the tool loop. Reasoning models count thinking tokens against this |
SPARK_MODEL |
google/gemini-2.5-pro |
Used only during active Spark sessions |
EMBEDDING_MODEL |
qwen/qwen3-embedding-8b |
Doc/history embedding model |
RERANK_MODEL |
cohere/rerank-4-fast |
Reranking model |
EMBEDDING_PROVIDER |
openai |
openai (remote) or local (sentence-transformers) |
LOCAL_EMBEDDING_MODEL |
sentence-transformers/all-MiniLM-L6-v2 |
Model when EMBEDDING_PROVIDER=local |
LOCAL_EMBEDDING_DEVICE |
cpu |
Device for local embeddings |
RERANK_ENABLED |
true |
Enable reranking (only when base URL is OpenRouter) |
RERANK_PROVIDER |
auto |
auto (remote on OpenRouter, local elsewhere), local (cross-encoder) or remote (OpenRouter /rerank) |
LOCAL_RERANK_MODEL |
cross-encoder/ms-marco-MiniLM-L-6-v2 |
Cross-encoder used when the rerank provider is local |
LOCAL_RERANK_DEVICE |
cpu |
Device for the local reranker (defaults to LOCAL_EMBEDDING_DEVICE) |
| Variable | Default | Description |
|---|---|---|
WEB_SEARCH_ENABLED |
true |
Enable Tavily web search |
TAVILY_API_KEY |
— | Required when web search is enabled |
CHAT_MENTION_ENABLED |
true |
Enable @mention chat mode |
CHANNEL_CONTEXT_MESSAGES |
10 |
Latest channel messages auto-injected as text and vision context into /ask, /chat, @mention and reply follow-ups (0 disables) |
CHANNEL_CONTEXT_IMAGES_ENABLED |
true |
Send images attached to channel-context and follow-up gap messages as vision parts (false for text-only models) |
CHANNEL_CONTEXT_IMAGE_MAX_AGE_MINUTES |
60 |
Only channel-context images newer than this are sent (0 = no age limit) |
CHANNEL_CONTEXT_REACTIONS_ENABLED |
true |
Show reactions on channel-context lines (history tool, recent window, follow-up gaps, message context) |
REACTION_TOOL_ENABLED |
true |
Give the chat model the add_reaction tool (current channel only, 3 per answer) |
REACTION_USERS_LIMIT |
5 |
Reactor names shown per reaction (one API call per reaction, cached while the count is unchanged; 0 = counts only) |
HISTORY_ENABLED |
true |
Enable server history RAG |
MEMORY_ENABLED |
true |
Enable persistent memory (cog not loaded when false) |
LOG_LEVEL |
INFO |
Python logging level |
| Variable | Default | Description |
|---|---|---|
HISTORY_VECTOR_STORE_DIR |
data/history |
Per-guild history storage dir |
HISTORY_DB_PATH |
data/history/history.db |
SQLite history DB (defaults to <HISTORY_VECTOR_STORE_DIR>/history.db) |
HISTORY_WINDOW_SIZE |
5 |
Sliding window of prior messages per chunk |
HISTORY_WINDOW_OVERLAP |
1 |
Overlap between consecutive chunks |
HISTORY_BACKFILL_LIMIT |
(none) | Max messages to backfill per channel per run (unset = all); restarts resume after the newest indexed message |
HISTORY_MAX_MSG_LENGTH |
800 |
Max chars per message in history chunks |
HISTORY_EXCLUDE_BOTS |
true |
Exclude bot messages from history |
HISTORY_INGEST_BATCH_SIZE |
10 |
Messages per embedding batch during live ingestion |
HISTORY_INGEST_FLUSH_SECONDS |
2.0 |
Max delay before a partial batch is flushed |
HISTORY_SNAPSHOT_INTERVAL |
300 |
Seconds between vector-store snapshots to disk |
HISTORY_QUERY_CACHE_SIZE |
200 |
LRU query result cache size |
HISTORY_DEDUPE_WINDOW_MINUTES |
10 |
Adjacent results in the same channel treated as one window and deduplicated (0 disables) |
HISTORY_RERANK_ENABLED |
true |
Rerank history search results |
HISTORY_RERANK_PROVIDER |
auto |
Rerank provider for history search (defaults to RERANK_PROVIDER) |
HISTORY_RERANK_MODEL |
LOCAL_RERANK_MODEL |
Remote rerank model for history search |
HISTORY_TIME_DECAY_LAMBDA |
0.0 |
Exponential decay on recency for recent sorting (0 disables) |
HISTORY_HYBRID_WEIGHT_SEMANTIC |
0.65 |
Semantic weight of the semantic/keyword hybrid blend |
HISTORY_HYBRID_WEIGHT_KEYWORD |
0.35 |
Keyword (FTS5) weight of the hybrid blend |
HISTORY_RRF_K |
60 |
Reciprocal-rank-fusion constant |
HISTORY_SQL_TOOL_ENABLED |
true |
Enable the read-only sql_history LLM tool |
HISTORY_SQL_TIMEOUT_SECONDS |
30 |
Query timeout enforced by the SQLite progress handler |
HISTORY_SQL_MAX_ROWS |
200 |
Max rows returned / shown to the LLM |
| Variable | Default | Description |
|---|---|---|
IMAGES_DIR |
data/images |
Where downloaded images are persisted |
IMAGE_MAX_SIDE |
640 |
Max dimension (px) of re-encoded JPEGs |
IMAGE_JPEG_QUALITY |
80 |
JPEG quality of re-encoded images |
IMAGE_RETENTION_SECONDS |
86400 |
Stored images swept after this many seconds |
CONVERSATIONS_IMAGE_TURNS |
3 |
Inline stored images only for the most recent N history turns (older become text markers) |
| Variable | Default | Description |
|---|---|---|
MEMORY_ENABLED |
true |
Enable the memory cog (cog is not loaded when false) |
MEMORY_DB_PATH |
data/memory.db |
SQLite path for memories + history |
MEMORY_LOG_CHANNEL_ID |
(none) | Channel ID for mutation log embeds (Revert button). Unset = console logging only |
MEMORY_BOT_WRITE_LIMIT |
20 |
Max bot (memory_write) writes per hour per guild |
MEMORY_INJECT_LIMIT |
5 |
Semantically selected memories injected per turn (pinned ride along separately) |
MEMORY_PIN_LIMIT |
12 |
Max pinned memories per guild (bot is capped; admins bypass) |
MEMORY_MAX_PER_SUBJECT |
200 |
Max active memories per scope (guild or user) |
MEMORY_SEMANTIC_MIN_SCORE |
0.35 |
Min cosine for semantic search/injection hits |
MEMORY_DEDUPE_THRESHOLD |
0.85 |
Cosine above which a create is refused as duplicate |
LORE_DB_PATH |
data/lore.db |
Legacy lore DB — one-time migration source for memory.db |
| Variable | Default | Description |
|---|---|---|
SUMMARY_ENABLED |
true |
Enable /resumo and the summarize_channel tool (cog not loaded when false) |
SUMMARY_MODEL |
CHAT_MODEL |
Model for the summary calls; a cheaper model is usually fine |
SUMMARY_MAX_MESSAGES |
1000 |
Max messages read per summary (the newest are kept) |
SUMMARY_MAX_DAYS |
7 |
Max lookback in days |
SUMMARY_SEGMENT_CHARS |
24000 |
Rendered message chars per LLM call; longer ranges are summarized map-reduce style |
| Variable | Default | Description |
|---|---|---|
VECTOR_STORE_PATH |
data/vectors.json |
Docs vector store path |
REINDEX_INTERVAL_HOURS |
6 |
Hours between automatic reindex checks |
| Variable | Default | Description |
|---|---|---|
CONVERSATIONS_DB_PATH |
data/conversations.db |
SQLite path for persisted conversations |
CONVERSATIONS_TTL_SECONDS |
1800 |
Conversation expiry after last activity |
CONVERSATIONS_MAX_STORED |
200 |
Max conversations kept per flow (chat/ask) |
CONVERSATIONS_MAX_TURNS |
24 |
Max turns stored per conversation |
CONVERSATIONS_HISTORY_TURNS |
16 |
Turns replayed to the LLM per request |
CONVERSATIONS_GAP_MESSAGES |
20 |
Max channel messages captured between two bot-directed turns as text and vision prior_context (0 disables) |
CONVERSATIONS_TRAJECTORY_ENABLED |
true |
Capture and replay the internal LLM trajectory (tool calls/results + exact user message) per turn |
CONVERSATIONS_TRAJECTORY_TURNS |
6 |
Most recent N turns replayed verbatim with their trajectory (older turns use the compact Q/A rendering) |
CONVERSATIONS_TRAJECTORY_STEP |
3 |
Window advances in steps of N turns — prefix-cache hysteresis so the replayed prefix only changes at deliberate boundaries |
CONVERSATIONS_TRAJECTORY_MAX_CHARS |
120000 |
Hard budget on the replayed trajectory payload; turns are collapsed to compact rendering when exceeded |
| Variable | Default | Description |
|---|---|---|
COOLDOWN_RATE |
1 |
Allowed uses per cooldown window |
COOLDOWN_PER |
30 |
Cooldown window in seconds (per user) |
MAX_CONTENT_SIZE |
5242880 |
Max bytes fetched from paste services (5 MB) |
MAX_LOG_CONTEXT |
12000 |
Max chars sent to LLM for log analysis |
Implemented in api_logger.py — patches the aiohttp/httpx transports so every outbound HTTP request (Discord REST, LLM API, Tavily, GitHub, plugin APIs) and every inbound Discord interaction is appended as one JSON line to a rotating file. See SETUP.md for sample lines and privacy notes.
| Variable | Default | Description |
|---|---|---|
API_REQUEST_LOG_ENABLED |
true |
Master switch |
API_REQUEST_LOG_PATH |
data/api_requests.log |
Log file path (parent dirs created) |
API_REQUEST_LOG_MAX_BYTES |
10485760 |
Rotate after this many bytes |
API_REQUEST_LOG_BACKUPS |
5 |
Rotated files to keep |
API_REQUEST_LOG_DISCORD |
true |
false omits outbound Discord REST traffic (inbound interactions are logged regardless) |
API_REQUEST_LOG_CONSOLE |
false |
Also mirror log lines to stdout / docker logs |
API_REQUEST_LOG_BODY |
openai,tavily |
Services whose request/response bodies are logged (all/none accepted) |
API_REQUEST_LOG_BODY_MAX_CHARS |
20000 |
Per-body truncation limit |
python main.pyOn startup the bot:
- Validates
BOT_TOKEN(exits if missing; warns ifTAVILY_API_KEYmissing while web search is enabled) - Loads cogs in order:
log_analyzer→history_rag→commands→plugins→spark→docs_rag→memory(first run imports legacylore.dbautomatically) - Loads cached vectors from disk or indexes all doc sources
- Starts periodic reindex loop + background history backfill
- Syncs slash commands with Discord
Cog load order matters —
log_analyzer'son_message(pattern matching) runs beforedocs_rag's (follow-up replies). Do not reorder without reviewing listener interactions.
docker compose up --build -d- Uses
python:3.11-slim, persistsdata/via volume mount, reads config from.env.
The suite lives in tests/ and runs with pytest. It needs no Discord token, network or models: Discord objects, the LLM client and embeddings are faked, and SQLite databases go in temporary directories.
pip install -r requirements-dev.txt
pytest # whole suite
pytest tests/test_summary.py # one file
pytest -k reindex # tests whose name matches| File | Covers |
|---|---|
test_api_logger.py |
Request logging for aiohttp/httpx, redaction, body capture, inbound interactions (runs in its own process) |
test_channel_context_images.py |
Images from automatically fetched channel context |
test_conversation_participants.py |
Participant scope for shared conversations and memory writes |
test_docs_rag.py |
Docs /reindex: partial and full reindex, failed sources kept |
test_history_backfill.py |
History backfill watermarks and the duplicate-append guard |
test_history_perf.py |
History search/indexing performance fixes (threading, FTS deletes, matrix buffer) |
test_history_search.py |
FTS search, filters, dedupe, authors table, message context, find_user |
test_history_sql.py |
The read-only sql_history tool and its sandbox |
test_mentions.py |
Which mentions in a reply become real pings |
test_message_media.py |
Stickers, GIFs, custom emojis, reactions and the add_reaction tool |
test_scheduler.py |
Cron job firing and rescheduling |
test_summary.py |
/resumo: period parsing, fetching, map-reduce, permissions |
test_tool_loop.py |
Tool errors reported to the model instead of aborting the turn |
test_trajectory_replay.py |
Trajectory capture, verbatim replay, text tool-call recovery |
tests/conftest.py puts the repo root on the import path, runs async def tests without a plugin, and provides a tmpdb fixture. tests/helpers.py has check(label, condition, detail), a labelled assert. New tests go in the file for the module they cover, or a new test_<module>.py.
See repository license. Documentation sources retain their original licenses.