This project implements an ADK-based agent that builds and maintains a persistent knowledge base (wiki) in Google Cloud Storage (GCS), following the "Active Knowledge Agent Wiki" pattern. It avoids traditional RAG (Retrieval-Augmented Generation) and vector search, relying instead on the LLM to actively synthesize and organize knowledge into interlinked markdown files, guided by an index. It also features a rich Web UI with an interactive graph view.
Unlike traditional RAG systems that retrieve raw document chunks at query time and synthesize answers from scratch every time, this agent:
- Incrementally builds a structured, interlinked collection of markdown files (the Wiki).
- Maintains consistency and cross-references as new sources are added.
- Uses an index file (
index.md) to navigate the wiki, avoiding the need for vector databases. - Captures Explicit Relationships: Defines typed connections (e.g., "regulated_by") in page frontmatter.
- Organizes with Tags: Assigns tags to pages for structured discovery.
- Dynamic Multi-Layer Hierarchy: Organizes files into logical directories and subdirectories based on domain, growing dynamically as needed.
The system consists of five main layers:
- Raw Sources: Files or URLs provided by the user (immutable).
- GCS Raw Data: Unmodified raw files stored in GCS under the
raw_data/folder (acting as the source of truth). - The Wiki: A directory of LLM-generated markdown files stored in GCS (configured via
WIKI_BUCKET_NAMEenvironment variable). - The Schema:
schema.md(also in GCS) defining rules and conventions for the agent. - The Web UI: A Next.js application providing:
- Tree View Sidebar: Dynamically generated navigation supporting arbitrary depth.
- Interactive Graph View: Visualizes links, explicit relationships, and tag clusters.
- Perspective Rendering: Filters graph to show only selected node and its neighbors.
The system is powered by a hierarchical multi-agent orchestration system built on top of the Google Agent Development Kit (ADK). Rather than a single agent attempting to execute all reasoning, verification, and compilation steps sequentially, tasks are delegated to specialized, autonomous sub-agents collaborating through a central orchestrator.
- Orchestrator Agent (agent.py): The root agent that receives input, coordinates sub-agents, and implements the dynamic response verification loop.
- Wiki Researcher Agent (wiki_researcher_agent.py): A specialized agent that executes the wiki graph retrieval, following links and source summaries to retrieve and synthesize raw data under
raw_data/. - Critic Agent (critic_agent.py): A verification agent that evaluates draft answers against wiki facts and original raw documents to output
APPROVEDorREVISEwith feedback. - Synthesizer Agent (synthesizer_agent.py): Processes raw text, creates/modifies markdown wiki pages, and defines frontmatter relationships.
- Reviewer Agent (reviewer_agent.py): Audits changes for schema compliance and factual contradictions.
- Librarian Agent (librarian_agent.py): Re-indexes wiki content, logs historical actions (
log.md), and tracks stub gaps (gaps.md). - Schema Manager Agent (schema_manager_agent.py): Evolves conventions in
schema.mdinteractively based on directory proposals.
graph TD
User([User]) -->|Browse & Audit| UI["Wiki Web UI"]
User -->|Ingest / Query / Lint / Schema Request| Server["FastAPI Backend Server"]
UI -->|REST API Requests| Server
Server -->|Read Wiki Content| GCS[(GCS Wiki Bucket)]
subgraph backend ["FastAPI Backend Server (ADK App)"]
Server --> Orch["Orchestrator (ADK Workflow)"]
Orch -->|Ingest| Extractor["Extractor Node/Tool"]
Extractor --> Synth["Synthesizer Agent"]
Synth --> Rev["Reviewer Agent"]
Rev --> Lib["Librarian Agent"]
Orch -->|Query| Res["Wiki Researcher Agent"]
Res --> run_critic{"Critic Node (Verification Loop)"}
run_critic -->|Revise| Res
run_critic -->|Approved| END([Done])
Orch -->|Schema| SchemaMgr["Schema Manager Agent"]
Orch -->|Lint| Linter["Linter Node"]
Linter --> Rev
end
Extractor -->|Upload Raw| GCS
Synth -->|Read/Write Pages| GCS
Rev -->|Read/Write Contradictions & Audits| GCS
Lib -->|Read/Write index.md, log.md & gaps.md| GCS
SchemaMgr -->|Read/Write schema.md & schema_proposals.md| GCS
Res -->|Read index.md & Pages| GCS
Linter -->|Read Pages for audit| GCS
subgraph gcs_bucket ["GCS Bucket Layout"]
schema[schema.md]
index[index.md]
log[log.md]
gaps[gaps.md]
proposals[schema_proposals.md]
raw_data[raw_data/]
hierarchy[Dynamic Hierarchical Folders...]
end
GCS --- gcs_bucket
The sequence below illustrates how the Orchestrator coordinates the ingestion of new resources, ensuring uncompromised data extraction, active cross-referencing, validation auditing, and automated directory registration:
sequenceDiagram
autonumber
actor User
participant Orchestrator as Orchestrator Agent
participant Extractor as Extractor Node / Tool
participant Synthesizer as Synthesizer Agent
participant Reviewer as Reviewer Agent
participant Librarian as Librarian Agent
participant GCS as GCS Storage Bucket
User->>Orchestrator: Ingest Request (Source URL/File)
Orchestrator->>Extractor: Extract content & upload
Extractor->>GCS: Upload original file to raw_data/
Extractor->>Orchestrator: Extracted source text
Orchestrator->>Synthesizer: Write wiki pages
Synthesizer->>GCS: Read existing pages & write updates
Synthesizer->>Orchestrator: Manifest of modified pages
Orchestrator->>Reviewer: Verify factual integrity
Reviewer->>GCS: Cross-check claims & verify directory schema
Reviewer->>Orchestrator: Factual audit & contradiction review report
Orchestrator->>Librarian: Re-index & Log
Librarian->>GCS: Update index.md, log.md, gaps.md & proposals
Librarian->>Orchestrator: Confirmation
Orchestrator->>User: Final ingestion summary & health delta
When new domains emerge from digested source documents, the system records proposed directory schemas in schema_proposals.md. The Orchestrator manages this schema evolution pipeline interactively via the Schema Manager Agent:
sequenceDiagram
autonumber
actor User
participant Orchestrator as Orchestrator Agent
participant SchemaMgr as Schema Manager Agent
participant GCS as GCS Storage Bucket
User->>Orchestrator: Schema Manage/Review Request
Orchestrator->>SchemaMgr: Check schema proposals
SchemaMgr->>GCS: Read schema_proposals.md
alt No pending proposals
SchemaMgr-->>User: No pending schema proposals found
else Pending proposals exist
SchemaMgr->>User: Present proposals for interactive approval
User->>SchemaMgr: Approve/Reject proposal selections
SchemaMgr->>GCS: Merge approved definitions into schema.md
SchemaMgr->>GCS: Clear approved entries from schema_proposals.md
SchemaMgr-->>User: Merge and update confirmation
end
When a user asks a question to retrieve or summarize information from the wiki (e.g., "Summarize the compliance frameworks for IAP"), the system bypasses traditional vector retrieval databases. Instead, the Orchestrator uses the central index file to locate highly relevant, hand-grounded documents and files:
sequenceDiagram
autonumber
actor User
participant Orchestrator as Orchestrator Agent
participant Researcher as Wiki Researcher Agent
participant GCS as GCS Storage Bucket
User->>Orchestrator: Q&A/Summary Request (e.g., "Summarize frameworks for IAP")
Orchestrator->>Researcher: Delegate query to researcher
loop Verification & Critic Loop (Capped at 2 revisions)
Researcher->>GCS: read_wiki_file("index.md") & "log.md"
GCS-->>Researcher: Index structure & chronological log
Note over Researcher: Navigate & read wiki pages,<br/>source summaries, & raw files
Researcher->>GCS: read_wiki_file("raw_data/original_file.pdf")
GCS-->>Researcher: Original GCS raw document
Note over Researcher: Generate draft response
Researcher-->>Orchestrator: Draft response
Orchestrator->>Critic: Run critic review
Critic->>GCS: Verify claims & citations
Critic-->>Orchestrator: STATUS (APPROVED or REVISE)
alt APPROVED
Note over Orchestrator: Finalize approved response
else REVISE (Increment Loop Count)
Note over Orchestrator: Append draft and critic's feedback<br/>to session history (await asyncio.sleep)
end
end
Orchestrator-->>User: Final response rendered as markdown chat bubble (message_as_output)
To fully appreciate the benefits of this active, compounding knowledge base, it is helpful to compare it directly with traditional Retrieval-Augmented Generation (RAG).
| Feature | Traditional RAG | Active Knowledge Agent Wiki Pattern |
|---|---|---|
| State & Memory | Stateless. Retrieves chunks on-the-fly for each query. Forgets what it synthesized last time. | Stateful. Actively integrates new information into an evolving, structured knowledge base. |
| Precision & Links | Fuzzy Similarity. Relies on vector distance, which can retrieve out-of-context or irrelevant text. | High-Precision Graph. Uses hard, semantic relationships (regulated_by, contradicts) and tags defined by the LLM in page frontmatter. |
| Auditability | Black Box. Vector store contains binary embeddings. Extremely difficult for humans to audit or manually correct. | Transparent. Made of clean, human-readable Markdown files in GCS. Humans can directly read and edit the agent's memory. |
| Infrastructure | High Complexity. Requires running a vector database, embedding APIs, chunking algorithms, and tuning parameters. | Zero Vector Cost. Relies entirely on standard cloud storage (GCS) and file system structures. No vector database needed. |
| Temporal Validity | Time Blind. Vector search cannot separate overlapping historical document versions, retrieving conflicting text from multiple years (e.g., 2024 vs. 2025 booklet chunks). | Time Aware. Pages are organized in versioned directories and tagged with validity dates in frontmatter, enabling the Orchestrator to query precise historical policy contexts. |
The Active Knowledge Agent Wiki architecture shines in complex, long-form knowledge environments where information is dynamic, highly interlinked, and requires human-in-the-loop verification.
- The Challenge: An insurance claims handler faces an influx of 100+ documents per claimβincluding police reports, medical bills, mechanic estimates, photos, and email exchanges. The claim evolves over weeks or months, and the handler needs to understand the chronological timeline, identify inconsistencies (e.g., medical treatments mismatching the police report), and build a final audit trail.
- Why Traditional RAG Fails: RAG retrieves disconnected fragments of text (e.g., a page from a medical report, a sentence from a policy). It cannot synthesize a cohesive timeline or recognize that a fact retrieved today directly contradicts a fact retrieved two weeks ago because it does not keep state.
- The Active Knowledge Agent Wiki Solution: The agent ingests incoming claim documents and actively maintains a compounding claim wiki.
- It builds a structured timeline (e.g.,
/claims/CLAIM-123/timeline.md), updating it chronologically. - It maps explicit relationships, such as tying
/claims/CLAIM-123/injury-report.mdto the/claims/CLAIM-123/medical-provider.mdviatreated_by. - The claims handler can review the resulting claims graph in the Web UI, instantly audit the LLM's synthesis, and correct any errors in the markdown files directly, ensuring perfect factual alignment before final payout approval.
- It builds a structured timeline (e.g.,
- The Challenge: Compliance officers in financial services or healthcare must track hundreds of fast-changing regulatory updates, internal policies, and audit reports. They need to map how a new state law impacts existing corporate rules.
- The Active Knowledge Agent Wiki Solution: The agent ingests new regulatory circulars and actively updates a corporate policy wiki. It creates tags for compliance areas and links policies directly to regulations (e.g.,
policy.md-[implemented_for]->regulation.md). This dynamic compliance graph lets officers immediately see the blast radius of any rule change.
- The Challenge: Developers or research teams trying to map out complex system architectures, codebase structures, or open-source protocols.
- The Active Knowledge Agent Wiki Solution: The agent maps out repositories, creates structural directories, extracts and connects concepts (e.g., mapping how MCP servers interact with Agent Platforms), and visualizes these relationships dynamically, creating a self-documenting codebase.
To keep the agents highly expandable and prevent token context bloat, the system uses the adk-progressive-skills library for progressive skill discovery:
.agents/skills/Directory: Local workspace directory containing custom ADK skills, each in its own sub-folder containing aSKILL.md(e.g., wiki-traversal, critic-evaluation, report-writing).- Progressive Discovery: Skills are scanned across four precedence paths (lowest to highest: global jetski, global user config, home directory, and local workspace
.agents/skills). - Dynamic Toolset Injection: The library patches the ADK
Agentconstructor to automatically instantiate and mount aSkillToolsetcontaining all discovered skills as tools, ensuring agents only retrieve skill contexts dynamically as needed during workflow execution.
To ensure answers are mathematically and factually precise, all queries run through a verification loop:
- Draft Generation: The researcher agent generates a draft response citing its findings.
- Factual Validation: The Critic Agent evaluates the draft response against index, pages, and raw files. It checks grounding (no hallucinations), citation formatting, and completeness.
- Actionable Feedback: If validation fails, the Critic outputs
STATUS: REVISEwith feedback. The orchestrator feeds the previous draft and feedback back to the researcher for correction. - Iteration Capping: To prevent infinite loops or excessive API usage, validation loops are capped at a maximum of 2 iterations. On the 3rd iteration, the draft is auto-approved and delivered with a notice.
- Cooperative Multitasking History: The loop uses
await asyncio.sleep(0)during revision to yield execution, ensuring the user revised draft is successfully saved to the persistent session database history.
Formatting rules are governance-defined under report-writing/SKILL.md:
- Structure: Every answer must be structured with an Executive Summary, Detailed Findings, Analysis & Context, and a dedicated Sources and Grounding Citations section.
- Grounding Citation Authority: Citations must map Wiki Pages and Source Summaries directly to the original raw files stored under
raw_data/in GCS (the ultimate grounding authority).
- To prevent the UI from displaying raw, unformatted markdown text boxes at the end of the query path, the
run_criticnode emits the final approved response as a standard modelcontentmessage withnode_info=NodeInfo(message_as_output=True). - This informs the ADK engine that the message content is the final output of the node, rendering it as a standard, beautifully formatted chat bubble.
agentwiki-adk/
βββ .agents/ # Local dynamic skill sheets folder (progressive discovery)
β βββ skills/
β βββ wiki-traversal/
β β βββ SKILL.md # Wiki navigation and raw GCS citation rules
β βββ critic-evaluation/
β β βββ SKILL.md # Response validation criteria and feedback rules
β βββ report-writing/
β βββ SKILL.md # Professional report template and grounding rules
βββ app/
β βββ __init__.py
β βββ agent.py # Defines the ADK Orchestrator Agent and workflow
β βββ agent_runtime_app.py # Entry point for Agent Runtime
β βββ agents/ # Specialized sub-agents
β β βββ __init__.py
β β βββ wiki_researcher_agent.py # Retrieval and Q&A researcher agent
β β βββ critic_agent.py # Grounding & citation validation agent
β β βββ synthesizer_agent.py # Wiki composition & GCS editing agent
β β βββ reviewer_agent.py # Schema & contradiction auditing agent
β β βββ librarian_agent.py # Bookkeeping, indexing, & gaps logging agent
β β βββ schema_manager_agent.py # Automated schema evolution manager agent
β βββ tools/
β βββ __init__.py
β βββ gcs_io.py # Tools for reading/writing to GCS
β βββ extractor.py # Tools for content extraction
β βββ health.py # Quantitative wiki health calculation tool
βββ frontend/ # Next.js Web UI
β βββ app/ # App router pages and API routes
β βββ components/ # React components (Graph, Sidebar, etc.)
β βββ ...
βββ pyproject.toml # Dependencies and project metadata
βββ schema.md # Initial schema (uploaded to GCS)
βββ index.md # Initial index (uploaded to GCS)
βββ log.md # Initial log (uploaded to GCS)
βββ README.md # This file
uvinstalled.agents-cliinstalled (uv tool install google-agents-cli).- Google Cloud SDK installed and configured.
Before running the agent or the Web UI, you must configure the GCS bucket where the wiki will be stored.
For local development, create a .env file in the project root and in the frontend/ directory to define the GCS bucket name:
WIKI_BUCKET_NAME=your-unique-gcs-bucket-name(Note: .env files are already configured in .gitignore and will not be committed.)
agents-cli installTo test the agent locally using the playground:
agents-cli playgroundImportant
Detailed step-by-step instructions for deploying this project to Google Cloud Platform (GCP) with secure direct Identity-Aware Proxy (IAP) integration are maintained in instructions.md.
Please refer to instructions.md for:
- Required GCP APIs and IAM Role configurations
- Cloud Run service setup with
--iapflag - GCS Bucket permission binding
- Multi-stage Docker builds and Artifact Registry push commands
