Read this in: English | 简体中文
Comprehensive intelligent search solutions from deep research to code search
openJiuwen Search is a collection of search capabilities within the openJiuwen open-source project, providing comprehensive intelligent search solutions from deep research to code search. This repository brings together multiple search agents targeting different scenarios, aiming to deliver enterprise-grade Agentic AI search capabilities to enterprises and developers.
DeepSearch is a knowledge-augmented, high-performance, high-precision deep retrieval and research engine. It leverages structured knowledge and large language models with various tools to deliver enterprise-grade Agentic AI search and research capabilities.
Core capabilities:
- Deep Research: Multi-step, multi-source validated, rigorously reasoned structured research report generation
- Deep Search: Intelligent Q&A based on state-space search with multi-step reasoning and tool calling
- Knowledge-augmented hybrid retrieval: Fusion of local knowledge bases (keyword, vector, graph retrieval) with web-augmented engines
- Rich reports with visuals: Report generation with embedded figures and charts, fully traceable content
- Segment-level provenance: Outputs with validated citations, supporting segment-level traceability and confidence assessment
- Collaborative and interactive: Natural-language feedback during the planning phase
Use cases:
- Financial analysis reports: Connect local investment and finance knowledge bases to generate investment and financial research reports
- Academic and policy research: Gather policy information and implementation details, then generate research reports
- Complex query search: Solve professional decision-making scenarios requiring multi-step, multi-source validation
Documentation: DeepSearch project docs
CodeSearch is an intelligent retrieval engine for code repositories. Given a problem description (e.g. a GitHub Issue, bug report, or feature request), it outputs "which files and which lines to look at to resolve it," providing precise code context for bug localization, code Q&A, and automated repair pipelines.
Core capabilities:
- Agentic multi-round retrieval: The retrieval agent makes autonomous decisions (browsing repo structure, multi-strategy search, expanding context, filtering commits), with dual-model collaboration to control costs
- Code-oriented hybrid indexing: Syntax-aware chunking (currently Python), dual-sparse BM25, incremental indexing with optional dense vector retrieval
- Dual-engine equivalent implementation: Workflow graph engine (default) + pure-code loop engine, with tests locking byte-identical output
- Service-oriented engineering: SDK, CLI, HTTP service, and container image form factors
Use cases:
- Bug localization: Turn an Issue into specific functions and code lines that need modification
- Code Q&A context supply: Provide line-precise code evidence for questions like "where is this feature implemented"
- Large repository navigation: Exchange natural language for relevant code slices in unfamiliar large codebases
Quick guide: CodeSearch quick guide
Documentation: CodeSearch project docs
openjiuwen-search-base is the shared foundation layer for the openJiuwen search product family, providing reusable LLM integration, vector retrieval, workflow, and other base components for DeepSearch, CodeSearch, and other products. This package has no dependencies on any product package; its core depends only on pydantic, with heavier dependencies offered as optional groups using guarded imports that can be enabled on demand.
Core modules:
- llm: LLM client protocol, standardized models for messages and tool calls, openJiuwen model adapter (with SSL certificate handling),
LLMConfig - embedding: OpenAI-compatible embedding client with local SQLite caching, limited retries with exponential backoff, persistent connection reuse
- milvus: Safe query expression construction (unified escaping), collection naming conventions (product prefix + schema version), general-purpose storage client
- workflow: Workflow node templates (three-phase) and branch routing construction
- logging_utils: Log management and sensitive information masking
- runtime: Run registry, passed via run_id through workflows; live objects never enter copyable workflow state
Design principles:
- Dependency direction: base depends on no product package; products depend on base
- Optional heavy dependencies: Core import paths never touch openjiuwen / pymilvus / aiohttp; pure logic parts remain importable and testable without the optional groups installed
- Namespace isolation: Collection names follow
{prefix}{name}__{schema_version}format, allowing multiple products to safely share the same Milvus instance
Quick install:
pip install -e . # core
pip install -e '.[workflow,milvus,embed]' # enable extras on demandDocumentation: base project docs
For secondary development, customized deployment, or source-level debugging:
Recommended model configuration:
- Verified models: Qwen3.7-Max (recommended), Qwen3.8-Flash, Qwen3.7-Plus, Qwen3-Max, GLM-5.3, GLM-5.3-Flash, GLM-5.2, GLM-5.1, GLM-5, DeepSeek V3.2, Kimi-K2.5
- Tip: Use more capable models for report generation to balance output quality and call stability
- Note: When using reasoning models, report generation time increases significantly; if speed matters, prefer non-reasoning models
# Index a local repository
pip install -e ../base -e '.[dev,milvus,llm,server]'
codesearch index --repo /path/to/your/repo --collection my_repo
# Search with natural language (configure CODESEARCH_LLM_API_KEY etc. in .env)
codesearch search --collection my_repo --query "TypeError when calling foo() with empty list"
Detailed usage quick start: CodeSearch quick start
- Product overview - Product positioning and core features
- Installation guide - Detailed installation and configuration
- Developer guide - Developer documentation
- FAQ - Frequently asked questions
- Product overview - Product positioning and core features
- Installation guide - Detailed installation and configuration
- Quick start - Get started quickly
- Developer guide - Developer documentation
- FAQ - Frequently asked questions
- openjiuwen-search-base - Module and design documentation for the shared foundation library
DeepSearch is built mainly on openJiuwen agent-core and can connect to different LLMs and tools. The system consists of the following components:
- Manager: Agent creation, workflow orchestration, and configuration management on the agent-core framework
- Query planning: Intent-based routing, structural planning, task decomposition, query rewriting, and other query understanding capabilities
- Knowledge retrieval: Offline knowledge construction and online retrieval, supporting multiple retrieval modes
- Understanding and analysis: Understanding of retrieval results and other contextual information, including evaluation, refinement, expansion, and fusion
- Result generation: Answers, report generation, interactive editing, and result provenance
┌──────────── Indexing (offline) ────────────────┐
│ Code repo → Syntax-aware chunking → Milvus │
│ dual-sparse index · Incremental: file hash │
│ dedup, unchanged files shared across versions │
└─────────────────────────────────────────────────┘
┌──────────── Retrieval (online) ─────────────────┐
│ Problem description → Retrieval agent │
│ (decision model · filter model · segment │
│ memory · five tool types) → files + line │
│ ranges │
└─────────────────────────────────────────────────┘
Contributions of code, bug reports, and feature ideas are welcome!
- Submit Issues and Pull Requests
- See the contribution guide
- Join community discussions
This project is licensed under Apache 2.0. See the LICENSE file.
- Project homepage: https://github.com/openJiuwen-ai/deepsearch
- Official website: https://www.openjiuwen.com/en/
- Issue tracker: Issues
This product serves as a workflow orchestration tool only and does not include AI model capabilities. Users are responsible for compliance with applicable regulations such as the EU AI Act when connecting AI models for specific business scenarios.
If this project is helpful to you, please give us a Star! Your support is our motivation for continuous improvement!
