Pinned Loading
-
semantic-cache-server
semantic-cache-server PublicTiered semantic cache for LLMs — Redis exact-match + pluggable vector search (Chroma/Qdrant/Pinecone/pgvector) to avoid redundant LLM calls. FastAPI, production-hardened with request coalescing, ci…
Python
-
local-voice
local-voice PublicOffline, privacy-first voice dictation for Windows. Global hotkey → on-device Whisper transcription → clipboard, with zero cloud calls. Rust (Tauri) + Python (FastAPI) + React.
Python 1
-
enterprise-DOC-Rag
enterprise-DOC-Rag PublicProduction-grade, multi-tenant document RAG platform with hybrid search (dense + BM25), reranking, semantic caching, incremental ingestion, and streaming citations. FastAPI + React + Qdrant + Postg…
Python
-
prompt-studio
prompt-studio PublicGitHub for prompts — version, test, and compare LLM prompts across providers, with a self-calibrating LLM-as-judge.
Python
-
local-mcp-crm
local-mcp-crm PublicLocal-first CRM built on the Model Context Protocol (MCP) — customer/project management exposed as MCP tools, usable from a custom LlamaIndex agent or directly in Claude Code or any other applicati…
Python 1
-
ai-gateway-service
ai-gateway-service PublicProduction-style multi-provider LLM gateway (OpenAI, Claude, Gemini, Ollama) with automatic fallback, SSE streaming, per-key rate limiting, and cost tracking — built with FastAPI + async SQLAlchemy.
Python
If the problem persists, check the GitHub status page or contact support.

