A Python tool that categorizes AI assistant conversations (Claude, ChatGPT) by topic using multilingual sentence embeddings and keyword-based classification. Analyze your conversation history to discover usage patterns across 12 topic categories with subcategory granularity.
🔍 Try it in your browser — a light, no-install version of the categorizer (keyword stage only, runs fully client-side, nothing uploaded) is live at mikekrohn.ai/tools/categorizer. Clone this repo for the full pipeline (embeddings, clustering, charts).
- Interactive web dashboard — Drag-and-drop JSON or CSV files in your browser to instantly visualize conversation patterns with charts and a searchable data table
- Multi-format ingestion — Parses conversation exports from both Claude and ChatGPT, handling nested JSON structures automatically
- Multilingual NLP — Generates language-agnostic embeddings with
paraphrase-multilingual-mpnet-base-v2, supporting English, Danish, and more - 12 topic categories — Classifies conversations into categories like Code Development, Learning/Education, and Creative & Ideation, each with fine-grained subcategories
- GPU acceleration — Automatic CUDA / Apple Silicon (MPS) detection with dynamic batch sizing and mixed-precision (FP16) inference
- Performance-first design — Parallel file I/O, persistent caching, compiled regex patterns, and chunked memory management
- Built-in benchmarking — Measure pipeline throughput and resource usage across all stages
JSON Exports ──► Data Processing ──► Embedding Generation ──► Categorization ──► Results
(Claude/ChatGPT) │ │ │ │
├─ Text extraction ├─ SentenceTransformer ├─ Regex scoring ├─ CSV output
├─ Language detect ├─ GPU/CPU auto-select ├─ 12 categories ├─ JSON metrics
└─ Parallel I/O └─ Embedding cache └─ Subcategories └─ Confidence
Stage 1 — Data Processing reads raw JSON conversation exports, extracts text from Claude (chat_messages) and ChatGPT (mapping) formats, detects language, and writes a cleaned CSV. Uses parallel I/O and language detection caching for speed.
Stage 2 — Embedding Generation encodes conversation text into dense vector embeddings using a multilingual SentenceTransformer model. Supports GPU acceleration, dynamic batching, and persistent caching to avoid redundant computation.
Stage 3 — Categorization scores each conversation against 12 topic categories using compiled regex keyword patterns. Assigns a primary category, subcategory, confidence score, and voice-conversation flag. Outputs a categorized CSV and JSON metrics summary.
- Runtime architecture
src/data_processing.py→ stage 1 ingestion/cleaningsrc/embedding.py→ stage 2 embedding generationsrc/clustering.py→ stage 3 categorization + metricssrc/app.py+src/static/index.html→ web dashboard/API
- CI and delivery workflows
.github/workflows/ci.yml→ lint + unit tests + coverage artifact.github/workflows/ui-test.yml→ Playwright UI tests (PR path-filtered + nightly schedule).github/workflows/pages.yml→ deploy static docs to GitHub Pages.github/workflows/codeql.yml→ CodeQL security analysis.github/workflows/dependency-scan.yml→ dependency vulnerability scan (pip-audit)
| Category | Example Subcategories |
|---|---|
| Learning/Education | Concept Understanding, How-to Learning, Academic Topics |
| Code Development | Bug Fixing, Feature Development, Code Review |
| Writing Assistance | Content Creation, Editing, Format/Style |
| Analysis/Research | Data Analysis, Research Review, Comparative Analysis |
| Creative & Ideation | Idea Generation, Visual Design, Innovation |
| Professional/Business | Strategy, Client/Customer, Business Analysis |
| Technical Support | Troubleshooting, Setup/Installation, Integration Issues |
| Personal Projects | Project Planning, Implementation, Review/Feedback |
| SoMe/Marketing | Content Creation, Event Announcements, Campaign Planning |
| DALL-E/Image | Image Generation, Image Editing, Style Transfer |
| Cooking/Food | Recipe Help, Ingredient Questions, Meal Planning |
| Information/Curiosity | General Knowledge, Cause/Effect, Research Requests |
- Python 3.9+
- pip
git clone https://github.com/Walliiee/GenAICategorizer.git
cd GenAICategorizer
pip install -e ".[web]"For GPU acceleration (CUDA):
pip install torch --index-url https://download.pytorch.org/whl/cu118
pip install -e ".[web]"The fastest way to explore your conversation data:
python -m src.appOpen http://localhost:8000 in your browser, then drag-and-drop your JSON or CSV files onto the page. The dashboard shows:
- Category distribution — horizontal bar chart of all 12 topic categories
- Language breakdown — doughnut chart of detected languages
- Complexity analysis — simple / medium / complex split
- Voice vs text — voice conversation detection
- Confidence scores — histogram of categorization confidence
- Searchable data table — filter by category, complexity, or free-text search; sortable columns; click to expand text previews
Build and run the web dashboard/API in a container — no local Python setup required:
docker build -t genai-categorizer .
docker run --rm -p 8000:8000 genai-categorizerThen open http://localhost:8000, or POST files to http://localhost:8000/api/analyze.
Notes:
- The image is CPU-only and pulls the CPU build of PyTorch. The
paraphrase-multilingual-mpnet-base-v2model (~569 MB) downloads on first request, so the first analysis is slower (subsequent ones are fast). - To persist the model between runs, mount a cache volume:
-v hf-cache:/app/.cache/huggingface.
For batch processing or scripting, run the three-stage pipeline directly:
-
Place your conversation JSON exports in
data/raw/. -
Run each stage:
# Stage 1 — Process raw JSON files into cleaned CSV
genai-categorizer-process
# Stage 2 — Generate sentence embeddings
genai-categorizer-embed
# Stage 3 — Categorize conversations
genai-categorizer-categorize- Find your results in
data/processed/:
| File | Description |
|---|---|
cleaned_conversations.csv |
Extracted text with language labels |
embeddings.npy |
Dense vector embeddings |
categorized_conversations.csv |
Final categorized output |
conversation_metrics.json |
Distribution and performance metrics |
genai-categorizer-benchmarkMeasures execution time, memory delta, and throughput for each pipeline stage. Results are saved as timestamped JSON files in outputs/benchmarks/.
GenAICategorizer/
├── src/
│ ├── __init__.py
│ ├── app.py # FastAPI web dashboard (drag-and-drop UI)
│ ├── data_processing.py # JSON parsing, text extraction, language detection
│ ├── embedding.py # Sentence embedding generation with caching
│ ├── clustering.py # Keyword-based topic categorization
│ ├── benchmark.py # Pipeline performance measurement
│ └── static/
│ └── index.html # Dashboard SPA (Chart.js, dark theme)
├── tests/
│ ├── test_data_processing.py
│ ├── test_clustering.py
│ └── test_embedding.py
├── data/
│ ├── raw/ # Input: conversation JSON exports (git-ignored)
│ └── processed/ # Output: CSVs, embeddings, metrics (git-ignored)
├── .github/workflows/ci.yml # GitHub Actions CI
├── pyproject.toml # Project metadata and dependencies
├── LICENSE
└── README.md
pip install -e ".[dev]"pytestTo run UI tests locally:
pip install -e ".[ui]"
playwright install chromium
pytest tests/test_ui.py -v -m uiruff check src/ tests/This project is licensed under the MIT License — see LICENSE for details.