AI-powered forensic face sketch generation and criminal identification system.
WitSketch enables law enforcement to generate forensic facial composites from witness descriptions and match them against a criminal database — in real time.
| Feature | Description |
|---|---|
| 🎨 Text-to-Sketch Generation | Generate forensic pencil sketches from natural-language descriptions using Stable Diffusion v1-5 |
| 🧩 Visual Builder | Drag-and-drop composite face builder using pre-drawn facial element assets |
| 🔍 Image-Based Matching | Upload a sketch or photo to match against the criminal database using FaceNet embeddings |
| 📝 Description-Based Matching | Match suspects via attribute vectors derived from a witness description |
| 📹 CCTV Analysis | Scan surveillance video footage to locate a specific suspect or identify all criminals in a crowd |
| 🗃️ Admin Panel | Add new criminal records with photos (auto-converted to sketches for embedding) |
| 🌍 Multi-language Support | Witness descriptions are auto-translated to English via Google Translate |
┌─────────────────────────────────────────────────────┐
│ FastAPI Backend (app.py) │
│ │
│ /generate Stable Diffusion (SD v1-5) │
│ └─ GAN fallback (DCGAN) │
│ /match FaceNet + Cosine Similarity │
│ /attribute_match 11D Attribute Vector Matching │
│ /cctv_upload MTCNN Face Detection + Matching │
│ /generate_from_builder Img2Img Refinement │
│ /admin/add_record Add new criminal to DB │
└─────────────────────────────────────────────────────┘
│
▼
┌─────────────────────┐ ┌────────────────────────┐
│ criminal_records │ │ dataset/ (CUFS photos) │
│ .json (embeddings) │ │ Face Sketch Elements/ │
└─────────────────────┘ └────────────────────────┘
Key Components:
diffusion_generator.py— Wrapsrunwayml/stable-diffusion-v1-5for text-to-sketch and img2img generationmodels.py— CustomAttributeSketchGenerator(DCGAN) as a fast fallback generatorcctv_matcher.py— MTCNN-based face detection pipeline for video scanning (single-suspect tracking + crowd identification)utils/face_encoder.py— FaceNet encoder producing 512D embeddings for similarity searchcreate_mock_db.py— Builds the criminal database from CUFS dataset photos
| Layer | Technology |
|---|---|
| Backend | Python, FastAPI, PyTorch |
| AI Models | Stable Diffusion v1-5, DCGAN, FaceNet (facenet-pytorch) |
| Face Detection | MTCNN (facenet-pytorch) |
| Image Processing | OpenCV, Pillow |
| Translation | deep-translator (Google Translate) |
| Frontend | HTML, CSS, Vanilla JS |
- Python 3.10+
- macOS with Apple Silicon (MPS) or CUDA GPU recommended
# Clone the repository
git clone <your-repo-url>
cd WitSketch
# Create and activate a virtual environment
python -m venv venv
source venv/bin/activate
# Install dependencies
pip install -r requirements.txtNote: First startup downloads ~4 GB of Stable Diffusion model weights to
~/.cache/huggingface.
python create_mock_db.pyThis generates criminal_records.json with photo embeddings and attribute vectors.
uvicorn app:app --reload --host 0.0.0.0 --port 8000Visit http://localhost:8000 in your browser.
| Role | Username | Password |
|---|---|---|
| Admin | admin |
admin123 |
| User | user |
user123 |
- Go to Generate → enter a witness description (e.g. "young male, short black hair, beard, oval face")
- Choose Diffusion (accurate, ~30s) or GAN (fast fallback)
- Optionally request multiple views (front, left/right profile)
- Image Match (
/match): Upload a sketch or photo - Description Match (
/attribute_match): Enter a text description - Filter results by location and risk level
- Single-suspect tracking: Upload a video + a target suspect image to find all timestamps where the suspect appears
- Crowd identification: Upload a video to identify all criminals in the footage against the custom database
Navigate to the Admin panel to add a new criminal record with a photo. The system automatically:
- Crops the face using MTCNN
- Converts the photo to a pencil sketch (Color Dodge pipeline)
- Computes a 512D FaceNet embedding
- Extracts 11D attribute vectors from the description
- Saves the record to
criminal_records.json
| Method | Endpoint | Description |
|---|---|---|
POST |
/generate |
Generate a forensic sketch from description |
POST |
/generate_from_builder |
Refine a composite builder image via img2img |
POST |
/match |
Match uploaded image against the criminal DB |
POST |
/attribute_match |
Match by witness text description |
POST |
/cctv_upload |
Scan video for a specific suspect |
POST |
/cctv_crowd |
Identify all criminals in a video |
POST |
/admin/add_record |
Add a new criminal record |
GET |
/admin/stats |
View system usage statistics |
GET |
/api/elements |
List available face sketch elements |
WitSketch/
├── app.py # FastAPI server + all API endpoints
├── models.py # DCGAN generator & discriminator
├── diffusion_generator.py # Stable Diffusion wrapper
├── cctv_matcher.py # Video face detection & matching
├── attribute_sketch_dataset.py # Attribute vector encoding
├── create_mock_db.py # Build criminal database
├── utils/
│ └── face_encoder.py # FaceNet 512D embedding encoder
├── static/ # Frontend HTML/CSS/JS
│ ├── login.html
│ ├── dashboard.html
│ ├── generate.html
│ ├── builder.html
│ ├── match.html
│ ├── cctv.html
│ └── admin.html
├── dataset/ # CUFS face photos
├── Face Sketch Elements/ # Facial composite assets
├── checkpoints_attribute/ # GAN model checkpoints
└── criminal_records.json # Criminal database with embeddings
- Uploaded image is converted to a pencil sketch (OpenCV Color Dodge) to normalise to the same domain as the database
- FaceNet extracts a 512D embedding
- Cosine similarity is computed against all database embeddings
- Final score = 95% embedding similarity + 5% risk level (tie-breaker)
- Witness description is translated to English and parsed for attributes (gender, hair, beard, glasses, face shape, age)
- An 11D attribute vector is computed
- Cosine similarity against all stored attribute vectors
- Final score = 85% attribute similarity + 15% risk level
This project is intended for academic and research purposes only.