-
Nokia
- Munich, Germany
- https://anubhabbanerjee.github.io/
- https://linkedin.com/anubhab-banerjee/
Popular repositories Loading
-
Annotated-LLM-Runtime
Annotated-LLM-Runtime PublicFrom-scratch, heavily-annotated CUDA inference runtime for Qwen2.5-Coder-7B on H100 (sm_90). Custom INT4 packer, fused GEMV, paged KV, split-KV attention, CUDA graph decode — every hot path comment…
-
WarpGroup-backend
WarpGroup-backend PublicA high-performance C++ backend for extreme-context LLM inference. It replaces item-count batching with dynamic, VRAM-aware First-Fit Decreasing (FFD) bin packing. By using PyBind11 for async queuei…
-
VRAM-Conductor
VRAM-Conductor PublicC++17 + CUDA orchestrator for running multiple LLM agents on one constrained GPU. `lmxd` daemon does NVML-seeded admission control to stop llama.cpp OOM crashes; `LayerStreamer` + `PinnedHostPool` …
C++ 3
-
inter-llm-tokf
inter-llm-tokf PublicInter-LLM knowledge handover framework leveraging Open Knowledge Format (OKF) for Qwen family. Bypasses text paring and tokenization by passing binary token arrays directly into model embedding lay…
-
sionna-munich-ai-ran
sionna-munich-ai-ran PublicGPU-accelerated Network Digital Twin for AI-RAN experiments in Munich. Powered by NVIDIA Sionna RT and PyTorch for ray-traced RF simulation. Includes real-world Munich simulation, Mitsuba 3 backend…
If the problem persists, check the GitHub status page or contact support.
