TritonParse: A Compiler Tracer, Visualizer, and Reproducer for Triton Kernels
-
Updated
Sep 30, 2026 - Python
TritonParse: A Compiler Tracer, Visualizer, and Reproducer for Triton Kernels
An end-to-end agent project for GPU kernel implementation, analysis, profiling, and iterative optimization. It helps an agent turn PyTorch logic or an existing kernel into a high-performance GPU kernel through a structured, profile-driven workflow.
可验证的 CUDA 学习主线:SGEMM、通用 GPU 算子、性能优化与轻量推理组件
Inference Optimization by Economic Kernel Search
Wire the 1CatAI Split-D D256 FlashAttention kernel (fishlikeX/sm70-attn, MIT) into NInfer on Tesla V100 sm_70: +36-41% prefill, TTFT -3min, decode unchanged. Measured data + integration guide. Published by the user with AI assistance.
Tesla V100 (sm_70) LLM inference on one card: sm70 decode kernel port + KV context-cache tuning for long-context agents, measured on a real 53-request Qwen3.8-27B agent session. 单卡 Tesla V100 32GB 跑 NInfer + Qwen3.8-27B:sm70 解码内核移植 + 上下文缓存调参,附 53 个真实 agent 请求的满载实测与全部原始日志(中文为主,含英文版)。Posted by an AI on behalf of the machine's owner.
Production-grade HIP kernel optimization lab — matrix ops, reductions, shared memory patterns
Single-pass Triton kernel fusing KV-cache eviction scoring into FlashAttention decode — K/V read once, ~2.4x vs 3-pass baseline on T4
GQA/MQA-extended FlashAttention-2 in Triton, with a KV-cache decode kernel — built on a public tutorial baseline, hardened for correctness (4 bugs found/fixed), and benchmarked against PyTorch SDPA on NVIDIA L4.
Hand-written Triton PagedAttention decode kernel with reproducible correctness & performance benchmarks against PyTorch baselines
寻找乐趣, 寻找生活 — ML / GPU Kernel / 技术笔记
To associate your repository with the gpu-kernel topic, visit your repo's landing page and select "manage topics."