Rutgers computer science and data science student. I work on inference latency, memory bandwidth, and throughput under real hardware constraints, and I contribute those fixes upstream.
portfolio · merged open-source PRs · LinkedIn
NVIDIA/cuda-python#2087FIPS-safe hashes for program cache keys |
NVIDIA/cuda-quantum#4688nvqpp: discriminate measured-register bool iteration |
huggingface/accelerate#4054Aggregate profiler memory example |
Dao-AILab/flash-attention#2622weights_only=True across all torch.load sites |
ai-dynamo/dynamo#10281HTTP 415 for unsupported image formats |
linkedin/Liger-Kernel#1157Guard save_for_backward on grad_bias in fused linear CE |
KVCacheForge-XKV-cache bottleneck lab. Measures TTFT, latency, throughput, HBM stalls, and GPU busy against baseline deltas. |
RoboFleetOpsAWS-native robotics fleet control plane on Lambda, DynamoDB, SQS, IoT Core, and API Gateway, deployed via CDK CI. |
PosCacheBenchLong-context benchmark for positional-attention failure modes under fixed KV-cache budgets. |
LunarLinkBenchMonte Carlo model of lunar comms links, separating relay passes from direct-to-Earth feasibility. |
satellite telemetry anomaly detection100K telemetry readings · 5 NASA/ESA fault modes · recurrence-plot CV · 0.91 F1 on Kepler-class wheel oscillation PDF · repo |
bell labs ml impact analysis71-paper corpus · semantic clustering · co-authorship networks · Gradient Boosting AUC 0.674 · SHAP attribution PDF · repo |
Undergraduate at Rutgers University, B.S. Computer Science and Data Science, class of 2028. Based in the New York and New Jersey area, open to relocation including San Francisco and the Bay Area, and to remote.
systems project evidence · KV-cache writing · aryanputta.com



