Inference optimization
Cutting latency and cost out of model serving — batching, KV-cache reuse, quantization, and semantic caching to shave every millisecond off time-to-first-token.
inference engine — demo
AI Systems Engineer · LLM inference & serving
response
> whoami
AI Systems Engineer. I build low-latency systems that serve large language models — from inference optimization and serving infrastructure to ML systems under real production load.
focus
Cutting latency and cost out of model serving — batching, KV-cache reuse, quantization, and semantic caching to shave every millisecond off time-to-first-token.
Packaging models as reliable, reproducible low-latency endpoints with vLLM, Truss, and Docker — shipped through CI/CD and built to stay up under load.
Understanding the stack end to end: autograd, transformer internals, and language models built from scratch — so optimization is grounded, not guessed.
Designing AI systems with software-engineering rigor and MLOps discipline — efficient, observable, and robust under real production load.
signal
I'm an AI Systems Engineer at Mursion, focused on the part of AI that decides whether it's actually usable in production: how fast it responds, under real load. My work sits where async systems, concurrency, and ML serving meet — and I'm deepening further into LLM inference-serving: throughput, KV-cache efficiency, and tail latency at scale.
work
Production-grade CNN inference service (P90 250ms) on FastAPI & NVIDIA Triton — GPU (TensorRT) and CPU (OpenVINO) backends, plus a full Prometheus/Grafana stack detecting data drift and model degradation in real time.
Contextual-RAG research assistant on AWS. Compresses embeddings with binary quantization and keeps the vector store in memory to minimize retrieval latency at scale.
Reproducible LLM serving with vLLM, packaged in Docker images for consistent, portable high-throughput inference.
Semantic cache, vector store, and RAG built on FAISS — cutting repeat-query latency and cost by serving from cache when meaning matches.
Model packaging and deployment with Truss — turning a trained model into a served endpoint with minimal glue.
A language model built from scratch in PyTorch, iteratively evolved from foundations toward current industry-standard architecture.
Reimplementing Karpathy's micrograd — a tiny autograd engine — to understand backprop from first principles.
research
A comparative study of how transformer fine-tuning, large language models, and recurrent baselines trade off accuracy against the cost of scoring at scale — with the fine-tuned approach exceeding prior benchmarks by 5.7%.
writing
What actually breaks when training falls off the single-GPU happy path — memory blowups, collective-communication failures, and the resilience patterns that keep large runs alive.
Read on DeepDivesDeconstructing the transformer from self-attention up through the KV cache — the mechanics that decode-time inference optimization is actually built on.
Read on DeepDives