inference engine — demo

Satyam
Chatrola

AI Systems Engineer · LLM inference & serving

scroll to boot

response

> whoami

AI Systems Engineer. I build low-latency systems that serve large language models — from inference optimization and serving infrastructure to ML systems under real production load.

focus

What I do

inference

Inference optimization

Cutting latency and cost out of model serving — batching, KV-cache reuse, quantization, and semantic caching to shave every millisecond off time-to-first-token.

serving

Serving infrastructure

Packaging models as reliable, reproducible low-latency endpoints with vLLM, Truss, and Docker — shipped through CI/CD and built to stay up under load.

internals

Systems from first principles

Understanding the stack end to end: autograd, transformer internals, and language models built from scratch — so optimization is grounded, not guessed.

scale

Scale & reliability

Designing AI systems with software-engineering rigor and MLOps discipline — efficient, observable, and robust under real production load.

signal

I turn models into
fast, dependable endpoints.

−71%p99 latency reduction
5×telemetry serialization speedup
5M+SKUs searched
+34%click-through lift

I'm an AI Systems Engineer at Mursion, focused on the part of AI that decides whether it's actually usable in production: how fast it responds, under real load. My work sits where async systems, concurrency, and ML serving meet — and I'm deepening further into LLM inference-serving: throughput, KV-cache efficiency, and tail latency at scale.

  • Cut p99 latency by 71% in production by moving blocking calls off the asyncio event loop, preserving session ordering.
  • Sustained throughput under heavy concurrent load with an async logging queue — watermark-based load-shedding and eventual-consistency backpressure to minimize GIL contention.
  • Shrank a deploy footprint from 14GB to 355MB and cut PR-to-prod from 30 minutes to under 7 by re-architecting multi-stage Docker builds.
  • Root-caused a production traffic-skew incident — recognized by engineering leadership — by isolating silent misrouting between builds during a Jenkins-to-ArgoCD migration.
  • Architected a hybrid retrieval system (two-tower NN + FAISS + BM25) scaling to 5M+ SKUs — lifting conversion 7.2% and click-through 34%, driving $300K+ ARR.
Quoted in Qualcomm's developer blog · On-Device AI Builders HackathonMS Computer Science · NYU

work

Selected projects

All repositories on GitHub

research

Published · peer-reviewed

Benchmarking Fine-Tuned Transformers, LLMs and LSTM Networks for Automated Essay Scoring

A comparative study of how transformer fine-tuning, large language models, and recurrent baselines trade off accuracy against the cost of scoring at scale — with the fine-tuned approach exceeding prior benchmarks by 5.7%.

writing

Deep dives

FeaturedQuoted by name on the power efficiency of on-device LLM inference. On-Device AI Builders Hackathon · Qualcomm × LM Studio × Microsoft, hosted at NYU Tandon