Rust + cuTile research prototype for paged latent-cache LLM decode attention, validated on an RTX 4060.
-
Updated
Jul 24, 2026 - Rust
Rust + cuTile research prototype for paged latent-cache LLM decode attention, validated on an RTX 4060.
A practical handbook for software engineers to learn AI, Large Language Models (LLMs), and Inference Engineering—from fundamentals to production systems.
LLM inference benchmarking dashboard: Python FastAPI backend with async orchestration, WebSocket live TTFT/TBT/throughput comparison across configs (512/128 to 4096/1024 tokens), Grafana + Docker Compose stack, GitHub Actions CI; 21/21 pytest passing.
Offline, evidence-first bottleneck hypotheses for LLM inference traces
Deterministic simulator for KV-cache admission, placement, movement, eviction, and recomputation policies
Trace-Aware Serving Controller: eval-gated inference policy optimization
Deep Agents and SvelteKit harness for authoring verifier-gated Bonsai workflow packs.
An interactive playground for learning inference engineering—explore LLM serving concepts, tune the stack, and graduate to production incidents.
Executable field guide and deterministic labs for LLM inference engineering across kernels, scheduling, KV cache, placement, and control.
Add a description, image, and links to the inference-engineering topic page so that developers can more easily learn about it.
To associate your repository with the inference-engineering topic, visit your repo's landing page and select "manage topics."