HardPhase 8 — LLM inference and performance engineering

Week 93 — Disaggregation

LLM inference and performance engineering · Read / diagram

Hard

Phase 8 — LLM inference and performance engineering

Overview

DistServe, Splitwise, Mooncake, autoscaling; Kubernetes topology (KServe + KEDA + HPA) and edge deployment (ONNX, TensorRT-LLM, WebLLM) as alternative topologies

Mode this week: read_diagram. Match the work to that mode (operating model): implement ships running code; read/diagram ships a doc; deploy/benchmark ships a measured run. Personal dates, checkboxes, and the RPG layer stay in a learner journal.

Core deliverable

Disaggregated serving plan + depth-mode sub-deliverable: deploy the W84 pipeline to K8s with HPA-driven autoscaling on a synthetic spike AND port one vLLM config to ONNX/TensorRT-LLM for edge benchmarking (Modal A10 fallback if no edge hardware)

Optional depth

Add a production-shaped benchmark, cost or safety analysis, and an architecture trade-off note.

Week loop (from the cookiecutter journal)

Forge theme: Week 93 - Inference boss

Leetcode Darbar (parallel)

This week's band (82-102): Heaps, priority queues, graphs, trees, range queries. Production connection: Inference schedulers, quotas, queues, tenant fairness.

Standard: one aligned problem. Auror: the five-slot queue on the Darbar page. Tracker stays in the learner journal.

Topics: graphs · heap · queue · range query · trees

Evidence contract

Every core week records all four (cookiecutter journal/weeks/week_093.md):

  1. Code / implementation (or the design artifact on a writing week)
  2. Benchmark / result
  3. Design doc / technical explanation
  4. Retrospective / learning note (Sailboat: destination, wind, anchor, rocks, heading)

On this site

Official Tensor-to-Tenant


Week 93 of Tensor-to-Tenant · Previous week · Next week.