Workload-specific systems work

LLM Inference Engineering

Measure, explain, and improve production inference behavior against the workload and SLO that matter to your application.

Workload first

Why workload characteristics matter

Prompt length, generated tokens, concurrency, and arrival patterns change the serving problem. A useful capacity or latency target starts with the real request distribution, not an isolated peak-throughput number.

Research evidence

A measured concurrency knee.

Our A100 investigation connects repeated vLLM serving measurements to scheduler behavior and GPU profiling. It shows why throughput and latency need to be read together.

Read the investigation
vLLM · A100 · ReproducibleBeyond Peak Throughput: Finding and Explaining vLLM’s Concurrency Knee on an A100Read the measured investigation

Capabilities

From serving trace to operating point.

Latency and throughput

Workload characterization, p95 TTFT, TPOT, throughput, concurrency, and latency SLOs.

Serving behavior

Batching, KV-cache behavior, scheduler decisions, and vLLM or SGLang configuration.

Capacity and design

GPU utilization, hardware/capacity planning, and dedicated deployment design.

Approach

Evidence before intervention.

  1. 01Measure workload
  2. 02Identify bottleneck
  3. 03Tune serving stack
  4. 04Validate against SLO
  5. 05Reproduce results

Technical conversation

Discuss an inference workload.

Share the model route, workload shape, latency or throughput target, serving stack, and constraints you can disclose. We will review the context and suggest a next step.