Skip to main content

62. LLM Serving

A trained model needs an inference system to serve users.

A typical architecture is:

Client
↓
API
↓
Inference Server
↓
LLM
↓
Response

Serving systems optimize for:

  • latency
  • throughput
  • concurrency
  • GPU utilization
  • memory usage