62. LLM Serving
A trained model needs an inference system to serve users.
A typical architecture is:
Client
↓
API
↓
Inference Server
↓
LLM
↓
Response
Serving systems optimize for:
- latency
- throughput
- concurrency
- GPU utilization
- memory usage
A trained model needs an inference system to serve users.
A typical architecture is:
Client
↓
API
↓
Inference Server
↓
LLM
↓
Response
Serving systems optimize for: