Production LLM traffic layerProduction
Envoy AI gateway
Envoy deployed as both general ingress and AI gateway, with model routing across Claude and GPT providers and all AI traffic telemetry flowing back into OpenObserve.
EnvoyIstioClaudeGPTOpenObserveKubernetes
architecturescroll to pan →
Problem
AI traffic needs a single control point: routing across model providers, rate limiting, and full cost visibility per model and per provider. Without it, spend and behavior are opaque.
My role
Deployed Envoy as both general ingress (nginx replacement) and AI gateway, and wired its telemetry into OpenObserve for AI observability.
What shipped
- +Envoy as general ingress and AI gateway
- +Model routing across Claude and GPT providers with rate limiting
- +All gateway and AI traffic telemetry flowing into OpenObserve
- +Per-model and per-provider cost tracking and traffic analysis
Outcome
- →AI SRE behavior verified by querying this telemetry
- →Customers query incidents to see what happened and why