← back to case studies
Production LLM traffic layerProduction

Envoy AI gateway

Envoy deployed as both general ingress and AI gateway, with model routing across Claude and GPT providers and all AI traffic telemetry flowing back into OpenObserve.

EnvoyIstioClaudeGPTOpenObserveKubernetes
architecturescroll to pan →
Clientsingress trafficEnvoyingress + AI gatewayrouterouteClauderate limitedGPTrate limitedOpenObserveper-model costtelemetry

Problem

AI traffic needs a single control point: routing across model providers, rate limiting, and full cost visibility per model and per provider. Without it, spend and behavior are opaque.

My role

Deployed Envoy as both general ingress (nginx replacement) and AI gateway, and wired its telemetry into OpenObserve for AI observability.

What shipped

  • +Envoy as general ingress and AI gateway
  • +Model routing across Claude and GPT providers with rate limiting
  • +All gateway and AI traffic telemetry flowing into OpenObserve
  • +Per-model and per-provider cost tracking and traffic analysis

Outcome

  • AI SRE behavior verified by querying this telemetry
  • Customers query incidents to see what happened and why