Principal Solutions Architectlive shell · type `help` · Ctrl+` to toggle

Chaitanya Sistla

chaitanya@prod: ~ — zsh
chaitanya@prod:~$
150+TB/day

Peak ingestion supported and scaled

Fortune 500 observability workloads

5PiB

Big data managed

Cloudera Hadoop, 300 nodes

1200+servers

Hardened overnight

Chef to Ansible, vulnerabilities removed

25nodes

Bare-metal Confluent Kafka

Terraform-imported, Ansible-managed

2DCs

Aerospike cross-DC replication

Data replicated between data centers

30-40%

Compute cost reduction vs Elastic

At equivalent workloads

2months

Competitive win to production

From POC kickoff, every criterion met

25%

Monthly cloud spend reduction

ClickOps to Terraform at Securonix

300nodes

Cloudera Hadoop cluster operated

Kerberos, TLS, Vault dynamic secrets

17+certs

Active certifications

AWS, Azure, Kubernetes, HashiCorp

70%

Routine operations automated

Big data platform at Early Warning

$500M→ $1B

Business scale supported

Compliant cloud infrastructure at ONEngine

$120Krevenue

Product built on Google ADK

Shipped and monetized

RAGengine

Context engine for cloud operations

Grounded in live infrastructure state

5systems

LLM systems in production

AI SRE, gateway, triage, and more

2gateways

AI gateways in production

Envoy and Portkey, Claude and GPT routing

chaitanya@prod:~$cat architecture/system-map# 00

# System Map

>How the platforms I ship fit together: telemetry in through OpenTelemetry, OpenObserve as the hub, production systems built on top.

Logs · metrics · tracesSecurity eventsIAM · net · CloudTrailWindows · Linux · VMwareOTelcollectorsOpenObserve150TB/day hubDashboards & alertsobservability-as-codeAI SREincident correlationSIEM detectionOCSF · Sigma → VRLEnvoy AI gatewayrouting · cost telemetryAI traffic telemetry flows back for cost tracking
chaitanya@prod:~$./atlas --list --anonymized# 01

# Deployment Atlas

>Every production deployment pattern I have shipped. Filter by platform or technology. Customers, partners, and competing vendors are anonymized on purpose.

platform
technology
17 deployments shown

Competitive observability win

Production

Large enterprise

POCtuned clustermigrationproduction

Sole technical lead, every success criterion met.

scale: POC to production in 2 months
Multi-cloudEKSOpenObserveTerraformKubernetesOpenTelemetry

150TB+/day scale engagements

Ongoing

Fortune 500

Elastic/GrafanamigrationOpenObserve150TB+/day

30-40% less compute than incumbent Elastic.

scale: 150TB+/day peak ingestion
EKSMulti-cloudOpenObserveClickHouseElasticsearchKubernetesTerraform

AI SRE incident correlation

Production

One of the world's largest retailers

alertscorrelationincidentrunbook

Alert storms consolidated into single incidents.

scale: Production on-call workflows
Multi-cloudEKSOpenObserveClaudeAnthropic SDKOpenTelemetryKubernetes

Envoy AI gateway

Production

Platform layer

clientsEnvoy / PortkeyClaude/GPTOpenObserve

Model routing with per-provider cost telemetry.

scale: Production LLM traffic
EKSMulti-cloudEnvoyPortkeyIstioClaudeGPTOpenObserveKubernetes

SIEM detection pipeline

Internal

Security engineering

IAM/net/CloudTrailOCSFSigma→VRLalerts

OCSF normalization with Sigma-to-VRL detections.

scale: Running internally, design partner engaged
Multi-cloudOpenObserveOCSFSigmaVRLThreat intel

Claude security triage automation

Internal

Security engineering

incidentsClaudevalidated output

Repeatable triage over incidents and findings.

scale: Repeatable internal automation
Multi-cloudClaudeAnthropic SDKOpenObserve

Enterprise onboarding breadth

Production

Cross-industry

EKS/AKS/OKESSO/OIDCOpenObserve

EKS/AKS/OKE/DOKS/bare metal, SSO, BYO bucket.

scale: Multi-platform, multi-tenant
EKSAKSOKEDOKSBare metalKubernetesOktaKeycloakEntra IDDex/OIDCOpenTelemetry

OTel collector fleet

Production

Cross-industry

Windows/Linux/VMwareOTelOpenObserve

Collectors across Windows, Linux, VMware, K8s.

scale: Heterogeneous infrastructure
Bare metalMulti-cloudOpenTelemetryLinuxVMwareKubernetes

GovCloud compliant ingestion

Production

Federal and defense

GovCloud/Azuretemplatescompliant ingestion

Compliant federal and Azure telemetry ingestion.

scale: Compliant federal workloads
Multi-cloudCloudFormationARMAWS GovCloudAzure

Multi-cloud SIEM/UEBA platform

Production

Security vendor

ClickOpsTerraform + Jenkins−25% spend

ClickOps to Terraform, 25% monthly spend cut.

scale: 25% monthly cloud spend reduction
Multi-cloudTerraformJenkinsAWSKubernetes

Bare-metal Confluent Kafka

Production

Regulated banking

25 bare-metalTerraform importAnsible servicesConfluent Kafka

Confluent Kafka on 25 bare-metal servers, imported into Terraform.

scale: 25 bare-metal servers
Bare metalOn-premConfluent KafkaTerraformAnsibleBare metal

Cloudera Hadoop big data cluster

Production

Regulated banking · big data

ingest300-node Cloudera HadoopKerberos/TLS/Vault5PiB

300-node Cloudera Hadoop cluster managing 5PiB of big data.

scale: 300 nodes, 5PiB managed
On-premBare metalCloudera / HadoopKerberosVaultBare metal

Aerospike cross-DC replication

Production

Regulated data (GDPR)

Aerospikereplication2 data centers

Aerospike with cross-data-center replication.

scale: Replicated across 2 data centers
On-premBare metalAerospikeGDPR

Solr to Elasticsearch migration

Production

Regulated data

SolrmigrationElasticsearch

Deployed Elasticsearch and migrated search off Solr.

scale: Search platform consolidation
On-premBare metalElasticsearchSolr

Chef to Ansible fleet hardening

Production

Regulated infrastructure

1200+ serversChef → Ansiblevulnerabilities removed

Converted 1200+ servers from Chef to Ansible overnight, removing vulnerabilities.

scale: 1200+ servers hardened overnight
On-premBare metalAnsibleChefLinux

Kerberos and mTLS machine identity

Production

On-prem security

KerberosmTLSmachine identity

Kerberos on-prem plus mTLS server-to-server, machine identity at scale.

scale: Machine identity at large scale
On-premBare metalKerberosmTLSVault

Jenkins mini cloud manager

Production

Platform engineering

ClickOpsJenkins automationself-service

Built a Jenkins mini cloud manager, eliminated most ClickOps.

scale: Eliminated most ClickOps
On-premMulti-cloudJenkinsTerraformAnsible
chaitanya@prod:~$ls ~/stack# 02

# Tech Stack

>The technologies I work with across the stack, grouped by domain.

Orchestration
KubernetesHelmArgoCD / GitOpsDocker
IaC
Terraform / OpenTofuCloudFormationAnsibleChefARM templates
Observability
OpenObserveOpenTelemetryPrometheusGrafanaDatadog
Cloud
AWSAzureGCPOCIGovCloud
Networking
IstioEnvoyVPC peeringPrivate LinkVPN / tunnelingCIDR planning
Security
VaultSSO / OIDCSOC 2 / ISO 27001VRLSigmaOCSFOPAGDPRKerberosmTLSLLM guardrailsPrompt-injection defenseAI gateway securityLLM output validationTrivy / tfsec / CheckovNuclei / Prowler / Custodian
Data
ClickHouseElasticsearchKafkaPostgresRedisMongoDBSnowflakeAerospikeSolrCloudera / HadoopDatabricksMySQL / SQL ServerTableau / Power BI
AI/LLM
Anthropic SDK / ClaudeGPT APIsPortkeypgvector / RAG
Languages
LinuxBashPythonGoJavaScriptReactRust
CI/CD
GitHub ActionsJenkinsHarness
chaitanya@prod:~$ls -la case-studies/# 03

# Case Studies

>The flagship engagements, each with an architecture diagram, the problem, my role, and the outcome. Read any of them end to end.

Large enterprise · sole technical leadProduction

Competitive observability win

A large enterprise evaluated OpenObserve against a leading observability vendor. I was the sole technical lead and took it from discovery to production in two months.

OpenObserveTerraformKubernetesOpenTelemetry
read case study →
Fortune 500 · scale and costOngoing

Observability at 150TB+/day

Fortune 500 POC engagements scaling to 150TB+/day ingestion, with a 30-40% compute cost reduction versus incumbent Elastic stacks at equivalent workloads.

OpenObserveClickHouseElasticsearchKubernetesTerraform
read case study →
One of the world's largest retailers · productionProduction

AI SRE: LLM-driven incident correlation

A production system that consolidates correlated alerts into a single incident, generates a root-cause summary, and produces a virtual runbook grounded in the customer telemetry.

OpenObserveClaudeAnthropic SDKOpenTelemetryKubernetes
read case study →
Production LLM traffic layerProduction

Envoy AI gateway

Envoy deployed as both general ingress and AI gateway, with model routing across Claude and GPT providers and all AI traffic telemetry flowing back into OpenObserve.

EnvoyIstioClaudeGPTOpenObserve
read case study →
Running internally, design partner engagedInternal

SIEM detection pipeline

A detection pipeline on OpenObserve with OCSF normalization, risk scoring, and Sigma rules converted to VRL, running internally with an enterprise design partner engaged, ahead of planned productization.

OpenObserveOCSFSigmaVRLThreat intel
read case study →
Repeatable triage automationInternal

Claude-based security triage

Repeatable Claude-based triage automation for security incidents and discovery findings at OpenObserve.

ClaudeAnthropic SDKOpenObserve
read case study →
Regulated banking · on-premProduction

Bare-metal Confluent Kafka platform

Confluent Kafka on 25 bare-metal servers, with the live physical servers imported into Terraform and every service managed by Ansible.

Confluent KafkaTerraformAnsibleBare metal
read case study →
Regulated banking · big data · 5PiBProduction

Cloudera Hadoop big data cluster

A 300-node Cloudera Hadoop cluster managing 5PiB of banking big data, secured with Kerberos, TLS, and Vault dynamic secrets, with 70% of routine operations automated.

Cloudera / HadoopKerberosVaultBare metal
read case study →
Regulated data · GDPR · multi-DCProduction

Cross-data-center data platform

Aerospike with replication across data centers, and a search platform migrated from Solr to Elasticsearch, in a regulated GDPR-scoped data environment.

AerospikeElasticsearchSolrGDPR
read case study →
On-prem security · Kerberos + mTLSProduction

Machine identity at scale

Kerberos implemented on premises and mTLS for server-to-server communication, establishing machine identity at large scale across a regulated on-prem fleet.

KerberosmTLSVault
read case study →
chaitanya@prod:~$cat ai-systems/*.md# 04

# LLM Systems in Production

>Production LLM systems, shown as architecture with the guardrails made explicit. Status language is precise on purpose.

AI SRE incident correlation

Production
logs·metrics·tracescorrelationincidentrunbook

Alert storms into one grounded incident.

grounded in telemetrytuned per environmentvalidated before on-call
view case study →

Envoy and Portkey AI gateway

Production
clientsEnvoy / PortkeyClaude/GPTOpenObserve

Routing and rate limiting with cost telemetry.

rate limitingper-provider cost trackingfull traffic telemetry
view case study →

Claude security triage

Repeatable internal automation
incidents/findingsClaudevalidated output

Repeatable first-pass triage automation.

structured validated outputscoped to triage
view case study →

ops0 Claude architecture

Founder project, past tense
findings graphClaudevalidated HCL/JSON

Grounded remediation with strict validation.

JSON/HCL validationSELECT-only SQLdynamic guardrailsenrichment-gatedtoken + cost metering

LLM Sigma-to-VRL conversion

In progress
SigmaLLMVRLvalidated

LLM converts Sigma rules to VRL.

validated against real telemetry
view case study →
chaitanya@prod:~$journalctl -u incidents --since prod# 05

# War Stories

>Production incidents, told straight: symptom, diagnosis, fix, and the lesson that stuck.

incident

EKS AMI upgrade crashing ingesters

AMI upgradeingesters crashnode image root causeresolved
symptom
After an EKS AMI upgrade, ingester pods began crashing and ingestion stalled.
diagnosis
Traced the failure to the node image change introduced by the AMI upgrade rather than the application itself.
fix
Diagnosed and resolved the crash so ingesters returned to healthy operation.
lesson
Node-level changes are application-level incidents. Treat AMI upgrades as first-class changes with the same scrutiny as app deploys.
incident

Missing HTTPRoutes causing 404s

404smissing HTTPRoutesroutes restoredresolved
symptom
Requests returned 404s where routes were expected to resolve.
diagnosis
Identified missing HTTPRoutes as the cause of the failed routing.
fix
Restored the missing HTTPRoutes and traffic resolved correctly.
lesson
Gateway routing config is part of the deploy surface. Missing route objects fail quietly as 404s, so verify routes as part of rollout.
incident

Query tier outage

query tier downled investigationrestoredresolved
symptom
The query tier went down, cutting off read access to telemetry.
diagnosis
Led the investigation into the query tier failure under live outage conditions.
fix
Led resolution and brought the query tier back to service.
lesson
Read-path availability matters as much as ingestion. An outage on query is an outage for every dashboard and alert that depends on it.
incident

Vulnerabilities across 1200+ servers, overnight

1200+ serversChef driftChef → Ansible overnightvulnerabilities removed
symptom
More than 1200 servers carried known vulnerabilities, with configuration drifting under Chef.
diagnosis
Remediation had to be fleet-wide and fast. The Chef-managed config had drifted enough that per-host fixes would not hold.
fix
Converted the entire fleet from Chef to Ansible overnight, bringing configuration under one consistent, idempotent layer and removing the vulnerabilities across 1200+ servers.
lesson
A uniform, idempotent config layer is a security control. Fleet-wide remediation is only possible when every host is managed the same way.
incident

Live physical servers, no rebuild allowed

physical serversno recreationterraform importmanaged as code
symptom
A regulated banking platform ran Kafka on 25 physical servers provisioned by hand, outside any version control.
diagnosis
The fleet had to come under infrastructure as code, but nothing could be recreated: the servers were live and in production.
fix
Imported the existing physical servers into Terraform state so they were managed as code in place, then used Ansible for services management, with no recreation.
lesson
Terraform import turns brownfield physical infrastructure into managed code without a risky rebuild. Import first, converge second.
chaitanya@prod:~$git log --oneline field-to-product/# 06

# Field to Product

>I turn one-off customer work into reusable assets. The loop runs from a pattern I hit in delivery to a product PR to something the next customer inherits.

01Customer patternA recurring need surfaces in delivery02Field buildI build it hands-on in the engagement03Product PRThe pattern becomes a product capability04Reusable assetCodified for the next customerthe next customer inherits it
Provider

OpenObserve Terraform provider

Built and maintain the Terraform provider and Kubernetes modules, published to the Terraform Registry.

Product PR

Observability-as-code product PRs

Authored product PRs enabling dashboards and alerts to export as observability-as-code.

Tooling

Datadog-to-OpenObserve migration tooling

Dashboard migration tooling to move off Datadog.

Library

Prebuilt dashboard library

A library of dashboards for common infrastructure and application monitoring.

Templates

GovCloud and Azure templates

CloudFormation templates for AWS GovCloud compliant ingestion, and ARM templates for Azure telemetry.

Extension

OpenObserve Lambda Extension

A Lambda extension project for telemetry from serverless workloads.

Product PR

Field feedback that changed the product

UX changes, trace pipeline improvements, usage dashboards, and permission fixes driven by delivery feedback.

compliance, run from the vendor side
  • +Achieved SOC 2 Type II and ISO 27001 for OpenObserve: continuous control monitoring, formal risk assessments.
  • +Partnered with customer CISOs on risk management and security assessments.
  • +Worked within GDPR obligations on regulated data platforms.
  • +HIPAA and FINRA adherence on regulated banking data platforms.
chaitanya@prod:~$ls ~/oss && cat certs.txt# 07

# Open Source and Founder Work

>Selected projects I built, plus the certifications behind the depth. ops0 is a founder and shareholder role, non-operational, described in past tense.

ops0

AI-native cloud security
ops0.com

Founder and shareholder, non-operational

An AI-native preventive cloud security platform. The positioning was fix and govern, not find and alert. What I built:

  • +Multi-cloud discovery pipeline: read-only scans of AWS, GCP, and Azure, normalized and enriched by 150+ provider enrichers, correlated into a relationship graph, matched against IaC inventory for managed-versus-orphan and drift.
  • +Security scanning across three engines (Nuclei, Prowler, Cloud Custodian), correlated by resource and vulnerability class into confidence verdicts, mapped to CIS, NIST, SOC 2, PCI, HIPAA, and ISO.
  • +Resource relationship graph: 100+ metadata-driven rules per provider, hard and soft dependencies, blast-radius analysis, and Kahn topological sort for safe Terraform import ordering.
  • +Claude-powered analysis on the Anthropic SDK: grounded single-turn remediation and Terraform generation over a knowledge graph of findings, with dynamic guardrails, strict output validation, and per-call token and cost metering.

+ 3 more

Oxid

Open source · Rust
oxid.sh

Creator

An open-source Rust-based parallel infrastructure-as-code execution engine.

  • +Parallel execution engine for infrastructure as code
  • +Written in Rust

ops0 CLI

Open source · Product Hunt
github.com/ops0-ai/ops0-cli

Creator

An open-source CLI, launched on Product Hunt.

  • +Open-source command line tool
  • +Launched on Product Hunt

OpenObserve Terraform provider

Terraform Registry

Author and maintainer

The Terraform provider and Kubernetes modules for OpenObserve, published to the Terraform Registry.

  • +Terraform provider for OpenObserve
  • +Kubernetes modules
  • +Published to the Terraform Registry
certifications · 17+ active certificationsfull list on LinkedIn →
AWS Solutions Architect ProfessionalAWS DevOps Engineer ProfessionalAWS Security SpecialtyAWS Advanced Networking SpecialtyAWS Machine Learning SpecialtyAzure Solutions Architect ExpertCertified Kubernetes Administrator (CKA)Certified Kubernetes Application Developer (CKAD)Certified Kubernetes Security Specialist (CKS)HashiCorp Terraform AssociateHashiCorp Vault AssociateSnowPro Advanced Architect
chaitanya@prod:~$./book-call --schedule# 08

# Contact

>Book a call directly, or reach out on LinkedIn or GitHub.

cal.com/chaitanyasistla/15min

No phone, no address: reach me through the channels above.