Robert Teah

DevOps Engineer

Cloud Engineer

Infra Engineer

Site Reliability Engineer

DevSecOps Engineer

Robert Teah

DevOps Engineer

Cloud Engineer

Infra Engineer

Site Reliability Engineer

DevSecOps Engineer

Senior DevOps / Platform / SRE Engineer with 10+ years of experience operating production infrastructure at enterprise scale across AWS and Kubernetes environments.

The 503 That Route53 Didn’t Cause

September 27, 2026 Uncategorized

At Luxottica, during a promotional campaign load spike, about 30% of product catalog requests started returning 503s — not all users, an intermittent pattern that made it harder to diagnose than a clean outage would have been.

What happened

The first hypothesis was HPA scaling lag, but scaling had actually completed four minutes before the campaign started — not the cause. ALB access logs showed the 503s were coming from the ALB itself, reason “TargetGroupNotFound.” The product catalog service had two target groups registered: one healthy with all pod IPs, one completely empty. A recent Terraform apply had created a new target group for a deployment but never destroyed the old one, and the Kubernetes Ingress annotation still referenced the old, now-empty target group’s ARN. Terraform managed the target group lifecycle; Kubernetes managed which ARN the ALB listener used — neither system had visibility into what the other was doing, and the stale reference sat there invisibly until traffic hit it.

The fix

Removing the stale annotation meant a kubectl apply against production Ingress during a live, revenue-impacting traffic window. I applied it directly via kubectl — two minutes, versus twelve to fifteen minutes through the full CI/CD pipeline — with the tradeoff explicitly logged and approved in the incident Slack channel. The Ingress controller reconciled within 60 seconds and the 503 rate dropped to zero within 90 seconds. Total incident duration: 34 minutes. Afterward, I added a CI/CD pipeline gate that validates every ALB target group ARN in an Ingress annotation has registered, healthy targets before any deployment is allowed to promote, and had Terraform’s output automatically update the Kubernetes Ingress annotation with the current ARN on every infrastructure apply.

The lesson

When two infrastructure systems manage overlapping resources — Terraform owning target group lifecycle, Kubernetes owning which target group the traffic actually points at — neither has full visibility into the state the other controls. That seam is where operational drift accumulates invisibly, and it doesn’t show up in either system’s own health checks. Validate state at the seam, not just within each system separately.

Write a comment