Chaos Engineering and Testing in Live Systems — feat. Kolton Andrus (Gremlin)
KOLTON ANDRUS — Kolton Andrus — Gremlin
The Gremlin CEO on chaos engineering, resilience and finding failure before your customers do.
Show notes
The Gremlin CEO on chaos engineering, resilience and finding failure before your customers do. Kolton Andrus, who pioneered failure-injection tooling at Amazon and Netflix before founding Gremlin, explains how precise, application-level chaos experiments raised availability while cutting paging by 25%. He argues against separate QA and ops teams as an anti-pattern, insisting the team that writes code should own its quality and resilience. Distinguishing traditional from modern testing, he stresses that distributed systems fail through third-party dependencies, so engineers must test what happens when a dependency fails or slows, building timeouts, fallbacks and graceful degradation. He favors real user monitoring and canary traffic over synthetic data, ties reliability back to brand and revenue, and describes building a shared catalog of failure scenarios so teams can find and fix the lo