Imported from personamanagmentlayer/pcl (
stdlib/qa/chaos-engineering-expert/SKILL.md). Install upstream withnpx skills add personamanagmentlayer/pcl --skill chaos-engineering-expert. Copyright stays with the author.
Chaos Engineering Expert
Core Concepts
Chaos Engineering Principles
- Hypothesis-Driven - Define expected system behavior
- Production Testing - Test in real environments
- Minimize Blast Radius - Start small, expand gradually
- Automation - Continuous chaos experiments
- Learn and Improve - Build resilience iteratively
- Observability - Monitor system behavior
Failure Types
- Network Failures - Latency, packet loss, partitions
- Resource Exhaustion - CPU, memory, disk
- Service Failures - Process crashes, unavailability
- Data Corruption - Corrupt files, bad data
- Time Drift - Clock skew, NTP failures
- Dependency Failures - Third-party service outages
Tools & Platforms
- Chaos Monkey - Netflix's random termination tool
- Gremlin - Enterprise chaos engineering platform
- Chaos Toolkit - Open-source chaos experiments
- Litmus - Kubernetes chaos engineering
- Pumba - Docker chaos testing
- Toxiproxy - Network condition simulation
Best Practices
Experiment Design
- Start with hypothesis
- Define steady-state metrics
- Begin with small blast radius
- Test in staging first
- Automate experiments
- Document learnings
Safety Measures
- Implement circuit breakers
- Set up monitoring/alerting
- Have rollback procedures
- Limit blast radius
- Run during business hours initially
- Get stakeholder buy-in
Observability
- Monitor golden signals
- Track error rates
- Measure latency (p50, p95, p99)
- Monitor resource utilization
- Log all chaos events
- Correlate metrics
Culture
- Foster blameless culture
- Share learnings openly
- Make chaos regular practice
- Train teams on chaos engineering
- Start with game days
- Celebrate failures as learning
Anti-Patterns
Common Mistakes
- Testing in production without preparation
- No rollback plan
- Ignoring blast radius
- Running attacks during incidents
- No monitoring in place
- Blaming teams for failures
Experiment Design Issues
- No clear hypothesis
- Undefined success criteria
- Too broad scope initially
- Missing steady-state verification
- No automation
- Poor documentation
Cultural Problems
- Blame-focused culture
- Resistance to controlled failure
- Lack of stakeholder support
- No learning from experiments
- Chaos for chaos sake
- Security concerns ignored
Reference Documentation
Detailed material lives alongside this skill and is read on demand:
- Implementation Examples — Chaos Toolkit Experiment, Gremlin Attack Scenarios, Custom Chaos Tool (Python), Kubernetes Chaos with Litmus
Resources
Official Documentation
- Principles of Chaos Engineering - Core principles
- Chaos Toolkit - Tool docs
- Gremlin Documentation - Platform guide
- Litmus Documentation - Kubernetes chaos
Learning Resources
- Chaos Engineering Book - O'Reilly
- Google SRE Book - SRE practices
- Awesome Chaos Engineering - Resources
- Chaos Engineering YouTube - Videos
Tools & Platforms
- Chaos Monkey - Netflix tool
- Gremlin Free - Free tier
- Chaos Mesh - Kubernetes platform
- Toxiproxy - Network simulator
Community Resources
- Chaos Engineering Slack - Community
- Reddit SRE - Discussions
- Chaos Conf - Annual conference
- LinkedIn Groups - Professional networks