Post-Mortem: Recovering from a Kubernetes Cluster Outage
A detailed breakdown of how a misconfigured HPA led to node exhaustion, and the steps we took to restore service and prevent future occurrences.
Technical write-ups, post-mortems, and architectural deep-dives.
A detailed breakdown of how a misconfigured HPA led to node exhaustion, and the steps we took to restore service and prevent future occurrences.
Best practices for using S3 backends with DynamoDB state locking, KMS encryption, and IAM role assumption for isolated environments.
How we reduced our container image sizes by 70% and sped up GitHub Actions build times using advanced multi-stage caching strategies.
The infrastructure challenges of deploying our Varicose Vein detection model to resource-constrained edge environments.