I'm an SRE / Platform Engineer — 16 years in IT, 8 of them building and running production platforms. I work with Kubernetes across AWS, Azure, GCP and Huawei Cloud, and I care most about the part most teams skip: SLIs, SLOs and error budgets that hold up when production disagrees with the plan.
Most of my recent work has been on a payments platform: multi-cluster Kubernetes, Terraform and Terragrunt at scale, and a full delivery migration from Azure DevOps to GitHub Actions.
| Orchestration | Kubernetes — EKS, AKS, GKE, Huawei CCE, ECS/Fargate |
| Infrastructure as code | Terraform and Terragrunt (since 2018) |
| Delivery | ArgoCD, GitHub Actions, Azure DevOps |
| Observability | Datadog, OpenTelemetry, Prometheus, Grafana, Loki |
| Reliability | SLIs, SLOs, error budgets, burn-rate alerting |
| Languages | Go, Python, Shell |
- LinkedIn — in/agnaldom
- Email — agnaldomarinho7@gmail.com




