Marcelo Bertho
Tech Lead SRE / DevOps / Platform Engineering · Cloud · AI for Infrastructure
Campinas, Brazil
Technology leader with 20+ years of experience and a strong recent track record in SRE, DevOps, Cloud and Platform Engineering. I've led cloud-native transformations, built engineering teams and operated mission-critical platforms across AWS, GCP and Azure.
Deep expertise in Kubernetes/EKS, Terraform, GitOps, CI/CD, observability, service mesh, reliability engineering and FinOps. I've operated systems under extreme traffic — major ticketing events, Black Friday, national TV campaigns and billion-message-per-month workloads.
Today my focus is building AI Agents for Infrastructure & SRE: automated troubleshooting, signal correlation across AWS, Datadog and Cloudflare, and intelligent incident response. It's the foundation of my AI Ops platform — taking AI from experiment to reliable operation.
Focus areas
- AI Agents for Infrastructure & SRE (AI Ops)
- SRE: SLA/SLO, MTTR/MTTD, high availability, DR (RPO/RTO)
- Platform Engineering & Internal Developer Platform (IDP)
- Multi-cloud: AWS · GCP · Azure
- Kubernetes/EKS, Terraform, GitOps (ArgoCD), service mesh (Istio)
- Observability: Datadog, Prometheus, Grafana, New Relic
- FinOps, incident management and postmortems
Experience
I lead cloud infrastructure, SRE and platform engineering for large-scale, high-concurrency ticketing systems (Brazil National Team qualifiers, Guns N' Roses, Festival de Parintins, Tardezinha). I build AI Agents for Infra & SRE for automated troubleshooting, cross-system analysis and faster incident response. Stack: AWS, EKS, Terraform, ArgoCD, Istio, Vault, Datadog, GitOps.
Led the cloud-native transformation for agribusiness (company controlled by LDC, ADM, Cargill and Amaggi), sustaining seasonal peaks and national TV campaigns with zero downtime. AWS, GCP, Kubernetes, Terraform, ArgoCD, Istio, New Relic.
Led the transformation into a cloud-native organization, with SRE, DevOps and DataOps on Kubernetes, AWS and GCP. Datadog, New Relic, Terraform, Crossplane, ArgoCD.
Prepared and operated infrastructure for Black Friday and national TV campaigns of Brazil's largest pet e-commerce, under extreme traffic. Kubernetes, Istio, GitOps, IaC and observability.
Led DevOps and DataOps initiatives in Brazil as part of a global team, building one of the largest enterprise data platforms. Azure, Azure DevOps, Key Vault, Talend, Power BI.
Operated messaging platforms at massive scale — ~1.2B SMS and 30M WhatsApp messages per month. 900+ multi-cloud instances (AWS, GCP, Azure, Oracle, VMware, XenServer).
SISFRON project (Integrated Border Monitoring System — Brazilian Army): highly available mission-critical infrastructure, CI/CD and a NOC for national communications monitoring.
Led infrastructure, CI/CD and IT operations for the rapid expansion of Samsung's Brazil R&D center — from ~40 to 450+ employees. CMMI Level 2/3 processes.
Earlier experience: IBM, FITec and CPQD as Configuration Management Analyst (2002–2006) — release management, SCM governance, CMMI Level 2/3, ClearCase/Perforce and Perl/CSH automation.
Education
Ready to put your AI into operation?
Explore the AI Ops platform or talk to me directly about your infrastructure and AI challenge.
Book a call