Professional Summary
Overview
Work History
Education
Skills
Certification
LANGUAGES
Timeline

RAFAEL CARDOSO DOS SANTOS

CARDOSO CLOUD & SRE LTDA
Joinville
2
Languages
1
Certification
16
years of professional experience

Senior SRE Consultant delivers cloud infrastructure architecture for enterprise clients across AWS, GCP, and Terraform, delivering 3–5 environments per year. Designs 2–4 CI/CD pipelines and Kubernetes environments that improve deployment speed, reliability, and operating consistency. Builds 2–3 Prometheus and Grafana monitoring stacks for SLA and SLO compliance while supporting safe releases through Jenkins and GitHub Actions.

Work History

Founder & Senior SRE Consultant

2 Months
CARDOSO CLOUD & SRE LTDA | 06.2026 - Current
  • • Cloud Infrastructure Architecture: Designing, building, and optimizing scalable, highly available architectures on AWS and GCP using infrastructure as Code (IaC) with Terraform.
    • Containerization & Orchestration: Managing and maintaining high-performance Kubernetes clusters (EKS/GCP) and Docker environments to ensure seamless deployment and scaling.
    • Observability & Monitoring: Implementing advanced monitoring, logging, and alerting systems using Prometheus and Grafana to track SLA/SLO compliance and accelerate root-cause analysis.
    • CI/CD Pipelines: Engineering, automation, and optimizing continuous integration and delivery pipelines (Jenkins, Docker) to guarantee safe, fast, and repeatable deployments.
    • System Resilience & Reliability: Proactively monitoring system health, managing incident response strategies, and optimizing system architecture to prevent downtime and guarantee reliability.
  • Improved operational efficiency, adopting new technologies that automated routine tasks.
  • Conducted root cause analyses on incidents, driving changes to prevent future occurrences.

Senior Site Reliability Engineer

1 Year 8 Months
Stone | 07.2024 - 03.2026
  • Managed corporate messaging platform with Apache Kafka (Confluent, MSK), Google Pub/Sub, and Amazon SNS, serving 50+ microservices
  • Rebuilt and standardized messaging infrastructure with Terraform, collaborating with product teams to map resources and remove obsolete ones, achieving significant cost savings
  • Led migration of SQS and Rabbit MQ services to internal Karavela platform (Backstage), centralizing management and streamlining provisioning processes
  • Improved incident management workflows by creating comprehensive documentation on troubleshooting procedures and common issues resolution steps.

Site Reliability Engineer

1 Year 1 Month
Ebury Bank | 02.2023 - 03.2024
  • Ensured high availability and reliability of critical financial applications for Accounts Squad, enhancing operational stability
  • Integrating systems with zero downtime to maintain business continuity
  • Refined monitoring strategy, eliminated false positives, and reduced alert fatigue, resulting in more effective incident response
  • Led incident response efforts, ensuring swift resolution of service disruptions and minimizing downtime.
  • Conducted root cause analysis on incidents, driving continuous improvement initiatives to prevent future occurrences.
  • Collaborated with development teams to ensure systems architecture supported scalability and reliability requirements.
  • Improved incident management workflows by creating comprehensive documentation on troubleshooting procedures and common issues resolution steps.
  • Developed custom scripts/tools as needed to automate routine tasks, increasing overall team productivity and efficiency.

Senior Site Reliability Engineer

1 Year 10 Months
Tembici | 03.2021 - 01.2023
  • Managed AWS/GCP environments for bike-sharing platform serving 1M+ daily users, ensuring high availability and performance
  • Implemented CI/CD pipelines with GitHub Actions, integrating SonarCloud, Dependabot, and OWASP ZAP to enhance security posture
  • Optimized Kubernetes HPA settings, achieving 35% increase in resource efficiency
  • Managed capacity planning efforts to ensure optimal resource allocation based on current demand projections and future growth expectations.

Site Reliability Engineer

1 Year 3 Months
TOTVS | 12.2019 - 03.2021
  • Ensured reliability of Fluig SaaS platform, supporting critical operations for enterprise customers
  • Led infrastructure planning for peak demand periods, maintaining 100% availability during Black Friday
  • Applied FinOps and right-sizing strategies, achieving 30% reduction in AWS costs ($120K yearly savings)
  • Managed capacity planning efforts to ensure optimal resource allocation based on current demand projections and future growth expectations.
  • Led incident response efforts, ensuring swift resolution of service disruptions and minimizing downtime.
  • Developed scripts for deployment automation, enhancing efficiency in application releases and updates.
  • Collaborated with development teams to ensure systems architecture supported scalability and reliability requirements.
  • Mentored junior engineers in best practices for system reliability and incident management processes.
  • Improved incident management workflows by creating comprehensive documentation on troubleshooting procedures and common issues resolution steps.
  • Developed custom scripts/tools as needed to automate routine tasks, increasing overall team productivity and efficiency.

Systems Administrator

9 Years 2 Months
TOTVS | 10.2010 - 12.2019
  • Managed JBoss, Tomcat, and WildFly servers for enterprise ERP systems, ensuring high availability in production environments
  • Managed JBoss, Tomcat, and WildFly servers for enterprise ERP systems in production environments
  • Administered Oracle, SQL Server, and Progress databases, maintaining 99.9% availability across all environments
  • Devised scripts and automation tools to improve system efficiency.

Education

Bachelor's Degree - Information Systems

UniSociesc | Joinville, SC, Brazil | 11.2017

Skills

AWS
Google Cloud Platform (GCP)
CI/CD
Kubernetes
Terraform
Docker
FinOps
Observability engineering
Incident management
Alert tuning
Capacity planning
Cloud architecture
Infrastructure automation
Platform modernization
Migration planning
Cost optimization

Certification

  • Copa AIOps - Matié (June 2026)
  • EF SET English Certificate 53/100 (B2 Upper Intermediate) - EF SET (June 2026)
  • Imersao Agentes IA - Hashtag Treinamentos (July 2026)
  • Automacao Inteligente com N8N para Iniciantes - Udemy
  • Fundamentos de Redes para DevOps - Udemy
  • Grafana: Dashboards Gerenciais + Monitoramento - Udemy
  • DevOps na Nuvem com IA - Hashtag Treinamentos

LANGUAGES

Portuguese: Native
English: B2 Upper Intermediate (EF SET Certified - 53/100)

Timeline

Founder & Senior SRE Consultant

CARDOSO CLOUD & SRE LTDA
06.2026 - CurrentRead More

Senior Site Reliability Engineer

Stone
07.2024 - 03.2026Read More

Site Reliability Engineer

Ebury Bank
02.2023 - 03.2024Read More

Senior Site Reliability Engineer

Tembici
03.2021 - 01.2023Read More

Site Reliability Engineer

TOTVS
12.2019 - 03.2021Read More

Systems Administrator

TOTVS
10.2010 - 12.2019Read More

UniSociesc

Bachelor's Degree from Information Systems
Read More
RAFAEL CARDOSO DOS SANTOS