Site Reliability Engineer
- Latin America, LATAM, United States
- POSTED 2 MONTHS AGO
About the job
We are looking for a Site Reliability Engineer based in Latin America to work on a long-term project for one of our clients, a software company based in San Ramon, California.
Our client provides cloud solutions trusted by governments worldwide to accelerate their digital transformation, deliver vital services, and build stronger communities.
Responsibilities
Lead deep technical analysis during critical incidents, coordinating cross-functional teams to identify root causes, restore services quickly, and drive long-term reliability improvements. Automation and Toil Reduction: Writing code and scripts (Python, Go, or Bash) to eliminate repetitive manual work and handle provisioning. Incident Management: Responding to system alerts, troubleshooting production issues, joining on-call rotations, and restoring services quickly. SLO and Error Budget Management: Establishing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to balance feature delivery speed with system stability. Monitoring and Observability: Building dashboards and setting up tools (such as Prometheus, Grafana, or Datadog) to track system health and latency. Post-Incident Reviews: Conducting blameless post-mortems to find root causes and prevent repeat failures. Capacity Planning: Analyzing resource usage trends to forecast future infrastructure and scaling needs. Requirements
Advanced Level of English. 5+ years of experience working as a Site Reliability Engineer 1+ years of experience within Microsoft Azure. Strong experience with Python, Bash or Go for automation purposes. Experience working with Monitoring and Observability tools such as Prometheus, Grafana, or Datadog. Ability to keep a focused mind during high-severity production outages to lead teams effectively. Proven experience explaining complex infrastructure failures to non-technical business stakeholders in clear, simple terms. Strong prioritization instincts, with the ability to quickly determine whether to resolve a live issue manually or engineer a permanent automated solution. Experience working with AI apps or agents. Bonus Points
Bachelor’s Degree in Computer Science, Systems Engineering or related fields. SaaS experience What we offer
Long term positions Compensation in USD Paid time off Cool clients and products Work with great engineers 4tech