Site Reliability Operations Engineer

  • Latin America, LATAM, United States
  • POSTED 1 MONTH AGO

About the job

We are looking for a Site Reliability Operations Engineer based in Latin America to work on a long-term project for one of our clients, a software company based in San Ramon, California.

Our client provides cloud solutions trusted by governments worldwide to accelerate their digital transformation, deliver vital services, and build stronger communities.

Responsibilities

Ensure the reliability, performance, and security of cloud-based applications and infrastructure, encompassing servers, networks, and databases. Respond quickly to any platform or performance-issue alerts to minimize downtime. Restore the application services ASAP. Collaborate to identify and resolve issues, enhancing the overall quality of applications. Develop and implement new alerts, as well as create standard operating procedure documents for supporting and maintaining web applications effectively. Requirements

Advanced Level of English. 2+ years of experience working as an SRO. 1+ years of experience working with Azure and AWS. 1+ years of experience working with Linux System Administration. Experience working with monitoring tools such as Datadog and PagerDuty. Experience working with Bash or PowerShell for scripting. Familiarity with scripts and basic automation. Experience working with AI apps or agents. Bonus Points

Bachelor’s Degree in Computer Science, Systems Engineering or related fields. What we offer

Long term positions Compensation in USD Paid time off Cool clients and products Work with great engineers 4tech