Site Reliability Operations Engineer
- Latin America, LATAM, United States
- POSTED 1 MONTH AGO
About the job
We are looking for a Site Reliability Operations Engineer based in Latin America to work on a long-term project for one of our clients, a software company based in San Ramon, California.
Our client provides cloud solutions trusted by governments worldwide to accelerate their digital transformation, deliver vital services, and build stronger communities.
Responsibilities
Ensure the reliability, performance, and security of cloud-based applications and infrastructure, encompassing servers, networks, and databases. Respond quickly to any platform or performance-issue alerts to minimize downtime. Restore the application services ASAP. Collaborate to identify and resolve issues, enhancing the overall quality of applications. Develop and implement new alerts, as well as create standard operating procedure documents for supporting and maintaining web applications effectively. Requirements
Advanced Level of English. 2+ years of experience working as an SRO. 1+ years of experience working with Azure and AWS. 1+ years of experience working with Linux System Administration. Experience working with monitoring tools such as Datadog and PagerDuty. Experience working with Bash or PowerShell for scripting. Familiarity with scripts and basic automation. Experience working with AI apps or agents. Bonus Points
Bachelor’s Degree in Computer Science, Systems Engineering or related fields. What we offer
Long term positions Compensation in USD Paid time off Cool clients and products Work with great engineers 4tech