Lead Infrastructure Operations Engineer
- Iowa - TMG Home Office · Remote US
- Full-time
- POSTED 2 DAYS AGO
About the job
Department:
Information Technology
Job Description:
The Lead Infrastructure Operations Engineer is responsible for enabling, coordinating, and improving the infrastructure capabilities needed to support AI-enabled use cases.
As TMG expands its AI capabilities, the AI CoE and delivery teams will develop large language model, retrieval-augmented generation, agent-based, machine-learning, and other AI-enabled use cases. This role will work within the Infrastructure team to help those teams define and coordinate their cloud, environment, connectivity, observability, access, security, capacity, and production-readiness needs.
This role focuses on the infrastructure and operational foundation required to move AI-enabled solutions safely and reliably into production. TMG’s managed services provider operates the underlying AWS cloud infrastructure. This role will translate solution requirements, coordinate infrastructure services and changes, monitor delivery, diagnose infrastructure-related issues, and validate that environments meet reliability, security, performance, and operational expectations.
This is a hands-on infrastructure operations role that works closely with AI engineering, application development, data, architecture, cybersecurity, infrastructure operations, and external managed-services partners. The role is not expected to design AI models or own AI engineering outcomes but will provide the infrastructure and operational support required for AI initiatives.
Work Arrangement:
Employees who live within 30 miles of the TMG home office are expected to follow a hybrid or in-office schedule. The initial training period may require additional in‑office days. Accountabilities:
Enable AI Infrastructure and Environments
Partner with AI engineering, application, data, security, architecture, and platform teams to understand the infrastructure and operational needs of new AI use cases. Translate AI solution designs into requirements for environments, compute, storage, networking, connectivity, identity, security, observability, capacity, and supporting cloud services. In Collaboration with other Infrastructure resources and managed service provider, coordinate infrastructure provisioning, configuration, access, and changes with TMG’s AWS managed-services provider and other technology partners. Support operational readiness for AI solutions transitioning into production. Identify infrastructure dependencies, constraints, risks, costs, and lead times early in the delivery lifecycle. Establish reusable infrastructure patterns, operational standards, dashboards, runbooks, and production-readiness requirements across AI use cases. Validate that environments and supporting services are appropriately configured, monitored, secured, scalable, and ready for production use. Implement new cloud functionality requirements in collaboration with architecture and AI engineering within approved architecture and guardrails. Follow and enforce established security, compliance, and operational controls across AI platforms and infrastructure. Operate and Improve Infrastructure Support for AI Solutions
Provide and support the infrastructure monitoring and telemetry platforms used by AI Ops and AI Engineering teams. Define and track infrastructure and service-health metrics covering reliability, availability, latency, cost, utilization, capacity, and operational performance. Work with AI engineering and application teams to diagnose infrastructure, connectivity, integration, and platform issues affecting AI solutions. Support root-cause analysis of infrastructure, platform, connectivity, and operational issues and coordinate resolution with AI engineers, application teams, platform teams, and managed-services providers. Support AI-related incident management, problem management, operational reviews, infrastructure readiness reviews, and production-readiness activities. Create and maintain service-health dashboards, runbooks, troubleshooting guidance, support procedures, and escalation paths. Support AI Risk and Operational Governance
Partner with AI and IT governance team to operationalize applicable controls and monitoring requirements. Support AI Governance and AI engineering teams by providing operational telemetry, infrastructure evidence, monitoring data, and production support information needed for control monitoring and review. Support the collection and retention of operational evidence, including infrastructure changes, access records, monitoring results, incidents, exceptions, and corrective actions. Escalate material operational, infrastructure, security, or reliability concerns to the appropriate AI, application, security, or governance owners for evaluation and resolution. Analyze trends and proactively identify capacity constraints, reliability risks, infrastructure degradation, support gaps, and operational issues before they become production incidents. Key Outcomes
AI teams receive timely and consistent infrastructure and operational support. AI use cases progress efficiently from experimentation to reliable production operation. Reusable infrastructure and operational patterns are applied across AI initiatives. Infrastructure and supporting services for AI solutions are observable, monitored, and operationally supported. Reliability, availability, cost, usage, capacity, and operational performance are consistently measured. Production issues are detected, diagnosed, and resolved more quickly. Releases result in fewer regressions and operational disruptions. AI-supporting infrastructure services are secure, compliant, operationally ready, and aligned with enterprise governance and risk requirements. Qualifications
Bachelor’s degree in computer science, engineering, information technology, or a related field, or equivalent practical experience. 8+ years of overall information technology experience, including 5+ years in infrastructure operations, cloud operations, site reliability engineering, DevOps, platform operations, application operations, or production engineering. Experience supporting infrastructure for cloud-hosted, distributed, data-intensive, AI-enabled, or business-critical production applications. Experience implementing and supporting observability, logging, tracing, monitoring, dashboards, alerting, and operational reporting. Working knowledge of cloud infrastructure, networking, APIs, integrations, identity and access management, security, and data pipelines. Experience coordinating infrastructure services, changes, dependencies, and issue resolution across internal teams, technology partners, and managed-services providers. Familiarity with AWS services and capabilities related to monitoring, logging, networking, security, identity, and infrastructure operations. Experience with automation, CI/CD pipelines, infrastructure-as-code, configuration management, and release-management practices. Experience supporting infrastructure for AI, machine-learning, data-intensive, or cloud-native applications is preferred. Familiarity with AI observability and operational characteristics of LLM, RAG, and agent-based solutions is helpful but not required. Familiarity with vector databases, model APIs, AI gateways, prompt-management platforms, or agent orchestration frameworks is preferred but not required. Strong analytical, documentation, communication, and cross-functional collaboration skills, with the ability to manage multiple priorities across concurrent technology initiatives. Experience in insurance, financial services, or another regulated industry, along with relevant AWS, cloud, infrastructure, DevOps, SRE, security, or AI certifications, is preferred. Pay Range:
Anticipated Hiring Range:
$130,000 - $150,000 annual base salary depending on experience, qualifications, and geographic location Benefits:
We are proud to offer our full-time regular employees a robust benefits suite that includes:
Competitive base salary plus incentive plans for eligible team members 401(K) retirement plan that includes a company match of up to 6% of your eligible salary Free basic life and AD&D, long-term disability and short-term disability insurance Medical, dental and vision plans to meet your unique healthcare needs Wellness incentives Generous time off program that includes personal, holiday and volunteer paid time off Flexible work schedules and hybrid/remote options for eligible positions Educational assistance
Bot-Filtering Mechanism
The Mutual Group uses a bot-filtering mechanism as part of its recruiting and hiring process. As a result, automated job applications or submissions may be rejected.
Equal Opportunity Employer
The Mutual Group is an Equal Opportunity Employer. It is our policy to recruit, hire, train and promote individuals in all job classifications without regard to race, color, religion, sex, national origin, age, veteran status, disability, sexual orientation, gender identity or any other characteristic protected by law.
Know Your Rights: Workplace Discrimination is Illegal Your Rights Under USERRA
Applicants requiring a reasonable accommodation due to a disability at any stage of the employment application process should contact Talent@themutualgroup.com.
Employment Verification
The Mutual Group participates in the E-Verify program and will provide the federal government with your Form I-9 information to confirm that you are authorized to work in the U.S. You are protected from employment discrimination based on your citizenship status and national origin.
E-Verify Program Overview
E-Verify Participation Poster
All offers of employment are contingent upon the successful completion of a background check.
#TMG