Staff/Sr. ML Infrastructure / Platform Engineer

  • Taipei
  • Full-time
  • POSTED 2 DAYS AGO

About the job

Join Trend ‧ Join New Generation

趨勢科技 - 全球雲端資安領航者 / 全亞洲最大軟體公司 / 企業版圖橫跨五大洲 / 趨勢全球研發基地在台灣 ===============================================================

About the Role

We are building a production-grade, GPU-accelerated LLM serving platform that powers multiple AI products at enterprise scale. You will be responsible for designing, building, and operating the infrastructure that serves large language models - from raw Kubernetes cluster management to multi-GPU inference optimization and autoscaling.

Required Qualifications

Model Serving & Inference

Operate multi-model LLM serving infrastructure Tune autoscaling policies to balance GPU cost and latency SLAs Kubernetes & GPU Infrastructure

Operate production K8s clusters with NVIDIA GPU nodes Handle GPU node lifecycle: NVIDIA driver setup Infrastructure as Code

Write and maintain Terraform/Terragrunt modules for AWS/GCP cloud Package platform components and model deployments as Helm charts Manage multi-environment configurations Observability & Performance

Maintain monitoring stack: Prometheus, Grafana, Build dashboards for GPU utilization, KV cache occupancy, TTFT/ITL latency, and cost per token Set up alerting for SLA violations and OOM events

Bonus Skills

These are not required, but candidates with these skills will stand out.

LoRA / PEFT fine-tuning workflows MLflow for experiment tracking, model registry, and automated adapter deployment Experience building LoRA adapter CI/CD pipelines (training → registry → serving) Experience with alternative inference frameworks such as SGLang or NVIDIA NIM, including deep Parameter Tuning for Continuous Batching, KV Cache management, and Speculative Decoding. =============================================================== 連結智慧 守護世界 --- Connected Intelligence for Securing a Connected World