Principal Infrastructure Engineer, AI Cluster Performance & Validation
Overview As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering and testing principles, you will focus on building and maintaining the control plane, tooling, and automation that supports performance and...
- Python
- PyTorch
- Linux
- Ansible
- Terraform
- Kubernetes
- Prometheus
- Grafana
- Docker
- C++