Infrastructure Engineer (GPU & Compute)
Lightning AI
- United States, United States
- Remote
- $180,000 - $200,000 a year
- Posted Jul 23, 2026
Job description
About the role
The senior GPU & Compute Infrastructure Engineer will own and evolve image management, deployment and validation systems across large‑scale bare‑metal GPU‑enabled infrastructure, ensuring reliability and performance for AI/ML workloads.
About the company
Lightning AI is the company behind PyTorch Lightning, building an end‑to‑end platform for developing, training and deploying AI systems.
Requirements
- 5+ years of experience in infrastructure engineering, systems engineering, or related roles
- Strong Linux systems experience in production environments
- Hands‑on experience with GPU‑enabled systems and tools such as NVIDIA DCGM
- Familiarity with bare‑metal provisioning and system bring‑up workflows
- Proficiency in Python or similar scripting/programming languages for automation
- Ability to debug complex issues across hardware, OS, GPUs, and system software
- Experience with high‑performance interconnects (e.g., InfiniBand, NVLink)
- Experience with PXE boot environments, LiveCD systems, or image‑based provisioning workflows
- Experience with hardware management interfaces such as iDRAC, IPMI, or Redfish
- Data center operations experience, including working with physical hardware
- Experience supporting AI/ML or HPC workloads at scale
- Experience with GPU validation frameworks or large‑scale hardware qualification processes