Infrastructure Engineer (GPU & Compute)

Lightning AI

  • United States, United States
  • Remote
  • $180,000 - $200,000 a year
  • Posted Jul 23, 2026
Sign up — let your agent apply Sign in

BestApply tailors your resume and applies for you.

Bare Metal ProvisioningLinuxPythonGPUNVIDIA DCGMHardware Debugging

Job description

About the role

The senior GPU & Compute Infrastructure Engineer will own and evolve image management, deployment and validation systems across large‑scale bare‑metal GPU‑enabled infrastructure, ensuring reliability and performance for AI/ML workloads.

About the company

Lightning AI is the company behind PyTorch Lightning, building an end‑to‑end platform for developing, training and deploying AI systems.

Requirements

  • 5+ years of experience in infrastructure engineering, systems engineering, or related roles
  • Strong Linux systems experience in production environments
  • Hands‑on experience with GPU‑enabled systems and tools such as NVIDIA DCGM
  • Familiarity with bare‑metal provisioning and system bring‑up workflows
  • Proficiency in Python or similar scripting/programming languages for automation
  • Ability to debug complex issues across hardware, OS, GPUs, and system software
  • Experience with high‑performance interconnects (e.g., InfiniBand, NVLink)
  • Experience with PXE boot environments, LiveCD systems, or image‑based provisioning workflows
  • Experience with hardware management interfaces such as iDRAC, IPMI, or Redfish
  • Data center operations experience, including working with physical hardware
  • Experience supporting AI/ML or HPC workloads at scale
  • Experience with GPU validation frameworks or large‑scale hardware qualification processes