Inference Software Engineer

Etched

  • San Jose, California, United States
  • Onsite
  • Posted Jun 17, 2025
Sign up — let your agent apply Sign in

BestApply tailors your resume and applies for you.

PyTorchC++RustLinuxNVLinkTPUGPUInfiniBandCompilersJAX

Job description

About the role

The role focuses on enabling state‑of‑the‑art AI models on Etched’s custom inference platform by porting models, expanding the multi‑node runtime, optimizing communication and routing, and using profiling tools to find performance and correctness issues.

About the company

Etched builds hardware and integrated systems for frontier AI inference, co‑designing chips, racks, software and manufacturing to deliver high‑throughput, low‑latency performance. Backed by hundreds of millions in funding and staffed by leading engineers, the San‑Jose company is redefining the infrastructure layer for the rapidly growing inference market.

Requirements

  • Proficiency in C++ or Rust.
  • Understanding of performance‑sensitive or complex distributed software systems such as Linux internals, accelerator architectures (GPUs, TPUs), compilers, or high‑speed interconnects (NVLink, InfiniBand).
  • Familiarity with PyTorch or JAX.
  • Experience porting applications to non‑standard accelerator hardware or platforms.
  • Experience developing low‑latency, high‑performance applications using both kernel‑level and user‑space networking stacks (nice‑to‑have).
  • Deep knowledge of distributed systems concepts, algorithms, and challenges including consensus protocols, consistency models, and communication patterns (nice‑to‑have).
  • Solid grasp of Transformer architectures, especially Mixture‑of‑Experts models (nice‑to‑have).
  • Experience building applications with extensive SIMD optimizations for performance‑critical paths (nice‑to‑have).