Staff AI Product Engineer
Nscale
- New York, New York, United States
- $220,000 - $293,333 a year
- Posted Aug 12, 2026
Job description
About the role
The Staff AI Product Engineer sets technical direction for Nscale’s AI services platform across 2‑4 teams, defining patterns, APIs and systems for AI capabilities, API gateway, developer experience, billing, identity and extensibility, driving scalability, developer experience and product velocity.
About the company
Nscale builds a vertically integrated GenAI cloud platform, owning data centres, software and applications that power the AI stack with sustainable technology; its culture emphasizes relentless innovation, ownership, accountability, openness, collaboration, adaptability and resilience.
Requirements
- 8–12 years of software engineering experience
- Proven ability to set technical direction for a product domain, including cross-team architectural patterns
- Deep expertise in API design, platform engineering, and large-scale distributed systems
- Experience designing cloud services with clear control plane / data plane separation and cell-based architecture for horizontal scalability and blast-radius isolation
- Experience building and operating developer-facing platforms used by large numbers of engineers or customers
- Proven ability to define the provisioning contract other teams onboard onto: typed inputs, readiness semantics, published outputs, and declared dependencies between provisioned services
- Experience with dependency-ordered composition across services owned by different teams — readiness gating, eventual consistency, and deciding what may be provisioned in parallel
- Track record of setting versioning and compatibility policy for customer-facing configuration surfaces, and of sequencing change across the schema, controller, packaging, and deployment layers that must land in order
- Sustained hands‑on production ownership at scale — has carried on‑call for systems they designed and fed that operational experience back into the architecture
- Track record of raising stability across a domain: SLO and error‑budget policy, incident review that produces systemic fixes, and reliability tracked as a measurable trend rather than per‑incident firefighting
- Experience making cost a first‑class engineering signal: usage attribution, cost‑per‑unit visibility, and guardrails that keep spend predictable as the platform scales
- Strong ability to resolve ambiguous, open‑ended technical problems at system scope
- Demonstrated ability to influence without formal authority — across teams, disciplines, and seniority levels
- Track record of creating durable technical standards and practices adopted across an organization
- Experience designing SDK and Terraform provider strategies that enable customers to extend and automate the platform
- Track record of building extensible platform layers: plugin systems, API versioning strategies, client library design
- Experience building AI/ML product platforms: inference APIs, fine-tuning UX, model management, evaluation tooling
- Experience building self-service, paved-path onboarding so teams can ship new deployable units without platform-team involvement
- Background in GPU cloud or compute platforms serving ML workloads; experience operating platforms with strict SLAs at scale