AI Compute Architect – Full-Stack GPU Software
We are partnering with an ambitious semiconductor company developing a new GPU architecture and complete AI software stack.
They are seeking a senior, hands-on AI Compute Architect to define how complete machine-learning workloads move from frameworks and computational graphs through compilers, runtimes and kernels onto the underlying SIMT hardware.
You’ll work alongside accomplished GPU, compiler and systems architects, providing technical direction while building consensus across a highly experienced engineering team.
What you’ll do
- Define the end-to-end architecture of a new GPU AI software stack
- Connect ML frameworks and complete model workloads to compiler, runtime and hardware requirements
- Establish interfaces between the compiler, runtime, driver, kernels and SIMT architecture
- Shape support for modern workloads across LLM inference, training and other AI applications
- Influence ISA, memory hierarchy, scheduling and execution architecture through software requirements
- Evaluate design trade-offs across performance, programmability, portability and developer experience
- Provide technical leadership across compiler, runtime, framework and architecture teams
What we’re looking for
- Extensive experience architecting software for GPUs, AI accelerators or high-performance computing platforms
- Technical ownership spanning several layers of an AI compute stack
- Deep knowledge of LLVM, MLIR, XLA, IREE or comparable compiler infrastructure
- Experience with CUDA, SIMT programming or equivalent parallel-compute architectures
- Familiarity with PyTorch, JAX, ONNX, Triton or other modern ML frameworks and programming models
- Understanding of GPU runtimes, drivers, kernel execution, memory movement and asynchronous scheduling
- Experience translating complete AI workloads into software and hardware architectural requirements
- Principal, Architect, Fellow or Distinguished-level technical leadership capability
