Principal / Distinguished Inference Performance Architect
I'm working with a well-funded AI compute company that is building a next-generation accelerator platform designed specifically for large-scale generative AI and foundation model workloads.
As they scale the platform, they are looking to hire a Principal or Distinguished-level Inference Performance Architect to define, build, and optimise the software stack that ultimately determines how efficiently large language models run on the hardware. This is a highly influential individual contributor role, sitting at the intersection of AI systems, compiler technology, runtime architecture, and accelerator performance.
You will be responsible for driving end-to-end inference efficiency, identifying and removing bottlenecks across the serving stack, and shaping the architectural decisions that maximise performance for real-world production deployments.
Key Areas of Ownership
You will work across multiple layers of the inference ecosystem, including:
- Transformer graph optimisation and execution planning
- Quantisation strategies, speculative decoding, and advanced inference techniques
- Compiler, runtime, and graph-lowering performance improvements
- Development and optimisation of high-performance GEMM, attention, and MoE kernels
- KV-cache design, memory hierarchy optimisation, and bandwidth efficiency
- Dynamic batching, scheduling, request routing, and serving infrastructure
- Tensor, pipeline, data, and expert parallelism strategies
- Multi-accelerator and multi-node inference scaling
- End-to-end profiling, performance analysis, and bottleneck identification
- Close collaboration with hardware, compiler, systems, and model teams to influence platform direction
What Success Looks Like
This role is measured by real production outcomes rather than theoretical benchmarks. You will be expected to deliver meaningful improvements across metrics such as:
- Tokens per second
- Time to first token (TTFT)
- Inter-token latency
- Throughput under high concurrency
- Accelerator utilisation
- Memory efficiency
- Cluster-level performance
- Cost per token and overall serving economics
What They're Looking For
- Principal, Distinguished, Fellow, or equivalent senior technical leadership experience
- Deep expertise in large-scale transformer inference optimisation
- Strong systems programming background with C++, alongside CUDA, Triton, or other accelerator programming frameworks
- Experience operating across multiple layers of the inference stack, from kernels and runtimes through to distributed serving systems
- A proven track record of personally delivering measurable improvements in latency, throughput, utilisation, or efficiency
- The ability to define technical direction and influence architecture while remaining highly hands-on in implementation and performance analysis
- Experience working on GPUs, AI accelerators, HPC platforms, or large-scale distributed AI infrastructure is highly desirable
Why This Opportunity?
This is a rare opportunity to define the core inference architecture of a new AI compute platform from the ground up. You'll work directly at the boundary between models, software, and silicon, helping shape how next-generation generative AI workloads are deployed and scaled while having a direct and measurable impact on performance at the hardware level.
