Principal Inference Performance Architect
Back to Job Search

Principal Inference Performance Architect

Reference: SE19
Location
San Francisco Bay Area, CA, USA
Salary
Up to $400k + equity
Contract Type
Permanent
Work Arrangement
In-Office (Full-Time)
Skill Requirements
  • Software Engineering
  • Embedded, Electronics & Semiconductors

Principal / Distinguished Inference Performance Architect

I'm working with a well-funded AI compute company that is building a next-generation accelerator platform designed specifically for large-scale generative AI and foundation model workloads.

As they scale the platform, they are looking to hire a Principal or Distinguished-level Inference Performance Architect to define, build, and optimise the software stack that ultimately determines how efficiently large language models run on the hardware. This is a highly influential individual contributor role, sitting at the intersection of AI systems, compiler technology, runtime architecture, and accelerator performance.

You will be responsible for driving end-to-end inference efficiency, identifying and removing bottlenecks across the serving stack, and shaping the architectural decisions that maximise performance for real-world production deployments.

Key Areas of Ownership

You will work across multiple layers of the inference ecosystem, including:

  • Transformer graph optimisation and execution planning
  • Quantisation strategies, speculative decoding, and advanced inference techniques
  • Compiler, runtime, and graph-lowering performance improvements
  • Development and optimisation of high-performance GEMM, attention, and MoE kernels
  • KV-cache design, memory hierarchy optimisation, and bandwidth efficiency
  • Dynamic batching, scheduling, request routing, and serving infrastructure
  • Tensor, pipeline, data, and expert parallelism strategies
  • Multi-accelerator and multi-node inference scaling
  • End-to-end profiling, performance analysis, and bottleneck identification
  • Close collaboration with hardware, compiler, systems, and model teams to influence platform direction

What Success Looks Like

This role is measured by real production outcomes rather than theoretical benchmarks. You will be expected to deliver meaningful improvements across metrics such as:

  • Tokens per second
  • Time to first token (TTFT)
  • Inter-token latency
  • Throughput under high concurrency
  • Accelerator utilisation
  • Memory efficiency
  • Cluster-level performance
  • Cost per token and overall serving economics

What They're Looking For

  • Principal, Distinguished, Fellow, or equivalent senior technical leadership experience
  • Deep expertise in large-scale transformer inference optimisation
  • Strong systems programming background with C++, alongside CUDA, Triton, or other accelerator programming frameworks
  • Experience operating across multiple layers of the inference stack, from kernels and runtimes through to distributed serving systems
  • A proven track record of personally delivering measurable improvements in latency, throughput, utilisation, or efficiency
  • The ability to define technical direction and influence architecture while remaining highly hands-on in implementation and performance analysis
  • Experience working on GPUs, AI accelerators, HPC platforms, or large-scale distributed AI infrastructure is highly desirable

Why This Opportunity?

This is a rare opportunity to define the core inference architecture of a new AI compute platform from the ground up. You'll work directly at the boundary between models, software, and silicon, helping shape how next-generation generative AI workloads are deployed and scaled while having a direct and measurable impact on performance at the hardware level.

Apply Now

Please fill in the form below to apply for this job.

Apply Now
Get in touch
Sebastian Eyre image
Sebastian Eyre
Similar Jobs
10th Aug 2026

Principal Software Architect

In-Office (Full-Time)Embedded, Electronics & SemiconductorsSoftware Engineering
14th Jul 2026

Lead Flight Software Engineer

In-Office (Full-Time)Software Engineering

Get in touch.

oho connects the future to your hands. Let us know what we can do for you.