Make each unit of compute do more useful work. You will optimize the kernels and compilation paths that matter to the user's real workloads.
Not open yet: Pilot expansion
This work starts when that stage arrives, so there is no application to submit today and we will not pretend otherwise. What is written below is what the role is for and what would make somebody right for it, published early on purpose so you can decide whether it is worth watching.
The work
Profile operators, memory traffic and scheduling; implement and validate kernels on selected backends. Work with numerical and model researchers on precision choices. Maintain portable abstractions where useful while documenting backend-specific behavior and limits.
The milestone
In your first 90 days, deliver one measured end-to-end improvement with correctness tests, hardware details and a reproducible comparison.
Evidence
Bring GPU programming, compiler or numerical computing expertise. Relevant experience may include CUDA, Metal, LLVM, MLIR, Triton or other accelerator toolchains; mastery of every backend is unnecessary.
Evidence, not credentials. We are describing work you can point at, in whatever form it exists.
The exercise
Show how a faster isolated kernel could still make the complete application slower, and design the correct benchmark.