Make useful intelligence run responsively on a computer the user owns. You will turn models into a dependable local service.
Not open yet: First team
This work starts when that stage arrives, so there is no application to submit today and we will not pretend otherwise. What is written below is what the role is for and what would make somebody right for it, published early on purpose so you can decide whether it is worth watching.
The work
Build model serving, memory management, quantization evaluation, batching and scheduling. Support the chosen accelerator backends and measure startup, time to first token, sustained throughput and quality. Keep interactive tasks responsive while background work uses spare capacity.
The milestone
In your first 90 days, ship a local inference service with a reproducible performance and quality report for the reference device.
Evidence
Bring strong performance engineering and practical experience serving machine-learning models. Understand memory capacity, bandwidth, context length and why theoretical compute figures do not predict user experience.
Evidence, not credentials. We are describing work you can point at, in whatever form it exists.
The exercise
Diagnose an inference slowdown as context grows and propose an improvement without quietly reducing output quality.