🤫husshhussh
🤫husshhusshOnePuppy
AI infrastructure · H27

AI Inference Systems Engineer

Make useful intelligence run responsively on a computer the user owns. You will turn models into a dependable local service.

See what is open todayAll roles

Not open yet: First team

This work starts when that stage arrives, so there is no application to submit today and we will not pretend otherwise. What is written below is what the role is for and what would make somebody right for it, published early on purpose so you can decide whether it is worth watching.

The work

What this person actually does

Build model serving, memory management, quantization evaluation, batching and scheduling. Support the chosen accelerator backends and measure startup, time to first token, sustained throughput and quality. Keep interactive tasks responsive while background work uses spare capacity.

The milestone

What it looks like when it is working

In your first 90 days, ship a local inference service with a reproducible performance and quality report for the reference device.

Evidence

What would show us you can do it

Bring strong performance engineering and practical experience serving machine-learning models. Understand memory capacity, bandwidth, context length and why theoretical compute figures do not predict user experience.

Evidence, not credentials. We are describing work you can point at, in whatever form it exists.

The exercise

How we would look at it together

Diagnose an inference slowdown as context grows and propose an improvement without quietly reducing output quality.

← All 72 roles in the catalog