🤫husshhussh
🤫husshhusshOnePuppy
AI infrastructure · H29

Machine Learning Training Infrastructure Engineer

Let users adapt models and run appropriately sized training jobs on resources they control. You will make training reproducible, interruptible and accountable.

See what is open todayAll roles

Not open yet: Pilot expansion

This work starts when that stage arrives, so there is no application to submit today and we will not pretend otherwise. What is written below is what the role is for and what would make somebody right for it, published early on purpose so you can decide whether it is worth watching.

The work

What this person actually does

Build training pipelines, checkpoints, optimizer-state handling, resource scheduling and distributed execution where the network supports it. Coordinate local inference with background training. Enforce dataset permissions and explicit limits before a job moves to remote compute.

The milestone

What it looks like when it is working

In your first 90 days, deliver a recoverable local fine-tuning workflow and a documented boundary between supported local jobs and remote workloads.

Evidence

What would show us you can do it

Bring practical ML training systems experience and distributed-computing fundamentals. Understand memory accounting, numerical stability and the difference between fine-tuning and large-scale pretraining.

Evidence, not credentials. We are describing work you can point at, in whatever form it exists.

The exercise

How we would look at it together

Plan a training run that must survive a power interruption while preserving the ability to compare results with the original baseline.

← All 72 roles in the catalog