Let users adapt models and run appropriately sized training jobs on resources they control. You will make training reproducible, interruptible and accountable.
Not open yet: Pilot expansion
This work starts when that stage arrives, so there is no application to submit today and we will not pretend otherwise. What is written below is what the role is for and what would make somebody right for it, published early on purpose so you can decide whether it is worth watching.
The work
Build training pipelines, checkpoints, optimizer-state handling, resource scheduling and distributed execution where the network supports it. Coordinate local inference with background training. Enforce dataset permissions and explicit limits before a job moves to remote compute.
The milestone
In your first 90 days, deliver a recoverable local fine-tuning workflow and a documented boundary between supported local jobs and remote workloads.
Evidence
Bring practical ML training systems experience and distributed-computing fundamentals. Understand memory accounting, numerical stability and the difference between fine-tuning and large-scale pretraining.
Evidence, not credentials. We are describing work you can point at, in whatever form it exists.
The exercise
Plan a training run that must survive a power interruption while preserving the ability to compare results with the original baseline.