Show when a private agent deserves a user's trust. You will build evidence about task completion, failure and human oversight.
Not open yet: Pilot expansion
This work starts when that stage arrives, so there is no application to submit today and we will not pretend otherwise. What is written below is what the role is for and what would make somebody right for it, published early on purpose so you can decide whether it is worth watching.
The work
Design evaluations for real workflows, adversarial inputs and distribution shifts. Use held-out tasks, blinded review where appropriate and calibrated human evaluation. Measure false completion, unauthorized actions, appropriate escalation and sustained usefulness over time.
The milestone
In your first 90 days, deliver a versioned evaluation suite and a release report that makes strengths and unresolved failures easy to inspect.
Evidence
Bring experimental design, statistics and hands-on AI evaluation. Show how you prevent benchmark contamination and avoid substituting a model judge for ground truth.
Evidence, not credentials. We are describing work you can point at, in whatever form it exists.
The exercise
Design an evaluation that detects an agent becoming more persuasive without becoming more accurate or reliable.