🤫husshhussh
🤫husshhusshOnePuppy
AI research · H38

AI Evaluation and Reliability Scientist

Show when a private agent deserves a user's trust. You will build evidence about task completion, failure and human oversight.

See what is open todayAll roles

Not open yet: Pilot expansion

This work starts when that stage arrives, so there is no application to submit today and we will not pretend otherwise. What is written below is what the role is for and what would make somebody right for it, published early on purpose so you can decide whether it is worth watching.

The work

What this person actually does

Design evaluations for real workflows, adversarial inputs and distribution shifts. Use held-out tasks, blinded review where appropriate and calibrated human evaluation. Measure false completion, unauthorized actions, appropriate escalation and sustained usefulness over time.

The milestone

What it looks like when it is working

In your first 90 days, deliver a versioned evaluation suite and a release report that makes strengths and unresolved failures easy to inspect.

Evidence

What would show us you can do it

Bring experimental design, statistics and hands-on AI evaluation. Show how you prevent benchmark contamination and avoid substituting a model judge for ground truth.

Evidence, not credentials. We are describing work you can point at, in whatever form it exists.

The exercise

How we would look at it together

Design an evaluation that detects an agent becoming more persuasive without becoming more accurate or reliable.

← All 72 roles in the catalog