Help an agent finish meaningful work while recognizing when it needs a person. You will study planning under uncertainty and limited authority.
Not open yet: Research program
This work starts when that stage arrives, so there is no application to submit today and we will not pretend otherwise. What is written below is what the role is for and what would make somebody right for it, published early on purpose so you can decide whether it is worth watching.
The work
Investigate task decomposition, tool selection, error recovery, uncertainty and learning from feedback. Use realistic environments with hidden failures and changing state. Keep authority outside model-generated text and work with product researchers on when asking a question improves the outcome.
The milestone
In your first 90 days, establish a task benchmark and demonstrate a planning improvement against a strong baseline with failure analysis.
Evidence
Bring research in reasoning, reinforcement learning, planning or decision systems. You should be comfortable explaining a method's assumptions and why a success metric may be misleading.
Evidence, not credentials. We are describing work you can point at, in whatever form it exists.
The exercise
Analyze an agent that appears successful because it declares completion before the external service confirms the action.