🤫husshhussh
🤫husshhusshOnePuppy
AI research · H36

Speech and Multimodal Intelligence Scientist

Make personal computing accessible through the ways people naturally communicate. You will study speech, vision and other inputs under real device constraints.

See what is open todayAll roles

Not open yet: Research program

This work starts when that stage arrives, so there is no application to submit today and we will not pretend otherwise. What is written below is what the role is for and what would make somebody right for it, published early on purpose so you can decide whether it is worth watching.

The work

What this person actually does

Develop or adapt models for speech, documents, images and grounded interaction. Evaluate diverse accents, environments and accessibility needs with appropriate participant consent. Measure local performance and make capture, retention and activation behavior understandable.

The milestone

What it looks like when it is working

In your first 90 days, deliver a bounded multimodal capability with evaluation across realistic conditions and a clear record of its limitations.

Evidence

What would show us you can do it

Bring research or advanced engineering in speech, vision or multimodal learning. Show how you distinguish apparent fluency from correct understanding.

Evidence, not credentials. We are describing work you can point at, in whatever form it exists.

The exercise

How we would look at it together

Design an evaluation for a voice-controlled task in a shared room where only one person has authorized the action.

← All 72 roles in the catalog