Apply
Fill out this form or shoot an email to [email protected].
Compensation: $150k to $350k base depending on experience, + equity
What you'll do
As our founding AI researcher, you will:
- Turn real research work into RL environments. Jupyter sessions, HPC jobs, simulation sweeps, literature triage, and analysis pipelines become gradeable tasks: recover a known result from held-out data, reproduce a figure from a paper's raw data, find the bug that invalidated a run, propose the next experiment the authors actually ran.
- Solve verification. Most scientific tasks don't have clean ground truth, so you decide what is actually gradeable and how: held-out outcomes, replication targets, tests over analysis code, execution feedback, expert preference. Getting this right is the whole game.
- Own the trace pipeline. Tuva instruments what researchers actually do, every task, the context it required, the output, and how project knowledge shifted as a result. You turn that stream into environment seeds, training data, and evals.
- Train and evaluate. Post-training and RL fine-tuning on open weights, distillation, rejection sampling, whatever earns its keep. You own the eval harness for long-horizon agents where a single episode is hours of work.
- Ship what you learn. Policies, scaffolds, and tool-use patterns that win in the environment should show up in the product, and product usage should feed the next environment. You'll work directly with our users to source tasks and check that env performance tracks real usefulness.
- Publish!
Who you are
- You love science and scientists. Maybe you have scientists in your family, have a technical degree and did research in college, or maybe you just feel a sense of wonder when you look up at the Milky Way on a dark, moonless night.
- You've built RL environments, post-training pipelines, or agent evals that other people depended on. Research or production, but shipped.
- You have a real opinion about where RL on LLMs works and where it's cargo cult, and you can defend it with things you've run.
- You're comfortable with the unglamorous half: rollout infra, data plumbing, harnesses, reproducibility. Environments are mostly engineering.
- You're a full-stack builder with product instincts. You can prototype in a day and own a feature from architecture to UX.
Prior (and/or ongoing) research experience, academic or industry, is a significant plus. A PhD in CS/ML or a scientific discipline, or equivalent depth, is a strong advantage.
Even if all of these don't apply, please reach out if our mission and attitude resonate. Humans are not checklists.