Software Engineer, RL Data

Cursor·San Francisco, California, United States·posted 19d ago · last seen 18m ago

Track this application

Get Started Free

Match score against your CV

Get Started Free

Tailor your resume to this job

Get Started Free

About interviewing at Cursor

Short and fast: a recruiter screen, then '2-3 short technicals' (Cursor's own job-posting wording), then the signature round — an onsite in their office where you build a real project. CEO Michael Truell has said every engineering and design hire spends two days in-office working on a project: you get a desk, a laptop, a choice of three projects, and a frozen older copy of the real codebase with the dev environment set up, then present what you built. Crucially, first technical screens BAN AI beyond autocomplete ('programming without AI is still a really great time-boxed test for skill and intelligence' — Truell), while the onsite expects you to work with full tooling including Cursor itself. There is no formal behavioral round — the 4-6 meals with the team during the trial are the culture interview. Decisions come fast; the whole loop targets weeks, not months.

Read the full Cursor interview process →

Description

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

Software Engineer, Reinforcement learning 
  • SpaceXAI is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective.

 

About the role

 
  • As a Software Engineer on the RL Data team at SpaceXAI, you'll create the tasks, rewards, and environments that train our coding agents. The team owns the data that goes into training: what the model is asked to do, how we score it, and the setups it learns in.

  

What you’ll do

 
  • Designing a task set that teaches a specific agent capability, then iterating on it from traces and evals until the model actually gets better.

  • Reading a pile of agent traces, finding a failure mode or a surprising behavior, and building a system that surfaces more of the same.

  • Turning a one-off recipe into something other teams can reuse: better rewards, cleaner environments, tighter data quality.

  • Partnering with research on whether a dataset is actually teaching the thing we think it is.

    

You may be a fit if

  • You write careful, fast code and have strong software engineering fundamentals.

  • You like setting tasks: breaking a fuzzy capability into something concrete you can measure.

  • You have an infra, data, or distributed systems background. RL experience is a plus, not a requirement.

  • You enjoy looking at messy real-world agent behavior and turning it into a dataset or a tool.

More engineering roles at Cursor