Member of Technical Staff (Answer Quality & Evals)
Track this application
Get Started FreeMatch score against your CV
Get Started FreeTailor your resume to this job
Get Started FreeAbout interviewing at Perplexity
One of the few AI startups with a fully published interview guide (perplexity.ai/hub/careers/interview-guide): online application (response within two weeks) → recruiter phone screen → a technical screen that for engineers is 'usually a standard technical programming interview' → a quickly-scheduled onsite of 4–5 interviews including a hiring-manager deep dive on past work and experience anecdotes → a final interview with a Perplexity founder or leader → decision within a week of the onsite. Coding leans Python and mixes LeetCode medium–hard with practical search-flavored tasks (ranking/filtering, concurrency, data handling); system design is AI-native (RAG pipelines, retrieval at scale, LLM serving cost/latency). Applicants are judged 'solely on merit and potential impact' and must show 'frontier knowledge and excellence in at least one area'; roles are broad by default with team matching happening during the onsite, every role — managers included — is hands-on, and building AI products isn't expected but fluency in using AI tools is required. In-person 4 days/week near an office; remote is case-by-case.
Read the full Perplexity interview process →Description
Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and specialized data sources. The Answer Quality team ensures that our prompts, tools, search systems, datasets, and models work together to create the best possible experience for our users.
As our product and agent capabilities evolve, we need evaluation systems that are fast, reliable, production-faithful, and actionable. In this role, you will build and improve the technical foundations that support Answer Quality across Perplexity. This includes our shared evaluation infrastructure and the platform used to replay and analyze agent traces. You will work closely with data scientists, engineers, and product teams to identify quality problems, measure their impact, and turn evaluation findings into product improvements.
Responsibilities
Build shared evaluation infrastructure that helps teams run reliable evals, analyze results, and make product and model decisions
Develop the platform for replaying and analyzing agent traces to reproduce production behavior and diagnose failures
Build and operate scalable systems for processing, storing, and monitoring interaction, trace, and evaluation data
Partner with data scientists, engineers, and product teams to turn answer-quality problems into evaluations, analyses, and product improvements
Operate in a small, high-impact team where your work directly shapes how Perplexity measures and improves Answer Quality
Qualifications
4+ years of software, data, or machine learning engineering experience shipping and operating production systems
Strong proficiency in Python and SQL, with solid fundamentals in system design, data modeling, and distributed systems
Experience building big-data systems, including distributed compute, large-scale storage, and high-volume pipelines
Demonstrated ownership of ambiguous technical projects from initial design through production operation
Ability to work effectively with data scientists, engineers, and product partners
Preferred Qualifications
Experience building evaluation, experimentation, observability, or machine learning infrastructure
Familiarity with LLM and agent systems, including tool use, execution traces, replay, and simulation
Experience building on top of large-scale data processing platforms such as Databricks, Snowflake, or ClickHouse
Salary Range: $200K - $350K