AIML - Sr Software Engineer - AI, Evaluation

Apple·Cupertino, California, United States·posted 3d ago · last seen 3m ago

Track this application

Get Started Free

Match score against your CV

Get Started Free

Tailor your resume to this job

Get Started Free

About interviewing at Apple

One of tech's least standardized loops: you interview for a specific team, and every stage belongs to it. Recruiter/hiring-manager screen, one to three 45–60 minute Coderpad coding screens, a system design round shaped by Apple's reliability and privacy constraints, and a behavioral round — then a panel debrief. There is no central question bank; questions map to the team's real stack.

Read the full Apple interview process →

Description

Summary

Do you want to help build the next generation of Apple AI products? Our Evaluation organization measures quality across Apple Intelligence, Siri, and newest generative features, and our measurements directly inform launch decisions. Our team builds the platforms behind it: LLM-as-judge autograders, tools to validate them against human judgment, and systems that run them at scale. We are looking for a senior engineer to own a critical part of that platform and take it to the next level.

Description

This is an ownership role. You will take on major systems in our evaluation platform, existing or new, and be accountable for its direction, quality, and adoption. This role sits at the intersection of AI modeling, software engineering, and product quality. We are expanding rapidly, so the ability to build a strong foundation and structure to any system is paramount. We are supporting features that encompass both product development and research, which requires a balance of respect for both the scientific process as well as the product release lifecycle. Our systems connect many platforms and serve teams across Apple, so judgment with people matters as much as technical depth.

Key Responsibilities

Own a core evaluation system end-to-end for multiple features: architecture, roadmap, reliability, and developer experience. Design testing and validation strategies that catch real failures before partners do. Build and maintain integrations with annotation, compute, reporting, and benchmarking platforms. Partner with teams building LLM judges on our platform, turning their recurring needs into product. Set the bar for using AI coding agents without compromising quality. Customer-oriented: treats internal users' needs as the job, never loses sight of Apple customers and product quality, and communicates clearly to align cross-functional partners.

Minimum Qualifications

BS/MS/PhD in Computer Science, Machine Learning, or a related field. 8+ years of software engineering experience, including owning production systems used by other teams. Exceptional Python skills and a strong engineering understanding of system design, API design, system testing and validation, debugging and monitoring. Experienced in shipping production code with AI coding agents while keeping quality high.

Preferred Qualifications

Proficient in building agentic systems, LLM-based applications, LLM-as-judge systems, and offline evaluations. Ability to interpret evaluation results, including human agreement, run-to-run consistency, and respecting constraints like data retention and privacy policies. Expert in building and maintaining technical integrations across multiple platforms in a large organization, with varying technical requirements and needs based off scaling, volume, throughput, etc. Product-minded, turning ambiguous requirements into a focused plan.

More engineering roles at Apple