Machine Learning Engineer, Evaluation, Agentic Search Capabilities

Apple·Santa Clara, California, United States·posted 12m ago · last seen 7m ago

Track this application

Get Started Free

Match score against your CV

Get Started Free

Tailor your resume to this job

Get Started Free

About interviewing at Apple

One of tech's least standardized loops: you interview for a specific team, and every stage belongs to it. Recruiter/hiring-manager screen, one to three 45–60 minute Coderpad coding screens, a system design round shaped by Apple's reliability and privacy constraints, and a behavioral round — then a panel debrief. There is no central question bank; questions map to the team's real stack.

Read the full Apple interview process →

Description

Summary

The Agentic Search team is redefining how hundreds of millions of people access information across Apple devices, with privacy built in from the ground up. As part of the Agentic Search Evaluation team, you will help advance Apple Intelligence through personalized search that powers experiences across Siri, Spotlight, Mail, Messages, and more. Your team researches and builds deep search systems for Personal Question Answering, enabling Siri to answer questions about a user's emails, messages, events, files, and more, while keeping personal data private.

Description

We are looking for a senior Machine Learning Engineer with a passion for building scalable, high quality solutions to evaluate personalized agentic search applications. In this role, you will directly impact how we assess and improve the quality of novel agentic search applications by driving solutions for data generation, data quality assessment, offline and online metrics, and methods for statistical model and search quality assessment. Your work will impact the quality of innovative Personalized Agentic Search experiences across Siri and Apple products. In this role you will have the opportunity to build scalable systems for comprehensive evaluation of personalized model and search systems, develop data generation and curation methodologies to support evaluations, and design online and offline metrics for product quality assessment. Our team is responsible for models and evaluation of search experiences that answer user’s questions using their personal documents with privacy at the forefront.

Key Responsibilities

Build scalable, automated systems for large scale, end-to-end evaluation of models and search powered systems. Design and implement novel data generation and data quality assessment methods to support modeling and evaluation of personal agentic search. Design and implement offline and online metrics to assess product and component level quality for personal agentic search. Collaborate with partner teams to define data and evaluation requirements and priorities, and to explore opportunities for enhancements to the personal search stack. Develop long-term technical vision for personal search quality; identify problem areas and drive solutions as part of a larger roadmap.

Minimum Qualifications

6+ years industry experience in building Machine Learning or ML evaluation systems at scale. Strong software engineering skills in mainstream programming languages, such as: Python, C/C++. Ability to quickly prototype ideas and solutions, and perform critical analysis. Strong communication skills and ability to drive solutions in collaboration with partner teams. Bachelors in Computer Science, Engineering, Statistics or related field.

Preferred Qualifications

Experience in design and building production ML systems and applications in search, NLP, recommendation systems, or information retrieval. Experience in data collection, data generation and/or data quality assessment of language, image or multi-modal data. Strong skills for quality metrics development; interpretation of evaluations; and presentation to executive audience. Advanced degree (Master’s or Ph.D.) in Computer Science, Engineering, Statistics, or related field, or equivalent industry work experience.

More engineering roles at Apple