Machine Learning Engineering Manager, Evaluation, Agentic Search Capabilities, Proactive

Apple·Santa Clara, California, United States·posted 2d ago · last seen 8m ago

Track this application

Get Started Free

Match score against your CV

Get Started Free

Tailor your resume to this job

Get Started Free

About interviewing at Apple

One of tech's least standardized loops: you interview for a specific team, and every stage belongs to it. Recruiter/hiring-manager screen, one to three 45–60 minute Coderpad coding screens, a system design round shaped by Apple's reliability and privacy constraints, and a behavioral round — then a panel debrief. There is no central question bank; questions map to the team's real stack.

Read the full Apple interview process →

Description

Summary

The Siri team is redefining how hundreds of millions of people access information across Apple devices, with privacy built in from the ground up. As part of the Agentic Search Evaluation team, you will help advance Apple Intelligence through personalized search, result ranking, and low-latency production services that power experiences across Siri, Spotlight, Safari, Messages, and more. Your team researches and builds deep search systems for Personal Question Answering, enabling Siri to answer questions about a user's emails, messages, events, files, and more, while keeping personal data private.

Description

We are looking for a Machine Learning Engineering Manager with a passion for building scalable, high-quality solutions to evaluate personalized search applications. In this role, you will directly shape how Apple assesses and improves the quality of Agentic Search — driving solutions across search quality, data quality, and both offline and online metrics to enable iterative product improvements. You will lead a team focused on evaluating Agentic Search experiences that power personalized Siri AI and enable answers to users' personal questions, with privacy at the forefront. Your team designs and builds evaluation systems, develops data generation and curation methods, and creates metrics frameworks for quality assessment and continuous improvement.

Key Responsibilities

Lead a team of applied machine learning and software engineers to design and build scalable, automated systems for end-to-end evaluation of personalized Agentic Search Contribute to the design and implementation of offline and online metrics to assess quality at the product and component level for Personal Question Answering, and drive iterative improvement Collaborate with partner teams to define data and evaluation requirements, set priorities, and identify opportunities to enhance the Personal Question Answering stack Define the long-term technical vision for personalized Siri Search quality, identify gaps, and drive solutions as part of a broader roadmap

Minimum Qualifications

6 or more years of industry experience building machine learning or machine learning evaluation systems at scale 2 or more years of industry experience in technical leadership or management Knowledge of machine learning fundamentals, neural network architectures, and software engineering in Python or C/C++ Communication and collaboration skills Bachelor's degree in Computer Science, Engineering, Statistics, or a related field

Preferred Qualifications

Experience designing and building production machine learning systems in search, natural language processing, recommendation systems, or information retrieval Experience in data collection, data generation, or data quality assessment of language, image, or multimodal data Advanced degree (Master's or Ph.D.) in Computer Science, Engineering, Statistics, or a related field, or equivalent industry experience

More engineering roles at Apple