Description
Summary
At Apple, we don't just build products — we build experiences fueled by world-class data. The Responsible AI team is looking for a senior operations program manager to own the partnerships, vendor relationships, and data operations behind our Safety Evaluation and Red Teaming programs, scaling the human-labeled data, workforce, and localizationinfrastructure that generative AI features across our Software ecosystem.
Description
In this role, you will manage Apple's relationship with internal crowd as well as a portfolio of external vendorpartners spanning evaluation, red-teaming, and data generation. You will own how this work gets resourced, scoped, and executed — from statements of work and workforce calibration to quality audits and ongoing partner performance — ensuring these partnerships can scale reliably as safety evaluation needs grow across products, features, and markets.You will also own the feedback loop and establish the processes that connect Red-Teaming, Evaluation, and Post-Ship Insights groups. Findings surfaced through red-teaming and evaluation need to inform what we monitor once features ship, and post-ship signals need to flow back into how we prioritize and design future red-teaming and evaluation coverage. You will build and operate the mechanisms — shared reporting, recurring syncs, and closed-loop tracking — that keep these three functional areas working from a single, current picture of risk rather than in isolation. Ensuring compliant and timely reporting across Apple Intelligence features is a key element of this role.You will shape the tools and infrastructure roadmap needed to power these functions. This includes maintaining clear, current documentation of processes, systems, and data flows across these functional areas, and identifying and prioritizing the tooling and infrastructure investments needed to support them as they scale — partnering with engineering and technical stakeholders to translate operational needs into build requirements.Lastly, you will also drive the operational scaling of our data infrastructure: how we manage, store, and govern the human-labeled goldsets that underpin safety benchmarking and LLM-judge validation, and how those pipelines extend to new languages, markets, and cultural contexts.
Minimum Qualifications
Bachelor's degree in a related field5+ years of experience in vendor/partner management, data operations, or program management, ideally supporting ML/AI evaluation, annotation, or data labeling programs
Demonstrated experience managing external vendor relationships and/or internal annotation operations partnerships at scale, including statements of work, budget, and performance management
Experience scaling human-labeled data pipelines, including dataset creation, quality calibration, storage, and governance
Experience standing up or scaling operations across multiple international markets, including working with multilingual and multicultural vendor workforces
Experience coordinating across functionally distinct but interdependent teams to close feedback loops and keep shared priorities aligned
Experience owning documentation and translating operational needs into tooling or infrastructure roadmaps, partnering with engineering to scope and prioritize build work
Strong ability to think strategically about operational tradeoffs while managing multiple concurrent initiatives in a fast-paced, evolving environment
Proven track record managing budgets and vendor relationships at scale
Proficiency with data tools (e.g., SQL, spreadsheets/BI tools) sufficient to track vendor performance, cost, and data quality metrics
Excellent written and verbal communication skills, with demonstrated ability to translate operational risk and status for both technical and non-technical stakeholders
Preferred Qualifications
Experience managing vendor or annotation partnerships specifically for AI safety evaluation, red-teaming, or Responsible AI programs
Familiarity with goldset design, labeling guidelines, or LLM-judge/evaluation workflows sufficient to partner effectively with technical teams
Experience building internationalization or localization operations for data labeling or evaluation programs
Familiarity with regulatory frameworks relevant to AI safety and content risk across international markets
Demonstrated success building operations or vendor programs from the ground up, including establishing new processes or partner relationships
Executive presence and experience presenting operational risk assessments and scaling plans to senior leadership