Site Reliability Engineer — Insight Team

Apple·Shanghai, Shanghai, China·posted 23d ago · last seen 44m ago

Track this application

Get Started Free

Match score against your CV

Get Started Free

Tailor your resume to this job

Get Started Free

About interviewing at Apple

One of tech's least standardized loops: you interview for a specific team, and every stage belongs to it. Recruiter/hiring-manager screen, one to three 45–60 minute Coderpad coding screens, a system design round shaped by Apple's reliability and privacy constraints, and a behavioral round — then a panel debrief. There is no central question bank; questions map to the team's real stack.

Read the full Apple interview process →

Description

Summary

The Insight team runs one of Apple's most critical Big Data ecosystems — an exabyte-scale, highly-available infrastructure that underpins manufacturing operations for every Apple product, globally. Every iPhone, iPad, and Mac has touched our systems. We advance technology by relying on each other's strengths and skills to build something bigger than ourselves. For this reason, team culture is central to our values. We value social skills and integrity as much as technical craft. We are looking for extraordinary DevOps with experience building large-scale data platforms, analytic tools and solutions which can help take our platform to the next level. Do you excel in a high-demand setting and exceed expectations, in an environment that requires time-management? The right person will prioritize tasks and complete assignments ahead of schedule. While being a great standout colleague, you will also work independently.

Description

The primary mission of this role is to ensure the stability and performance of our production environment, including the secure and reliable execution of all deployment and monitoring processes. We expect this role to go beyond operations by fully embracing a dev-and-ops mindset. Beyond standard support and troubleshooting, you will leverage AI tools and automation to drive efficiency, applying your engineering expertise to continuously upgrade and optimize our Big Data and microservices ecosystem.

Key Responsibilities

You will operate and improve services within a very large-scale, highly available Big Data ecosystem supporting Exabytes level of data with sustained, rapid growth. Your work directly enables the engineering and operations teams that build every Apple product. - Own the reliability, performance, and scalability of services and cloud infrastructure within the Insight ecosystem - Build and advance AIOps capabilities across the ecosystem — developing AI-driven alerting, anomaly detection, LLM-assisted operational tooling, and automated incident triage that measurably improve observability and response - Instrument services for deep observability: dashboards, meaningful alerts, runbooks, and clearly defined SLOs and SLIs that give the team and stakeholders an accurate view of ecosystem health - Partner with incident management to drive effective incident response and conduct thorough post-incident reviews; translate findings into architectural and operational improvements that prevent recurrence - Collaborate with cross-functional engineering teams across Apple's global manufacturing services, communicating effectively across time zones and cultures

Minimum Qualifications

BS or MS in Computer Science, Software Engineering, or an equivalent technical discipline Good programming skills in Java, Python or Go Hands-on experience in SRE, DevOps, or platform engineering, with demonstrated ability to work on technical initiatives Experience applying AI/ML techniques to development & operations Proficiency in English for clear technical communication, documentation, and global collaboration Availability to participate in SRE on-call rotations during China morning hours (8:00 AM CST during US Daylight Saving Time, and 9:00 AM CST during US Standard Time)

Preferred Qualifications

Cloud Technologies: Hands-on experience with cloud platforms (AWS or GCP), containerization and orchestration (Docker, Kubernetes), and multi-region infrastructure design including data residency considerations and IAM. Distributed Systems & Big Data Platforms: Strong background in microservices, APIs, and messaging/streaming (Kafka, Solace, Pub/Sub), combined with production experience in relational databases (MySQL, PostgreSQL) and big data/storage technologies (Redis, Bigtable, Druid, Elasticsearch, ClickHouse). CI/CD & Observability: Proven ability to build and maintain CI/CD pipelines or GitOps workflows (ArgoCD, Jenkins, GitHub Actions), and operate observability systems (Grafana, Prometheus, Kibana) including service instrumentation, dashboard design, and alert tuning. AI tooling & Problem Solving: Forward-looking experience in developing MCP and AI Agentic flows, paired with a flexible and creative approach to solving complex technical problems.

More engineering roles at Apple