Manager, TechOps / Site Reliability Engineering
Track this application
Get Started FreeMatch score against your CV
Get Started FreeTailor your resume to this job
Get Started FreeAbout interviewing at Apple
One of tech's least standardized loops: you interview for a specific team, and every stage belongs to it. Recruiter/hiring-manager screen, one to three 45–60 minute Coderpad coding screens, a system design round shaped by Apple's reliability and privacy constraints, and a behavioral round — then a panel debrief. There is no central question bank; questions map to the team's real stack.
Read the full Apple interview process →Description
Summary
Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless. Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here. Customer Systems is part of IS&T and drives the technology behind Apple's customer support experience — from contact center operations to the software powering the iconic Genius Bar. The team also builds and operates AppleCare's online support platform, which handles 6 billion visits per year, delivering seamless, high-quality support to Apple customers around the globe.Description
The Technical Operations Engineer will support and improve business-critical Customer Systems services used across Apple’s global customer-support organizations. The role combines advanced production support, automation, reliability engineering and Operational Excellence. The engineer will be responsible for service health, incident response, root-cause analysis and continuous improvement of operational processes, tooling and platform reliability.Key Responsibilities
Operate globally distributed, business-critical customer-support applications and platforms. Monitor service availability, latency, capacity, throughput and error rates. Lead troubleshooting of complex application, infrastructure, database, integration and network issues. Participate in or lead major-incident response, service restoration and stakeholder communication. Perform root-cause analysis and drive permanent corrective actions. Identify recurring incidents, operational gaps and reliability risks, and convert them into improvement initiatives. Build automation, diagnostics and self-service tools to reduce manual operational work. Develop dashboards, health checks, actionable alerts and automated remediation. Drive Operational Excellence across incident, problem, change, release and capacity management. Define and track service metrics, operational KPIs, SLIs and SLOs. Support Java applications, REST APIs, microservices and asynchronous integrations. Coordinate production releases, configuration changes, deployment validation and rollback planning. Lead production-readiness and operational-readiness reviews for new services and major changes. Partner with software engineering teams to improve application operability, resilience and supportability. Maintain runbooks, troubleshooting guides and service documentation. Support capacity planning, resilience testing and disaster-recovery exercises. Participate in on-call rotations and support critical launches and business events. Mentor junior engineers and promote a culture of ownership and continuous improvement.Minimum Qualifications
7+ years of experience in Technical Operations, production engineering, application support or Site Reliability Engineering. Strong troubleshooting skills across Linux, applications, APIs, databases and networks. Experience deploying, monitoring and supporting Java-based enterprise applications. Experience with scripting or software development using Python, Java, Go or shell. Experience with logging, monitoring and observability tools such as Splunk. Understanding of HTTP, DNS, TCP/IP, load balancing and distributed systems. Experience with incident, problem, change and release-management processes. Demonstrated ability to use operational metrics and incident trends to drive service improvements. Strong communication, documentation and stakeholder-management skills.Preferred Qualifications
Experience with Kubernetes, container platforms and Alicloud services. Knowledge of Oracle, PostgreSQL, MongoDB, Kafka or similar technologies. Experience building operational automation and automated remediation. Experience with high availability, disaster recovery and production-readiness practices.More engineering roles at Apple
Apple TV Business Manager - FR & BLXNew
Apple·Paris
Software Development Engineer, IS&T Ai & Data PlatformsNew
Apple·Shanghai
Language Engineer (Hindi)New
Apple·Hyderabad
Senior Backend Platform Software Engineer - Special ProjectsNew
Apple·Cupertino
Platform Software Engineer, Audio Software Integration
Apple·Cupertino