Network Reliability Engineer, Infrastructure Services
Track this application
Get Started FreeMatch score against your CV
Get Started FreeTailor your resume to this job
Get Started FreeAbout interviewing at Apple
One of tech's least standardized loops: you interview for a specific team, and every stage belongs to it. Recruiter/hiring-manager screen, one to three 45–60 minute Coderpad coding screens, a system design round shaped by Apple's reliability and privacy constraints, and a behavioral round — then a panel debrief. There is no central question bank; questions map to the team's real stack.
Read the full Apple interview process →Description
Summary
Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless.” Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here. Infrastructure Services is part of IS&T and the foundation of Apple's global network operations — managing data center equipment and systems to deliver compute, storage, and networking services for teams across Apple, including its internal developer community. From individual facilities to a worldwide network, Infrastructure Services ensures the technology underneath everything works without question.Description
We are seeking an experienced and visionary Network Reliability Engineer to drive the technical strategy and execution for ensuring the availability, performance, scalability, and resiliency of Apple's global network services. In this role, you will work as a technical leader solving complex networking challenges at massive scale, partnering with engineering, infrastructure, and operations teams across Apple to deliver reliable, fault-tolerant systems. As a technical leader within the Cloud Networking organization, you will define and drive the reliability and resiliency architecture for Apple's network platform services. You will be responsible for establishing SRE and SWE best practices, architecting fault-tolerant network control and data planes, and championing data-driven decision-making through observability and automation. You will drive resilient cloud networking solutions that operate reliably across multiple cloud providers and global regions, handling failures gracefully and maintaining service availability. Your technical leadership will ensure Apple's network services meet demanding availability, latency, resilience, and security requirements while continuously improving operational maturity. We are looking for a technical expert who deeply understands cloud networking at scale, is passionate about operating mission-critical, globally distributed infrastructure, preventing outages through proactive engineering, and driving long-term reliability improvements through architectural excellence.Key Responsibilities
Define and drive the long-term technical vision, architecture, and reliability strategy for large-scale cloud networking platforms spanning control plane and data plane systems. Architect and evolve fault-tolerant, highly available network services, ensuring graceful degradation and consistent performance under partial and systemic failure scenarios. Establish platform-wide resiliency patterns including service discovery, health checking, automated failover, rate limiting, circuit breaking, and traffic management across multi-region and multi-cloud environments. Lead the design of network configuration management, routing state distribution, traffic engineering, and capacity planning systems, balancing scalability, correctness, and operational simplicity. Serve as a senior technical authority and architectural reviewer, influencing critical design decisions across multiple teams and ensuring network failure modes are explicitly addressed. Build and champion automation-first reliability solutions, including topology discovery, deployment safety mechanisms, self-healing systems, and operational tooling that reduce toil and improve MTTR. Define and own reliability metrics and observability standards (SLIs, SLOs, error budgets), using data to drive engineering trade-offs, reliability investments, and incident response improvements. Multiply impact through cross-team technical leadership, embedding reliability early in design, mentoring engineers, and sharing deep technical knowledge through documentation and technical talks.Minimum Qualifications
Extensive experience in software engineering, systems engineering, or infrastructure engineering. Strong background in designing, operating, and supporting highly available, fault-tolerant distributed systems at hyper scale. Strong systems programming skills including multi-threading, concurrency, caching, batching Solid understanding of network infrastructure and software-defined networking (SDN). Ability to lead cross-functional collaboration and influence technical decisions across teams.Preferred Qualifications
Expert knowledge of API design and interface technologies (JSON, ProtoBuf, REST, RPC, XML, etc) In depth knowledge of K8s, OpenStack, system virtualization, build systems and infrastructure as code Strong knowledge of observability systems (metrics, logging, tracing) and qualification engineering. Broad knowledge of networking solutions across OSI layers 3 through 7. Excellent written and verbal communication skills with the ability to clearly articulate risk, reliability trade-offs, and operational priorities. Proven ability to manage competing priorities, drive initiatives to completion, and deliver results in fast-paced environments.More engineering roles at Apple
Software Engineer, Apple Services Engineering - CommerceNew
Apple·Cupertino
Software Development Engineer in Test – Apps, Identity Management ServicesNew
Apple·Sunnyvale
Machine Learning (ML) Data Scientist - ISE Analytics and User Studies, Input ExperienceNew
Apple·Cupertino
Software Screening & Integration Engineer, Generative AI Experiences Software - Special ProjectsNew
Apple·Cupertino
AIML - Machine Learning Researcher, MLRNew
Apple·Cambridge