Data Center Production Operations Engineer
Track this application
Get Started FreeMatch score against your CV
Get Started FreeTailor your resume to this job
Get Started FreeAbout interviewing at Meta
Recruiter screen, a technical screen (60-minute asynchronous challenge or paired live coding), a short culture-fit questionnaire, then a virtual onsite of about four rounds: fast-paced coding with multiple problems per round, system design scoped to level, a lighter behavioral, and an AI-enabled coding round assessing how you work with a coding assistant.
Read the full Meta interview process →Description
About
Meta is seeking a Data Center Production Operations Engineer to support the reliability, efficiency, and scalability of our global data center infrastructure. In this role, you will be responsible for the day-to-day operational health of server fleets and production systems, working at the intersection of hardware lifecycle management, systems troubleshooting, and operational process improvement. You will partner closely with hardware engineering, capacity planning, and infrastructure teams to ensure Meta's data centers operate at peak performance, directly enabling the products and services that connect billions of people worldwide.
Responsibilities
- Monitor and maintain the operational health of large-scale server fleets and production infrastructure across data center environments
- Diagnose and resolve hardware and systems failures, coordinating with engineering teams to drive root cause analysis and implement corrective actions
- Execute and refine server deployment, decommissioning, and lifecycle management processes to support capacity and reliability goals
- Develop and maintain operational runbooks, escalation procedures, and documentation to standardize production operations workflows
- Collaborate with hardware engineering and capacity planning teams to identify systemic issues and propose infrastructure improvements
- Track and analyze operational metrics and failure trends to surface insights that improve fleet reliability and reduce mean time to resolution
- Support the qualification and rollout of new server hardware generations by validating operational readiness and identifying deployment risks
- Partner with cross-functional teams including network engineering, facilities, and software infrastructure to resolve complex production incidents
- Identify opportunities to automate repetitive operational tasks and contribute to tooling improvements that increase operational efficiency
- Provide technical guidance to peers on production operations best practices, hardware troubleshooting methodologies, and process standards
Minimum Qualifications
- 2+ years of experience in data center operations, production operations, or systems administration in a large-scale infrastructure environment
- Experience troubleshooting server hardware components including CPUs, memory, storage, and networking hardware in a production setting
- Experience developing or improving operational processes, runbooks, or standard operating procedures for data center or infrastructure teams
- Experience analyzing operational data or failure metrics to identify trends and drive reliability improvements
- Experience collaborating with cross-functional engineering teams to resolve production incidents and implement systemic fixes Background in capacity planning, hardware lifecycle management, or server deployment operations for hyperscale data centers
- Experience supporting hardware qualification or new server platform bring-up in a data center production environment
- Experience with fleet management tooling, asset tracking systems, or infrastructure monitoring platforms at scale
- Familiarity with scripting languages such as Python or Bash for automating operational workflows and data analysis tasks