System DEvelopment Engineer, Automation, RTS-Solution Performance & Automation
Track this application
Get Started FreeMatch score against your CV
Get Started FreeTailor your resume to this job
Get Started FreeAbout interviewing at Amazon
Recruiter screen and (for most SDE roles) an online assessment, one technical phone screen, then a single-day interview loop of four to six 45–60 minute sessions mixing coding, system design, and Leadership-Principle behavioral rounds. A Bar Raiser from outside the hiring team joins the loop and can veto the hire.
Read the full Amazon interview process →Description
Key job responsibilities
Build and operate production monitoring services that evaluate robotics telemetry and performance signals for AR solutions, and trigger defined mechanisms (e.g., Andon, escalation, incident workflows) when conditions degrade.
Implement signal evaluation and automation frameworks that are reusable across solutions; default to “build once, standardize, reuse everywhere” and document exceptions with rationale.
Develop closed-loop operational automation including incident creation, routing, enrichment, and escalation integration with partner systems; reduce manual triage and repeated investigative effort.
Improve signal quality and operator trust by reducing false positives/alert noise, tuning mechanisms, and instrumenting measurement for alert precision and operational outcomes.
Build internal tooling that accelerates investigation and diagnosis for field/support stakeholders (explainability views, diagnostics tooling, guided triage workflows), grounded in “mechanisms over dashboards.”
Own the end-to-end lifecycle of your services (design, implementation, testing, deployment, operations). When systems fail, ensure contributing causes are identified and eliminated with permanent fixes, not just mitigations.
Be active in engineering and operational review mechanisms including code reviews, operational readiness reviews (ORRs), correction-of-errors (COEs), and post-incident analyses; use these mechanisms to drive resilience improvements and to coach peers.
Balance constraints and explicitly manage short-term workarounds: avoid them where possible, replace them with long-term solutions, or escalate over-use when it creates systemic risk.
Partner with Data Engineering on pipeline SLAs and reliability. Coordinate with solution/product engineering and field stakeholders through defined interfaces.
Produce clear, accurate, inclusive documentation (technical runbooks, playbooks, operational procedures, and design notes) so systems and artifacts can be maintained and extended by engineers unfamiliar with them.
About the team
Solution Performance & Automation (SPA) is RTS’s dedicated team for network monitoring mechanisms and automation across AR solutions. We provide a single, governed point of ownership for performance definitions, operational monitoring standards, and cross-team adoption. We build closed-loop monitoring mechanisms that define signals, automate Andon and escalation, and enable support scale through faster, informed decision-making.
Basic Qualifications
- Bachelor's degree in computer science or equivalent- Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
- Experience in automating, deploying, and supporting large-scale infrastructure
- Experience with Linux/Unix
- Experience with version control systems and CI/CD pipeline implementation
- Experience in automation or monitoring frameworks, deployment or development
- 4+ years of systems design, software development, operations, automation, and process improvement experience
- Experience troubleshooting and debugging technical systems
Preferred Qualifications
- Experience with distributed systems at scale- Experience in any of the following: Cloud Architecture, Systems Design, Software Development, Infrastructure Architecture, Data Engineering or DevOps
- Experience using data, reporting, or tools to measure performance and make adjustments accordingly
- Experience in complex work environments, including (but not limited to robotics, automation, diagnostic and test equipment)
- Experience working in a collaborative team environment to deliver high-quality design solutions
- Experience building monitoring/observability systems, alerting mechanisms, and signal evaluation frameworks at scale
- Experience with cloud infrastructure (AWS or equivalent), CI/CD, and operational readiness practices (ORRs/COEs/post-incident mechanisms)
- Demonstrated ability to reuse/extend existing systems, make pragmatic tradeoffs, and reduce operational load through durable mechanisms
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, MA, North Reading - 129,200.00 - 174,800.00 USD annually