- Remote
- Full time
Job description
## What you'll do architect systems with reliability designed in from the outset, define the SLIs and SLOs that determine whether we are meeting customer expectations, and build the automation that keeps our infrastructure efficiently ahead of capacity and performance demand. At the Staff SRE you work without day-to-day guidance, applying deep subject-matter knowledge and industry-leading practice to improve the products, processes, and services that Twilio runs on. You will own the reliability posture of significant parts of our production estate. Your impact will be felt across multiple teams rather than within one. This is a hands-on engineering role. You will write and deploy code that improves service reliability, orchestrate complex changes across systems, lead the response when production is degraded, and raise the quality bar for the engineers around you. Responsibilities In this role, you’ll: - Own the reliability posture of production services in your area — availability, latency, capacity, efficiency, performance, and the monitoring and alerting that makes them visible - Define, instrument, and operate against SLIs and SLOs, and use error budgets to drive engineering priorities - Identify trends and problem areas that threaten stability, and provide a path forward to mitigate risk before it reaches customers - Drive down repair items and prevent classes of incidents rather than resolving them one at a time - Improve detection, response, and recovery — reducing time to acknowledge, engage, mitigate, and restore, with fewer people pulled in - Design for failure: strengthen failure domains, validate recovery paths, and make production changes safer to ship and safer to roll back - Participate in on-call for the services you support, and lead the response when production is degraded - Write post-mortems that identify true root causes, and drive the follow-up work to completion - Oversee efforts to identify, diagnose, report, and document production problems across all reliability dimensions - Write, configure, and deploy code that measurably improves service reliability — maintainable, reviewed, documented, and well tested - Orchestrate complex changes across systems and services, documenting design changes, technical decisions, migration plans, and upgrades - Lead debugging, troubleshooting, and analysis of service architecture and design - Use code review to drive up the quality of your coworkers' code - Reduce the operational overhead required to run infrastructure and services - Drive projects from conception to completion for efforts spanning the concerns of your team - Coordinate across programs, collaborating with others to estimate and communicate delivery timelines - Break projects into milestones and tasks, track progress, and communicate updates to stakeholders - Identify and communicate changes that may impact stability …
Requirements
See the listing for full requirements.
Related jobs
More roles you might like
- Remote
- Full time
- Remote
- Full time