Manager, Site Reliability Engineering - Storage Layer Service
MongoDB
Job Description
MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, ## What you'll do partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Cork for our hybrid working model. Responsibilities - Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers - Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs - Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges - Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you - Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams - Possess a customer-focused mindset, treating internal developers as your primary users - Value efficiency in processes and operations, and have a track record of optimizing team workflows - Prefer automation over manual processes, fostering a culture of building software solutions to eliminate toil - Have deep technical familiarity with Kubernetes ecosystems, containerization technologies, and modern IaC tooling (e.g., Terraform, Crossplane, or Operators) so you can effectively guide the team's technical decisions - Have operated or supported stateful storage or database systems at scale and are comfortable with durability, consistency and recovery trade-offs - Excel at translating complex business and engineering …
Requirements
See the listing for full requirements.
Related jobs
More roles you might like
<h3>Description:</h3> <p>The key objectives of the TPRM Program are to: </p> <ul> <li>Assess the risk of third-party relationships which drive the rigor of risk management activities both during the third-party onboarding and ongoing… ## What you'll do </h3> <p>Risk Asses…
The worldwide data management software market is massive, forecasted to grow from approximately $82 billion in 2023 to approximately $137 billion in 2027. ## What you'll do have gained a deep understanding of MongoDB and its partner ecosystem and completed New Hire Training -…
The data management software market is transforming how organisations build and run applications. MongoDB is the leading developer data platform and the first database provider to IPO in more than 20 years. ## What you'll bring Excellent customer communication, prioritisation,…