Role overview
Doctolib is hiring for the Staff Site Reliability Engineer (x/f/m) role in Berlin, Berlin, DE. It is a position, Lead level, in the Tech sector. It was posted yesterday.
On TalentyGo you can review this job and apply more effectively: Charlie prepares an ATS-optimized resume and a cover letter tailored to "Staff Site Reliability Engineer (x/f/m)" at Doctolib in about a minute. Before you apply, you can also check how well your profile fits, with a match score based on skills, experience, location and seniority.
- Role
- Staff Site Reliability Engineer (x/f/m)
- Company
- Doctolib
- Location
- Berlin, Berlin, DE
- Work mode
- On-site
- Seniority
- Lead
- Sector
- Tech
- Posted
- yesterday
Description
<h3><strong>Your Impact</strong></h3>
<p>We are looking for a <strong>Staff Site Reliability Engineer</strong> to join our SRE team dedicated to platform reliability within Platform Engineering.</p>
<p>Your mission will be to act as a technical leader driving Doctolib's reliability and scalability at a European scale, ensuring our platform remains reliable, debuggable, and resilient across infrastructure, observability, and cross-cutting reliability initiatives. You will play a pivotal role in a team driving reliability standards across 170+ applications, contributing directly to supporting 520,000 health professionals and 90 million patients in their daily healthcare journey.</p>
<p>This role sits at the intersection of infrastructure, developer experience, and product engineering. You'll act as a technical leader and strategic partner to SREs, software engineers, and product teams, guiding decisions, mentoring engineers, and driving cross-cutting initiatives that elevate our operational maturity.</p>
<h3><strong>What you'll do</strong></h3>
<p>Your responsibilities include but are not limited to:</p>
<ul>
<li>Lead large-scale cross-cutting reliability initiatives across the platform, spanning infrastructure automation, observability, and incident management</li>
<li>Identify and drive improvements to incident detection, response, and postmortem analysis capabilities</li>
<li>Define and evolve SLOs, error budgets, and alerting standards across multiple product teams</li>
<li>Take part in the on-call rotation, and actively contribute to improving our on-call experience by refining alerting, reducing noise, and ensuring actionable telemetry</li>
<li>Serve as a mentor and technical coach to senior engineers, helping elevate the craft of reliability engineering across the company</li>
<li>Influence strategic decisions by providing technical guidance to leadership and representing reliability engineering in architectural reviews and platform discussions</li>
<li>Partner with software engineering teams to embed reliability practices early in the development lifecycle</li>
</ul>
<h2><strong>Who you are</strong></h2>
<p>Before you read on: if you don't have the exact profile described below, but you feel this job description matches your skill set, we still encourage you to apply.</p>
<p><strong>You'll be a great fit if you:</strong></p>
<ul>
<li>Have extensive experience (8+ years) in SRE, platform engineering, or infrastructure roles within a large-scale, multi-team production environment</li>
<li>Have proven experience with cloud platforms such as AWS, GCP, or Azure</li>
<li>Have strong experience with containerization and orchestration technologies, Kubernetes is a must, its deployment and scaling strategies ecosystem</li>
<li>Have implemented and operated SLIs, SLOs, and error budgets in production</li>
<li>Have experience managing on-call rotations and leading incident response in high-stakes environments</li>
<li>Have a strong systems engineering background with fluency in at least one backend programming language (e.g., Go, Python, Ruby)</li>
<li>Have a proven ability to lead through influence: setting technical direction, driving consensus, and mentoring engineers across teams</li>
<li>Are comfortable balancing long-term architecture work with fast, iterative improvements</li>
<li>Have clear, concise communication skills, both written and verbal, with the ability to drive alignment in ambiguous environments</li>
<li>Partner with feature teams to accelerate their production readiness, providing hands-on guidance on reliability best practices, launch reviews, and operational standards before go-live</li>
<li>Are fluent in English</li>
</ul>
<p><strong>It would be fantastic if you:</strong></p>
<ul>
<li>Have deep expertise in observability tooling and architecture (logging, tracing, metrics)</li>
<li>Have experience designing and operating high-scale telemetry pipelines and working with developers to improve instrumentation quality</li>
<li>Appreciate working in regulated environments, healthcare, fintech, or similar; at Doctolib, data privacy and compliance are part of every engineering decision</li>
<li>Care about reliability enablement, golden paths, runbooks, shared libraries</li>
</ul>
<h2><strong>Life at Doctolib Tech</strong></h2>
<ul>
<li>Our solutions are built on a single fully cloud-native platform that supports web and mobile app interfaces, multiple languages, and is adapted to country and healthcare specialty requirements.</li>
<li>Our stack is composed of Rails, TypeScript, Java, Python, Kotlin, Swift, and React Native.</li>
<li>We leverage AI ethically across our products to empower patients and health professionals. Discover our AI vision<a href="https://careers.doctolib.com/ai-careers/"> here</a>.</li>
</ul>
<p>Want to learn more about our tech culture and environment? Visit the<a href="https://careers.doctolib.com/tech-doctolib/"> Doctolib Tech site</a>.</p>
<h2><strong>What we offer</strong></h2>
<ul>
<li>A Deutschlandticket (Germany-wide public transport pass) fully paid for by Doctolib</li>
<li>28 vacation days + 1 additional day for each full calendar year of employment (up to a maximum of 30 days)</li>
<li>Work from abroad for up to 10 days per year thanks to our flexibility days policy</li>
<li>Company health insurance with great supplementary benefits through our partner Allianz</li>
<li>Company pension scheme (bAV) through Allianz with an employer subsidy of 40% (15% within the probationary period)</li>
<li>The Doctolib Parent Care program, which includes one month additional parental leave and much more</li>
<li>Enrollment in Doctolib's long-term employee value sharing plan called DoctoGrowth</li>
<li>Free mental health and coaching services through our partner Moka.care</li>
<li>Subsidized sports membership through our partner Urban Sports Club</li>
<li>A flexible workplace policy offering both hybrid and office-based mode</li>
<li>Alongside healthy snacks and our regular breakfast buffet, we provide a subsidized meal benefit</li>
<li>For caregivers and workers with disabilities, a package including an adaptation of the remote policy, extra days off for medical reasons, and psychological support</li>
<li>Relocation support in case of international mobility</li>
<li>Access to the best AI tools for coding, development and dedicated training</li>
</ul>
<h2><strong>Our interview process</strong></h2>
<ul>
<li>Recruiter Interview</li>
<li>System Design Interview</li>
<li>Technical SRE Interview</li>
<li>Behavioral Interview</li>
<li>At least one reference check</li>
</ul>
<p>We want your experience to be clear, respectful, and transparent. Learn more about our hiring process on our<a href="https://careers.doctolib.com/tech-doctolib/getting-hired/"> candidate experience page</a>.</p>
<h2><strong>Job details</strong></h2>
<ul>
<li>Permanent position</li>
<li>Tech stack: Kubernetes, Terraform, AWS / GCP, Prometheus, OpenTe
TalentyGo is an aggregator of job postings from public sources. Always verify information directly with the company. Applications go through the original company website; TalentyGo does not manage hiring processes.