Senior Site Reliability Engineer (Performance and Scalability) at Digital Zone
Poland
<p>Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not by owning every service yourself.</p><p>What you'll do</p><ul><li>Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load.</li><li>Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results.</li><li>Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale.</li><li>Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC.</li><li>Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on.</li><li>Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems.</li></ul><p></p> <p>What you'll bring</p><ul><li>5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at 10s of thousands of requests per minute.</li><li>A track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adopted.</li><li>Deep AWS experience and a solid grasp of Postgres performance and scaling.</li><li>Fluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similar.</li><li>A calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeep.</li></ul> <ul><li>Immediate, large-scale impact on a high-growth business</li><li>Top-of-the-market compensation packages</li><li>Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more</li></ul>
Apply Now