Site Reliability Engineer (SRE) - GCP at Devsu
Worldwide
<p></p><p>We are seeking a Site Reliability Engineer (SRE) with deep expertise in monitoring, observability, and reliability engineering to support systems running across on-premises infrastructure and Google Cloud Platform (GCP).</p><p>This role is primarily responsible for designing, operating, and improving monitoring, alerting, and observability platforms, with a strong focus on Grafana and Kubernetes environments.</p><p>As a secondary responsibility, this role provides backup coverage for the Application Support team during periods of resource constraints or major incidents, offering L2/L3 technical support when required.</p><p></p><p></p>Responsibilities<p>Monitoring & Observability (Core Focus)</p><ul><li>Own and operate the monitoring and observability stack across on-prem and GCP environments</li><li>Design, build, and maintain Grafana dashboards for infrastructure, Kubernetes, and applications</li><li>Define, tune, and maintain alerts to ensure high signal-to-noise ratio</li><li>Establish observability standards and best practices across teams</li><li>Improve visibility into system health, performance, and reliability</li></ul><p></p><p>Site Reliability Engineering</p><ul><li>Apply SRE principles to improve availability, performance, and resilience</li><li>Define and track SLIs, SLOs, and error budgets</li><li>Participate in on-call rotations and SEV incident response</li><li>Lead or contribute to incident investigations and root cause analysis (RCA)</li><li>Drive preventative actions to reduce repeat incidents</li></ul><p></p><p>Kubernetes & Platform Reliability</p><ul><li>Support and monitor Kubernetes environments (GKE and on-prem clusters)</li><li>Monitor cluster health, capacity, and resource utilization</li><li>Troubleshoot platform-level issues impacting application reliability</li><li>Collaborate with Platform and Engineering teams on reliability improvements</li></ul><p></p>Secondary Responsibilities (Backup Application Support)<ul><li>These responsibilities are activated as needed, not part of day-to-day operations.</li><li>Provide L2/L3 application support coverage during:</li></ul><ul><ul><li>Support team resource shortages</li><li>High-severity incidents (SEVs)</li><li>Peak support periods or escalations</li></ul></ul><ul><li>Triage and troubleshoot application issues using existing runbooks and dashboards</li><li>Collaborate with Application Support and Engineering teams during incidents</li><li>Ensure all actions, findings, and resolutions are documented in ServiceNow (SNOW)</li></ul> <ul><li>Strong experience as a Site Reliability Engineer or Reliability Engineer</li><li>Deep hands-on expertise with Grafana (dashboards, alerting, troubleshooting)</li><li>Solid experience with monitoring and observability systems</li><li>Production experience operating Kubernetes environments</li><li>Experience supporting systems in GCP and on-prem environments</li><li>Strong Linux systems and troubleshooting skills</li><li>Fluent English (written and spoken).</li><li>Ability to work in PST time zone.</li><li>Ability to participate in an on-call rotation that includes coverage for one weekend day. Time worked during the weekend is compensated with one day off during the week, in accordance with the established work schedule.</li></ul><p></p><p>Technology Stack:</p><ul><li>Observability: Grafana, Prometheus, logging platforms</li><li>Containers: Kubernetes (GKE and on-prem)</li><li>Cloud: Google Cloud Platform (GCP)</li><li>Operations: Linux, networking, infrastructure monitoring</li><li>Incident Tools: PagerDuty, ServiceNow, Slack (or equivalents)</li></ul><p></p><p>Nice to have: </p><ul><li>Experience supporting application teams during SEV incidents</li><li>Knowledge of capacity planning and performance tuning</li><li>Scripting skills (Python, Bash, etc.)</li><li>Experience with hybrid infrastructure environments</li></ul><p></p> <p>At Devsu, we believe in creating an environment where you can thrive both personally and professionally. By joining our team, you’ll enjoy:</p><ul> <li>A stable, long-term contract with opportunities for career growth</li> <li>Private health insurance</li> <li>A remote-friendly culture that promotes work-life balance</li> <li>Continuous training, mentorship, and learning programs to keep you at the forefront of the industry</li> <li>Free access to AI training resources and state-of-the-art AI tools to elevate your daily work</li> <li>A flexible Paid Time Off (PTO) policy as well as paid holiday days</li> <li>Challenging, world-class software projects for clients in the US and LatAm</li> <li>Collaboration with some of the most talented software engineers in Latin America and the US, in a diverse work environment</li> </ul><p>Join Devsu and discover a workplace that values your growth, supports your well-being, and empowers you to make a global impact.</p>
Apply Now