[8SN] Senior Site Reliability / Production Support Engineer at Software Mind
Canada
<section class="job-section" id="st-companyDescription"><div><p class="googlejobs-paragraph--empty"></p><h2 class="title">Company Description</h2></div><div class="wysiwyg spl-wysiwyg"><p>We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!<br>  </p><p><strong>About the Client</strong></p><p>Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies.</p><p>You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production.</p><p>#LI-DNI</p></div></section><section class="job-section" id="st-jobDescription"><div><p class="googlejobs-paragraph--empty"></p><h2 class="title">Job Description</h2></div><div class="wysiwyg spl-wysiwyg" itemprop="responsibilities"><p><strong>About the Role</strong></p><p>We are looking for a <strong>Senior Site Reliability / Production Support Engineer</strong> to support the deployment, operations, and ongoing reliability of a production UI service running on Kubernetes.</p><p>This role is focused on maintaining highly available cloud-native applications, troubleshooting production issues, and improving operational excellence. You will work closely with engineering teams to monitor service health, investigate incidents, and ensure reliable service delivery.</p><p>While this role supports a UI-based service, it is <strong>not a frontend development position</strong>. Basic knowledge of Web Components is sufficient to perform first-level debugging when necessary.<br>  </p><p><strong>What You'll Do</strong></p><ul><li>Support deployment, operations, and ongoing maintenance of a production service running on Kubernetes.</li><li>Monitor application health, availability, and performance.</li><li>Investigate and resolve production incidents using logs, monitoring, and debugging tools.</li><li>Perform log analysis using Splunk to identify root causes and troubleshoot service issues.</li><li>Collaborate with software engineers to improve service reliability and operational efficiency.</li><li>Participate in incident response and production support activities.</li><li>Assist with first-level debugging of UI-related issues involving Web Components.</li><li>Contribute to continuous improvements in automation, monitoring, and operational processes.</li><li>Support CI/CD pipelines and cloud-native deployment practices.</li></ul></div></section><section class="job-section" id="st-qualifications"><div><p class="googlejobs-paragraph--empty"></p><h2 class="title">Qualifications</h2></div><div class="wysiwyg spl-wysiwyg" itemprop="qualifications"><p><strong>Required Qualifications</strong></p><ul><li>5+ years of experience in <strong>Site Reliability Engineering, DevOps, Platform Engineering, or Production Operations.</strong></li><li>Strong hands-on experience with <strong>Kubernetes </strong>in production environments.</li><li>Experience supporting cloud-native applications.</li><li>Experience monitoring production systems and troubleshooting complex incidents.</li><li>Strong knowledge of <strong>Splunk </strong>for log analysis and debugging.</li><li>Experience working in <strong>Linux </strong>environments.</li><li>Understanding of networking fundamentals and distributed systems.</li><li>Experience collaborating with software engineering teams to resolve production issues.</li><li>Strong troubleshooting and root cause analysis skills.</li><li>Excellent written and spoken English (B2+).</li></ul></div></section><section class="job-section" id="st-additionalInformation"><div><p class="googlejobs-paragraph--empty"></p><h2 class="title">Additional Information</h2></div><div class="wysiwyg spl-wysiwyg" itemprop="incentives"><p><strong>Preferred Qualifications</strong></p><ul><li>Experience with CI/CD pipelines.</li><li>Experience with cloud platforms such as AWS, Azure, or GCP.</li><li>Familiarity with container technologies such as Docker.</li><li>Exposure to observability tools (Prometheus, Grafana, OpenTelemetry, etc.).</li><li>Basic understanding of Web Components and frontend architecture.</li><li>Experience supporting high-availability enterprise SaaS platforms.</li><li>Knowledge of infrastructure automation or Infrastructure as Code (Terraform, Helm, Ansible, etc.) is a plus.</li></ul><p><strong>What We Offer</strong></p><ul><li>Competitive salary and laptop</li><li>Professional development and training opportunities</li><li>Work with cutting-edge cloud and container technologies</li><li>Flexible work arrangements and collaborative team environment</li><li>Impact on organization-wide digital transformation initiatives</li></ul></div></section>
Apply Now