Site Reliability Engineer (SRE)

1 month from now
CA > Vancouver
Managed Services Team

Job Description

WHAT YOU’LL DO

Monitor platform health across multiple client environments using tools like Grafana and Prometheus, or other monitoring tools

Respond to and triage incidents, following established runbooks and escalation paths

Participate in post-incident reviews and contribute to postmortem documentation

Support the Service Desk team with technical triaging, incident classification, and resolution

Maintain and improve observability dashboards, alerts, and SLI, and SLO tracking

Write and maintain runbooks, operational documentation, and knowledge base articles

Identify recurring issues and propose automation or process improvements to reduce toil

Participate in on-call rotation covering weekends (alternating schedule — one weekend on, one weekend off)

Collaborate with the EMEA team during shift overlap to ensure smooth handoffs and continuity

Support root cause analysis and contribute to continuous improvement initiatives

WHAT WE’RE LOOKING FOR:

Strong proficiency in English (written and verbal communication) is required

2–3 years of experience in SRE, platform operations, or a technical Service Desk role

Experience with monitoring and observability tools such as Grafana, Prometheus, or equivalent

Solid understanding of incident management processes (triaging, escalation, postmortems)

Experience supporting e-commerce platforms

Scripting skills in shell and/or Python for automation and operational tasks

Familiarity with containerization concepts (Docker, Kubernetes) at an operational level

Experience working in Agile environments and using ticketing tools (e.g. Jira)

Comfort working independently during early-morning shifts with minimal supervision

Strong documentation habits and attention to detail

Experience with Agile processes, testing, and code review

Strong experience with scripting - shell, Python, etc.

Excellent customer service attitude, communication skills (written and verbal), and interpersonal skills

Excellent analytical and problem-solving skills

Ability to communicate effectively with technical and non-technical stakeholders. You should feel comfortable explaining technical concepts in simple terms

Experience working in fast-paced, Agile environments, balancing priorities across multiple projects

**NICE TO HAVE: **

Experience with Google Cloud Platform (GCP) or other major cloud providers (AWS, Azure)

Familiarity with CI/CD pipelines (GitHub Actions, GitLab CI)

Basic experience with Infrastructure as Code tools such as Terraform

Basic knowledge of AIOps concepts and their application in operational workflows

SRE or cloud certifications (Google Cloud, AWS, Kubernetes)

Not sure? Upload your CV!
Quick Match

Let us do the work—upload your CV and get matched to jobs automatically.

We'll only use your CV to match you to jobs. No spam.

Related Jobs You Might Like 🔥