WitnessAI Logo

WitnessAI

Site Reliability Engineer - Platform Engineering

Posted 8 Days Ago
Be an Early Applicant
7 Locations
Mid level
7 Locations
Mid level
This role involves maintaining and enhancing system reliability, managing AWS infrastructure, optimizing Kubernetes clusters, developing CI/CD pipelines, and promoting SRE practices across the team.
The summary above was generated by AI

Job Title: Site Reliability Engineer (SRE), Platform Engineering

About Us: WitnessAI is a leader in providing innovative networking solutions designed to enhance security, performance, and reliability for businesses of all sizes.  We are seeking a highly skilled Site Reliability Engineer (SRE) with a strong background in Linux administration, AWS, and Kubernetes for our Platform Engineering team. The ideal candidate will help ensure the reliability, scalability, and performance of our systems while driving a culture of automation and continuous improvement.

Key Responsibilities

System Reliability & Operations

  • Maintain and improve the reliability, availability, and performance of our services and infrastructure.

  • Monitor system health, troubleshoot issues, and respond to incidents with a focus on reducing mean time to recovery (MTTR).

Infrastructure Management

  • Administer and optimize Linux-based systems across development, staging, and production environments.

  • Design and manage scalable, secure, and cost-effective solutions on AWS.

  • Build, maintain, and monitor Kubernetes clusters to support containerized applications.

Automation & Tooling

  • Develop and maintain CI/CD pipelines to streamline deployments.

  • Automate operational tasks using tools such as Terraform, Crossplane, or custom scripts.

  • Create and enhance monitoring, alerting, and logging systems to improve observability.

  • Build ad-hoc, reusable automation solutions where required.

Collaboration & Best Practices 

  • Partner with engineering teams to integrate SRE principles into the software development lifecycle.

  • Advocate for best practices in incident response, post-mortem reviews, and capacity planning.

  • Share knowledge with team members and contribute to a culture of continuous improvement.

Security & Compliance

  • Implement security best practices for cloud and containerized environments.

  • Ensure compliance with organizational and industry standards.

Requirements

Technical Skills

  • Proven expertise in Linux system administration (e.g., Ubuntu, CentOS, or similar).

  • Deep understanding of AWS services and architecture (e.g., EC2, S3, RDS, VPC, IAM).

  • Strong experience managing Kubernetes clusters in production.

  • Hands-on experience with infrastructure-as-code tools like Terraform or CloudFormation

  • Proficiency in scripting or programming languages (e.g., Python, Bash, or Go).

  • Demonstrated experience in app development for ba lend automation solutions.

  • 3+ years of experience in a Site Reliability Engineer, DevOps Engineer, or similar role working for a SaaS or Cloud bases company.

Operational Expertise

  • Familiarity with monitoring and logging tools such as Prometheus, Grafana, ELK, or Datadog

  • Experience designing and maintaining CI/CD pipelines (e.g., Jenkins, GitLab CI, or CircleCI).

  • Understanding of networking concepts (e.g., DNS, load balancing, firewalls).

Problem Solving & Collaboration

  • Strong analytical and troubleshooting skills.

  • Ability to work effectively in a collaborative, team-oriented environment.

  • Excellent written and verbal communication skills.

Education

Bachelor’s degree in Computer Science, Engineering, or equivalent work experience.

Nice-to-Have Skills:

  • Experience with service meshes and other CNCF technologies (e.g., Istio or Linkerd).

  • Knowledge of database systems (e.g., MySQL, PostgreSQL, or NoSQL databases).

  • Familiarity with cloud-native technologies and tools (e.g., Helm, ArgoCD, Spinnaker).

Benefits:

  • Hybrid work environment

  • Competitive salary.

  • Health, dental, and vision insurance.

  • 401(k) plan.

  • Opportunities for professional development and growth.

  • Generous vacation policy.

Salary range:

$170,000-$200,000

Top Skills

AWS
Bash
CircleCI
Datadog
Elk
Gitlab Ci
Go
Grafana
Jenkins
Kubernetes
Linux
Prometheus
Python
Terraform

Similar Jobs

6 Days Ago
Toronto, ON, CAN
Senior level
Senior level
Big Data • Cloud • Internet of Things
The Staff Software Engineer will enhance the reliability and performance of systems, automate infrastructure management, and support data engineering efforts.
Top Skills: Apache BeamApache KafkaSparkAWSAzureCloudFormationDatadogDockerGithub ActionsGCPGoogle Cloud DataflowGrafanaJavaJenkinsKubernetesPrometheusPythonTerraform
Senior level
Big Data • Cloud • Internet of Things
Seeking a Senior Platform Engineer & SRE to enhance system reliability, performance, and scalability through automation, monitoring, and collaboration with teams to implement DevOps best practices.
Top Skills: Apache BeamApache KafkaSparkAWSAzureCi/CdCloudFormationDatadogDockerGithub ActionsGCPGoogle Cloud DataflowGrafanaJavaJenkinsKubernetesPrometheusPythonTerraform
2 Hours Ago
Remote
Hybrid
6 Locations
Mid level
Mid level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
As a UI Engineer, you will build and maintain single-page web applications with Ember.js, collaborate with diverse teams, and drive technical aspects of product delivery.
Top Skills: AWSCSSEmberGitHTMLJavaScriptMochaPostcssQunitSassTailwind Css

What you need to know about the Vancouver Tech Scene

Raincouver, Vancity, The Big Smoke — Vancouver is known by many names, and in recent years, it has gained a reputation as a growing hub for both tech and sustainability. Renowned for its natural beauty, the city has become a magnet for professionals eager to create environmental solutions, and with an emphasis on clean technology, renewable energy and environmental innovation, it's attracted companies across various industries, all working toward a shared goal: advancing clean technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account