Social Discovery Group Logo

Social Discovery Group

Site Reliability Engineer (SRE)

Posted 8 Hours Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Canada
Mid level
In-Office or Remote
Hiring Remotely in Canada
Mid level
Own and improve production infrastructure reliability, deployments, Infrastructure-as-Code, Kubernetes environments, automation, CI/CD, monitoring, alerting, and observability. Investigate incidents, optimize system performance, maintain documentation and runbooks, and support DNS, WAF, CDN, and caching infrastructure. The role requires strong Linux administration, Bash scripting, networking, Git, and containerization skills, with independent ownership and collaboration across development and operations teams.
The summary above was generated by AI

Social Discovery Group (SDG) is a group of social discovery companies. SDG solves the problems of loneliness, isolation, and disconnection - transforming virtual intimacy into the new normal. SDG’s products redefine the way people interact and connect with one another.

Our portfolio includes social entertainment platforms designed to connect people online across different cultures and regions of the world.

We bring together a team of like-minded people and IT professionals who specialize in creating and developing globally impactful social discovery products. Our international team of digital nomads works remotely from all over the world.

We’re proud to be a two-time “Great Place to Work” winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).

We are looking for a Site Reliability Engineer (SRE) passionate about infrastructure reliability, automation, and the development of scalable production systems.

Your main tasks will be:

  • Own and improve production infrastructure reliability and stability
  • Prepare, execute, and support deployments and infrastructure changes
  • Build and maintain Infrastructure-as-Code solutions using Ansible and Terraform
  • Support and optimize Kubernetes-based and containerized environments
  • Develop automation scripts and internal operational tooling
  • Monitor system health, investigate incidents, and proactively improve observability
  • Participate in CI/CD improvements together with Development, QA, DevOps, and SRE teams
  • Work with monitoring and alerting systems to reduce downtime and improve system performance
  • Maintain technical documentation, runbooks, and operational procedures
  • Support DNS, WAF, CDN, and caching infrastructure where required

We expect from you:

  • 3+ years of experience in SRE, DevOps, System Administration, or Build/Release Engineering
  • Strong Linux administration and troubleshooting skills
  • Hands-on experience with Kubernetes and containerization technologies (Docker/Podman)
  • Experience with CI/CD pipelines, preferably GitLab CI
  • Practical experience with Infrastructure-as-Code and configuration management tools (Ansible and/or Terraform)
  • Experience with observability and monitoring tools such as Prometheus, Grafana, Zabbix, or VictoriaMetrics
  • Good understanding of networking fundamentals, DNS, HTTP/HTTPS, load balancing, and troubleshooting
  • Experience with Git and modern software delivery workflows
  • Ability to work independently, take ownership, and proactively improve infrastructure
  • Fluent Russian level for technical documentation and team communication

Nice to have:

  • AWS or GCP experience
  • RabbitMQ / AMQP experience
  • Cloudflare, Akamai, WAF, CDN experience
  • Experience with tracing and advanced observability tooling

What do we offer:

  • REMOTE OPPORTUNITY to work full-time;
  • Vacation 28 calendar days per year;
  • 7 wellness days per year (time off) that can be used to deal with household issues, to lie down and recover without taking sick leave;
  • Bonuses up to $5000 for recommending successful applicants for positions in the company;
  • 50% payment for professional training, international conferences, and meetings;
  • Corporate discount for English lessons;
  • ​Health benefits. According to the paychecks, if you are not eligible for corporate medical insurance, the company will compensate you with up to $1,000 gross per year per employee. This can be spent on self-purchase of health insurance or on doctor’s fees for yourself and close relatives (spouse, children);
  • Workplace organization. The company provides all employees with an equipped workplace and all the necessary equipment (table, armchair, wifi, etc.) in our offices or co-working locations. In the other locations, the company provides reimbursement of workplace costs up to $1000 gross once every 3 years, according to the paychecks. This money can be spent on the rent of the co-working room, on equipping the working place at home (desk, chair, Internet, etc.) during those 3 years;
  • Internal gamified gratitude system: receive bonuses from colleagues and exchange them for our merchandise, team building activities, massage certificates, etc.

Sounds good? Join us now!
The initial pay level or pay range for this role will be shared with candidates during the recruitment process and before the commencement of employment.

Similar Jobs

2 Days Ago
Remote
British Columbia, BC, CAN
Senior level
Senior level
Software
Design, deploy, maintain, and troubleshoot cloud-based AI infrastructure built on Kubernetes and open-source technologies. Optimize container infrastructure for reliability, performance, scalability, and security; resolve complex networking, storage, Linux, and Kubernetes issues; develop AI-driven DevOps automation; collaborate with distributed teams and customers; participate in code reviews; define technical strategies; and mentor team members and customers.
Top Skills: Ci/CdCncfDcgmDistributed SystemsGoGpu SchedulingInfinibandJavaScriptKubernetesLinuxMicroservicesMig/VgpuNvidia Ai EnterpriseNvidia-Certified HardwareNvlinkOpenshiftOpenstackPythonRancherRdma/RoceVMware
2 Days Ago
Remote
Canada
Senior level
Senior level
Information Technology • Cybersecurity
Designs and maintains scalable cloud and on-premise infrastructure, Infrastructure as Code, Kubernetes environments, CI/CD pipelines, data streaming, caching, observability, and incident response systems. The role owns AWS reliability and cost efficiency, supports progressive deployments, troubleshoots complex production issues, partners with software teams, and drives automation and continuous improvement across platform operations.
Top Skills: Alert ManagerAmazon RdsApache KafkaArgocdAWSChaossearchConfluent CloudElasticsearchGithub ActionsGoGoogle Cloud PlatformGrafanaHelmIstioJenkinsKubernetesKustomizeLaunchdarklyAzureNode.jsOpensearchOpsgeniePagerdutyPosthogPrometheusPythonRedisTerraformTerragrunt
3 Days Ago
In-Office or Remote
Senior level
Senior level
Cloud • Fintech • Software • Analytics
Lead reliability engineering for AWS and Kubernetes-based systems. Build secure CI/CD pipelines, improve observability, define SLOs, optimize distributed systems, manage incidents, and strengthen platform resilience. Partner with engineering, security, platform, and product teams while mentoring engineers and reducing operational toil through automation and developer tooling.
Top Skills: Amazon EksAWSAws CdkAws CloudwatchAws X-RayCi/CdCloudfrontGateway ApiGithub ActionsIamInfrastructure As CodeIstioKubernetesLambdaMtlsNode.jsOpentelemetryS3TypescriptVpc

What you need to know about the Vancouver Tech Scene

Raincouver, Vancity, The Big Smoke — Vancouver is known by many names, and in recent years, it has gained a reputation as a growing hub for both tech and sustainability. Renowned for its natural beauty, the city has become a magnet for professionals eager to create environmental solutions, and with an emphasis on clean technology, renewable energy and environmental innovation, it's attracted companies across various industries, all working toward a shared goal: advancing clean technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account