OpusClip Logo

OpusClip

Senior Back-end Engineer (Infrastructure & Reliability)

Posted 5 Days Ago
Be an Early Applicant
In-Office
Burnaby, BC
Senior level
In-Office
Burnaby, BC
Senior level
The Founding SRE will ensure the reliability and scalability of the OpusClip platform, design infrastructure, optimize performance, and lead incident response efforts.
The summary above was generated by AI

🎨 OpusClip is the world's No.1 AI video agent, built for authenticity on social media.

We envision a world where everyone can authentically share their story through video, with no expertise needed. Within just 18 months of our launch, over 10 million creators and businesses have used OpusClip to enhance their social presence.

We have raised $50 million in total funding and are fortunate to have some of the most supportive investors, including SoftBank Vision Fund, DCM Ventures, Millennium New Horizons, Fellows Fund, AI Grant, Jason Lemkin (SaaStr), Samsung Next, GTMfund, Alumni Ventures, and many more.

Check out our latest coverage by Business Insider featuring our product and funding milestones, and our recognition as one of The Information's 50 Most Promising Startups in 2024.

Headquartered in Palo Alto, we are a team of 100 passionate and experienced AI enthusiasts and video experts, driven by our core values:

  • Be a Champion Team

  • Prioritize Ruthlessly

  • Ship fast, Quality Follows

  • Obsess over customers

Be a part of this exciting journey with us!

The Mission

We are looking for a hands-on Founding SRE to own the reliability and scalability of the OpusClip platform. You will stabilize our processing clusters, design isolated environments for our largest Enterprise partners, and serve as the technical bridge between infrastructure and our 15M+ users.

You will engineer the infrastructure strategy that underpins our trust and reliability in the market. You will help setup oncall rotation and own the full incident lifecycle from minimizing Time-to-Detect to tracking post mortem actions ensuring that our high-velocity growth never compromises our performance.

Key Responsibilities

Infrastructure Architecture & Cluster Operations

  • Architect Dedicated Environments: Lead the design and implementation of high-throughput, isolated processing clusters for Enterprise clients. You will build the "paved road" to ensure strict High Availability (HA) without noisy neighbor interference.

  • Scale Production: Drive general improvements in our Temporal clusters and production Kubernetes environments. You will operationalize scaling strategies that support both self-serve consumers and high-touch Enterprise contracts.

  • Technical Execution: Be hands-on with the stack to optimize resource allocation, reduce latency, and enforce isolation strategies for critical accounts.

Monitoring, Alerting & Detectability

  • Beat the Customer to the Alert: Overhaul our Datadog observability suite to aggressively reduce Time-to-Detect (TTD). You ensure we identify latency spikes and stalled projects before users do.

  • Threshold Tuning: tune alert thresholds to eliminate noise and focus on "symptom-based" alerts that reflect the actual user experience.

  • External SLO Ownership: Define and report on Service Level Objectives (SLOs), acting as the internal guarantor that we are meeting the targets we sold.

Incident Command & "Extreme Ownership"

  • First Responder & Mitigation: Serve as the first line of defense during outages. You will own immediate mitigation, including cluster debugging and manual scaling intervention if Horizontal Pod Autoscalers (HPA) fail or lag.

  • Drive Recovery Metrics: You are accountable for shortening Time-to-Mitigation (TTM) and Time-to-Recover (TTR). Your priority is to stop the bleeding first, then fix the wound.

  • Root Cause Analysis: Lead the post-mortem process to determine Time-to-Root Cause and implement systemic fixes. You will translate these technical findings into clear updates for Customer Experience (CX) and Leadership.

  • Accountability: Work collaboratively with Engineering Owners to track improvements against the reliability roadmap. You are responsible for flagging risks early and resetting expectations on platform performance when necessary.

  • Cross-Functional Bridge: Serve as the primary technical voice to the Customer Experience (CX), Sales, and Leadership teams. You will translate technical constraints and roadmaps into clear, honest updates for stakeholders.

Qualifications

  • Production K8s & Temporal: Expert-level ability to debug Kubernetes internals (HPA logic, node scaling) and operate stateful workflow engines (Temporal) at scale.

  • Incident Command: Proven track record as a primary first responder, demonstrating the ability to aggressively reduce Time-to-Mitigation (TTM) and Time-to-Recover (TTR).

  • Observability Architecture: Experience architecting Datadog SLOs and tuning alerts to distinguish system noise from actual user pain.

  • Automation: Strong proficiency in Python or Bash to automate manual recovery and operational tasks.

  • ⭐️ Bonus: Experience scaling GPU/video rendering workloads or thriving in early-stage, high-velocity startups.

The Tech Stack

  • Orchestration & Compute: Kubernetes (GKE), Docker, Horizontal Pod Autoscaling (HPA).

  • Workflow Engine: Temporal

  • Observability: Datadog (APM, Custom Metrics, Alerting).

  • Infrastructure as Code: Terraform or similar IaC tools.

  • Scripting & Backend: Python (primary), Bash.

  • Data & Storage: Redis, Milvus (Vector DB), Postgres.

Compensation

  • The base salary range for this position in the selected city is CAD$160k - $235k annually.

  • This role is also eligible for performance-based bonus/commission, with an annual range of CAD$16k - $58k.

  • Actual compensation may vary based on factors such as a candidate's qualifications, skills, experience, and geographic location.

  • On top of cash compensation, we offer competitive equity packages (ISOs) because if you're building the company, you should own a piece of it.

🎁 Our Benefits

  • Comprehensive medical, vision, and dental coverage support you and your wellbeing.

  • Flexible paid time off that empowers you to recharge, reset, and come back stronger.

  • Generous equipment, software, and office furniture budget, including MacBook, 4k monitor, standing desks and more, get what you need to be creative and productive.

  • Access to the latest technologies and tools.

  • Free lunch and dinner, plus a wide variety of global snacks and beverages for onsite positions in North America.

  • Visa sponsorship available, subject to eligibility.

EEO

OpusClip is proud to be an equal opportunity employer. We do not discriminate in hiring or any employment decision based on race, color, religion, national origin, age, sex (including pregnancy, childbirth, or related medical conditions), marital status, ancestry, physical or mental disability, genetic information, veteran status, gender identity or expression, sexual orientation, or other applicable legally protected characteristics. OpusClip considers qualified applicants with criminal histories, consistent with applicable federal, state and local law. Opus Clip is also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures.

Top Skills

Bash
Datadog
Docker
Gke
Kubernetes
Milvus
Postgres
Python
Redis
Temporal
Terraform

Similar Jobs

2 Hours Ago
Easy Apply
Remote or Hybrid
2 Locations
Easy Apply
Senior level
Senior level
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
The Engagement Manager leads enterprise deployments of Superhuman's AI productivity platform, managing the project lifecycle from scoping to go-live, while collaborating with cross-functional teams and ensuring customer success.
Top Skills: AsanaCodaConfluenceJIRAMiroMondayMs ProjectNotionSmartsheetWrike
14 Hours Ago
Hybrid
Burnaby, BC, CAN
Senior level
Senior level
Gaming • Information Technology • Mobile • Software • Esports
The Senior Level Designer will create a player-centric game world, collaborating across teams to drive the game's vision and design. Responsibilities include defining the vision, creating game experiences, and ensuring strong player interaction.
Top Skills: Unreal Engine
18 Hours Ago
In-Office or Remote
8 Locations
Senior level
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
As a Senior ASIC Test & Characterization Engineer, you will validate ASIC mining products by developing automation test scripts, conducting tests, and collaborating with cross-functional teams to ensure product quality.
Top Skills: AsicCElectronic Measurement EquipmentFpgaPython

What you need to know about the Vancouver Tech Scene

Raincouver, Vancity, The Big Smoke — Vancouver is known by many names, and in recent years, it has gained a reputation as a growing hub for both tech and sustainability. Renowned for its natural beauty, the city has become a magnet for professionals eager to create environmental solutions, and with an emphasis on clean technology, renewable energy and environmental innovation, it's attracted companies across various industries, all working toward a shared goal: advancing clean technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account