Mirantis Logo

Mirantis

Senior Data Platform Engineer — Kafka & PostgreSQL

Posted 2 Hours Ago
In-Office or Remote
Hiring Remotely in Remote, OR
Senior level
In-Office or Remote
Hiring Remotely in Remote, OR
Senior level
Operate PostgreSQL and Kafka data platforms on Kubernetes across cloud and bare-metal environments. Own GitOps, infrastructure as code, automation, upgrades, observability, high availability, backups, disaster recovery, security, multi-tenancy, and cross-region replication. Build self-service database and topic provisioning, operate CDC pipelines, support incident response, and develop tooling in Go or Python. Partner with application teams on schema standards, performance, and reliable stateful-system operations.
The summary above was generated by AI
Company Description

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.  https://www.mirantis.com/

Job Description

We are looking for a Senior Data Platform Engineer to run the PostgreSQL and Apache Kafka platform behind k0rdent-ai — our multi-tenant control plane for enterprise GPU infrastructure. Every cluster provisioned, every GPU-hour consumed, and every tenant action is recorded on this platform, so it must be reliable, secure, and recoverable.

This is a hands-on DevOps / SRE role for stateful systems. You will deploy and operate PostgreSQL (CloudNativePG) and Kafka (Strimzi) on Kubernetes using operators, Helm, and GitOps — across cloud and bare-metal clusters, in a global control plane and multiple regions. You will own automation, upgrades, observability, backups, and disaster recovery, and give product teams self-service access to databases and topics through code.

Main Responsibilities

  • Run PostgreSQL and Kafka on Kubernetes: Deploy, upgrade, scale, and operate CloudNativePG and Strimzi clusters with Helm charts and operator custom resources, across cloud and bare-metal Kubernetes.
  • GitOps & Infrastructure as Code: Manage everything declaratively with Argo CD (or Flux), Helm, and Terraform; build CI/CD pipelines for database and Kafka changes, schema migrations, and operator upgrades.
  • High Availability & Disaster Recovery: Design and operate high availability within each region and cross-region disaster recovery — PostgreSQL replica clusters, backups and point-in-time recovery, Kafka MirrorMaker 2 — with defined RPO/RTO targets and regular failover drills.
  • Self-Service Platform: Give product teams self-service provisioning of databases, users, topics, and ACLs through code, with secure defaults and guardrails.
  • Observability & Operations: Build monitoring and alerting (Prometheus, Grafana, OpenTelemetry) for replication lag, backups, consumer lag, and capacity; participate in on-call and lead incident reviews.
  • Security & Multi-Tenancy: Implement TLS/mTLS, secrets management (Vault/OpenBao), network policies, Kafka ACLs, and per-tenant isolation; integrate database and Kafka access with our identity platform (Keycloak / OIDC).
  • Data Movement & CDC: Operate the outbox and change-data-capture pipelines (Debezium, Kafka Connect) that keep databases and event streams consistent between the global control plane and regions.
  • Automation & Tooling: Write automation and small services in Go or Python to remove manual work, and partner with application teams on schema standards and performance troubleshooting.

Qualifications

Must Have

  • Experience: 8+ years in DevOps, SRE, platform, or infrastructure engineering, including 3+ years running stateful systems (databases or message streaming) in production.
  • Kubernetes: Strong hands-on Kubernetes operations — StatefulSets, persistent volumes and storage classes, pod disruption budgets, scheduling and affinity, network policies, and cluster upgrades — on cloud and/or bare-metal clusters.
  • Operators & Helm: Production experience running PostgreSQL and/or Kafka through Kubernetes operators (CloudNativePG, Zalando, or Crunchy PGO; Strimzi, Confluent for Kubernetes, or similar), and authoring and maintaining Helm charts. Depth with one operator matters more than the specific product.
  • GitOps & IaC: Day-to-day use of Argo CD or Flux, Terraform, and CI/CD pipelines (GitHub Actions, GitLab CI, or similar) in a declarative, review-driven workflow.
  • PostgreSQL Operations: Practical PostgreSQL administration — replication and failover, backup and point-in-time recovery, connection pooling (PgBouncer), version upgrades, and performance troubleshooting.
  • Kafka Operations: Running Kafka in production — brokers, topics and partitions, replication, consumer groups, monitoring, and capacity planning.
  • Reliability & DR: Designing and testing high availability and disaster recovery for stateful systems, with defined RPO/RTO.
  • Automation: Strong scripting and programming in Go (preferred) or Python, plus Bash.

Nice to Have

  • Cross-region replication (CloudNativePG replica clusters, Kafka MirrorMaker 2) and global/regional multi-site architectures.
  • Debezium, Kafka Connect, and transactional outbox / CDC patterns.
  • Identity integration for data platforms (Keycloak, OIDC, LDAP/Active Directory).
  • Observability stacks (Prometheus, Grafana, OpenTelemetry, VictoriaMetrics).
  • Secrets and security tooling (Vault/OpenBao, cert-manager, mTLS).
  • Bare-metal infrastructure, Ceph or object storage, or GPU/AI infrastructure.
  • Contributions to CloudNativePG, Strimzi, or related open-source projects.
  • Experience in SOC 2 or ISO 27001 environments.

Education and Experience

Bachelor's degree in Computer Science or a related field, or equivalent practical experience.

Additional Information

What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

We are a Leader for Container Management in G2 (#2 after AWS)!

Similar Jobs

11 Minutes Ago
Remote or Hybrid
Senior level
Senior level
Information Technology
The Senior Project Manager leads complex technology projects for enterprise and government clients from initiation through closure. Responsibilities include managing scope, schedules, budgets, profitability, risks, contracts, vendors, resources, quality, compliance, and stakeholder communications. The role requires coordinating cross-functional teams, managing government contract requirements, overseeing project changes and deliverables, mentoring team members, and presenting updates to executives and clients. Travel up to 25% and occasional work outside standard hours may be required.
Top Skills: AgileFedrampFismaMS OfficeMicrosoft ProjectWaterfall
12 Minutes Ago
Remote or Hybrid
Mid level
Mid level
Information Technology
Support accounts receivable operations, including collections strategies, cash application, rebills, reconciliations, refunds, and credit management. The role focuses on optimizing cash flow, reducing bad debt and DSO, maintaining accounting and SOX controls, improving processes and systems, ensuring data accuracy, communicating with stakeholders and customers, and mentoring junior team members.
Top Skills: Enterprise Erp Systems
14 Minutes Ago
Remote or Hybrid
Expert/Leader
Expert/Leader
Information Technology
Owns enterprise AI security assessment, architecture, and control programs. The role evaluates AI systems, vendors, platforms, and shadow AI; designs secure architectures, guardrails, identity models, sandboxes, gateways, and runtime controls; and develops detection and response capabilities for prompt injection, jailbreaks, data leakage, and tool misuse. The engineer partners with governance, identity, platform, engineering, and security operations teams, advises senior leaders, drives remediation, and mentors security professionals.
Top Skills: Ai GatewaysAi GuardrailsAi ObservabilityAi SandboxesAnthropicAWSAzureAzure Ai ServicesData Loss PreventionGCPIdentity And Access ManagementIso/Iec 42001LangchainMcpMicrosoft Copilot StudioMitre AtlasNist Ai RmfOpenaiOwasp Top 10 For Llm ApplicationsSemantic KernelZero Trust

What you need to know about the Vancouver Tech Scene

Raincouver, Vancity, The Big Smoke — Vancouver is known by many names, and in recent years, it has gained a reputation as a growing hub for both tech and sustainability. Renowned for its natural beauty, the city has become a magnet for professionals eager to create environmental solutions, and with an emphasis on clean technology, renewable energy and environmental innovation, it's attracted companies across various industries, all working toward a shared goal: advancing clean technology.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account