Back to All Open Roles
CLOUD INFRASTRUCTURE & SRE JOB ID: IEC-SRE-2026 2 Openings Available

Lead DevOps & Cloud Platform Architect

Architect Multi-Cloud Kubernetes Meshes, Terraform GitOps & 99.99% SRE Telemetry

We are looking for a Lead DevOps & Cloud Platform Architect to manage our multi-cloud infrastructure spanning AWS, Google Cloud, and private GPU clusters. You will write modular Terraform code, automate zero-downtime ArgoCD GitOps pipelines, and ensure 99.99% uptime with sub-minute incident response.

ROLE SNAPSHOT

Join This Squad

  • Work Model: 100% Remote Async
  • Direct Review: Founders & Tech Lead
100% Confidential • Fast 48h Response
Mission & Impact

Key Architectural Challenges You Will Solve

01

Multi-Cloud Kubernetes

Deploy and maintain production AWS EKS and GCP GKE clusters with automated autoscaling and node optimization.

02

Declarative Terraform IaC

Manage 100% of cloud resources via version-controlled Terraform modules with zero manual console drift.

03

GitOps CI/CD Automation

Automate Canary and Blue-Green zero-downtime deployments with ArgoCD, Helm, and GitHub Actions.

04

Full-Stack Observability

Configure Prometheus, Grafana, Loki, and OpenTelemetry with custom SLI/SLO dashboards and instant PagerDuty alerts.

Stack & Tooling

Production Tech Ecosystem

Cloud Providers
AWS (EKS, VPC, RDS, S3, IAM) Google Cloud Platform (GKE, BigQuery) Azure Cloudflare
Containers & Orchestration
Kubernetes (K8s) Docker Helm ArgoCD Istio Service Mesh vLLM GPU Clusters
Infrastructure as Code
Terraform OpenTofu Terragrunt Ansible Packer
Monitoring & SRE
Prometheus Grafana Loki Jaeger / OpenTelemetry Datadog PagerDuty
Responsibilities

What You Will Own & Deliver

  • Architect, provision, and maintain production-grade Kubernetes clusters across AWS EKS and Google Cloud GKE.
  • Author modular, reusable Infrastructure as Code (IaC) using Terraform following least-privilege security and VPC isolation.
  • Build automated zero-downtime deployment pipelines using ArgoCD, GitHub Actions, and container image vulnerability scanners.
  • Set up comprehensive observability stacks (Prometheus, Grafana, Loki) with custom latency, error rate, and saturation alerts.
  • Lead incident response reviews, post-mortem analyses, and disaster recovery replication failover tests.
  • Collaborate with engineering leads on cloud cost governance, AWS Graviton4 migrations, and FinOps budget optimization.
Qualifications

Required Experience & DNA

  • 5+ years of DevOps / SRE experience managing high-traffic cloud infrastructure in enterprise environments.
  • Deep expertise in Kubernetes administration, networking (CNI), ingress controllers (Nginx/Traefik), and security policies.
  • Advanced proficiency in Terraform, infrastructure state management, and multi-environment module design.
  • Strong Linux systems engineering background with proficiency in Bash and Python for infrastructure automation.
  • Hands-on experience setting up Prometheus metrics collection, Grafana alerting rules, and centralized log aggregation.

High-Value Bonus Signals

  • Official AWS Certified Solutions Architect (Professional) or Certified Kubernetes Administrator (CKA).
  • Experience managing GPU compute clusters (NVIDIA CUDA / vLLM) for LLM inference workloads.
  • Experience with FinOps cloud cost reduction tooling (Kubecost, Infracost).
Rapid Hiring Protocol

Our Transparent 4-Stage Interview Loop

01
Discovery Call
20 min culture and compensation alignment.
02
Architecture Review
45 min real-world systems deep dive. No trivia.
03
Founder Sync
30 min roadmap and equity alignment.
04
Offer in 48h
Formal offer letter & hardware delivery.
DIRECT FOUNDER INBOX

Apply for Lead DevOps & Cloud Platform Architect

Fill out this short application. Our engineering squad leads review all submissions directly and respond within 24 business hours.