Mission & Impact
Key Architectural Challenges You Will Solve
01
Multi-Cloud Kubernetes
Deploy and maintain production AWS EKS and GCP GKE clusters with automated autoscaling and node optimization.
02
Declarative Terraform IaC
Manage 100% of cloud resources via version-controlled Terraform modules with zero manual console drift.
03
GitOps CI/CD Automation
Automate Canary and Blue-Green zero-downtime deployments with ArgoCD, Helm, and GitHub Actions.
04
Full-Stack Observability
Configure Prometheus, Grafana, Loki, and OpenTelemetry with custom SLI/SLO dashboards and instant PagerDuty alerts.
Stack & Tooling
Production Tech Ecosystem
Cloud Providers
AWS (EKS, VPC, RDS, S3, IAM)
Google Cloud Platform (GKE, BigQuery)
Azure
Cloudflare
Containers & Orchestration
Kubernetes (K8s)
Docker
Helm
ArgoCD
Istio Service Mesh
vLLM GPU Clusters
Infrastructure as Code
Terraform
OpenTofu
Terragrunt
Ansible
Packer
Monitoring & SRE
Prometheus
Grafana
Loki
Jaeger / OpenTelemetry
Datadog
PagerDuty
Responsibilities
What You Will Own & Deliver
- Architect, provision, and maintain production-grade Kubernetes clusters across AWS EKS and Google Cloud GKE.
- Author modular, reusable Infrastructure as Code (IaC) using Terraform following least-privilege security and VPC isolation.
- Build automated zero-downtime deployment pipelines using ArgoCD, GitHub Actions, and container image vulnerability scanners.
- Set up comprehensive observability stacks (Prometheus, Grafana, Loki) with custom latency, error rate, and saturation alerts.
- Lead incident response reviews, post-mortem analyses, and disaster recovery replication failover tests.
- Collaborate with engineering leads on cloud cost governance, AWS Graviton4 migrations, and FinOps budget optimization.
Qualifications
Required Experience & DNA
- 5+ years of DevOps / SRE experience managing high-traffic cloud infrastructure in enterprise environments.
- Deep expertise in Kubernetes administration, networking (CNI), ingress controllers (Nginx/Traefik), and security policies.
- Advanced proficiency in Terraform, infrastructure state management, and multi-environment module design.
- Strong Linux systems engineering background with proficiency in Bash and Python for infrastructure automation.
- Hands-on experience setting up Prometheus metrics collection, Grafana alerting rules, and centralized log aggregation.
High-Value Bonus Signals
- Official AWS Certified Solutions Architect (Professional) or Certified Kubernetes Administrator (CKA).
- Experience managing GPU compute clusters (NVIDIA CUDA / vLLM) for LLM inference workloads.
- Experience with FinOps cloud cost reduction tooling (Kubecost, Infracost).
Rapid Hiring Protocol
Our Transparent 4-Stage Interview Loop
01
Discovery Call
20 min culture and compensation alignment.
02
Architecture Review
45 min real-world systems deep dive. No trivia.
03
Founder Sync
30 min roadmap and equity alignment.
04
Offer in 48h
Formal offer letter & hardware delivery.
DIRECT FOUNDER INBOX
Apply for Lead DevOps & Cloud Platform Architect
Fill out this short application. Our engineering squad leads review all submissions directly and respond within 24 business hours.