We’re adding a hands-on DevOps Engineer to own and improve the infrastructure, deployment, and operational foundations of a growing software platform. This is an ownership-driven role: you’ll work closely with engineering and delivery, turn operational needs into reliable systems, and make sure teams can ship safely and predictably.
You’ll be responsible for the full operational lifecycle — infrastructure as code, CI/CD, cloud environments, observability, security, reliability, and incident response. The environment is centred on Google Cloud Platform and its supporting services, so we need someone who is comfortable working from incomplete requirements, making sensible technical decisions, and improving systems incrementally without creating unnecessary complexity
Requirements * 5+ years of hands-on DevOps, Platform Engineering, SRE, or cloud infrastructure experience. * Strong Linux administration and production troubleshooting skills. * Practical production experience with Google Cloud Platform. * Strong hands-on experience with Terraform. * Hands-on experience building and maintaining CI/CD pipelines with GitHub Actions. * Production experience operating containerised workloads on Google Cloud Platform. * Experience implementing monitoring, logging, alerting, and operational dashboards for workloads running on Google Cloud Platform. * Solid understanding of networking fundamentals, DNS, TLS, load balancing, ingress, firewalls, and private connectivity. * Experience with secrets management, identity and access controls, backups, and disaster-recovery practices. * Ability to work independently, make practical decisions from partial requirements, and communicate operational risks clearly
Responsibilities * Own cloud infrastructure and runtime environments across development, staging, UAT, and production. * Build and maintain infrastructure as code using Terraform. * Design, maintain, and improve CI/CD pipelines so application changes can be tested, deployed, and rolled back safely. * Operate and support containerised application workloads on Google Cloud Platform. * Implement monitoring, logging, alerting, and service-level indicators that give the team clear visibility into system health. * Improve platform reliability, availability, scalability, performance, and cost efficiency. * Establish secure access, secrets management, backup, recovery, and environment-management practices. * Lead operational troubleshooting and incident response; document root causes and turn incidents into concrete improvements. * Partner with developers to improve deployability, observability, and operational readiness from design through production. * Treat infrastructure and internal tooling as products — reliable, maintainable, documented, and easy for engineers to use.
Would be a plus * Experience building or operating internal developer platforms and self-service deployment workflows. * Experience designing GitHub Actions pipelines and managing GitOps deployment workflows with Flux. * Experience with GCP managed application services, container registries, IAM, and storage buckets. * Experience managing Cloudflare CDN, Tunnel, and DNS configurations. * Experience defining service-level objectives, error budgets, capacity plans, and reliability metrics. * Experience improving cloud cost visibility and reducing unnecessary infrastructure spend. * Security or compliance experience in regulated environments. * Software development or scripting experience in Python, Go, or another general-purpose language