We are looking for a Senior DevOps Engineer to take ownership of cloud infrastructure, CI/CD, security, monitoring, and deployment architecture for an enterprise AI platform.
You will work closely with the CTO and engineering team to build and scale reliable infrastructure supporting AI-powered applications and agentic systems.
This is a hands-on role with significant ownership. You will help define infrastructure standards, improve deployment processes, and contribute to building the DevOps/Infrastructure function as the engineering organization grows.
Responsibilities * Own and develop GCP-based cloud infrastructure, including Cloud Run, GKE, IAM, DNS, VPCs, and Load Balancers. * Design and maintain scalable CI/CD pipelines supporting multiple applications and environments. * Build and manage containerized and serverless production environments. * Implement infrastructure automation using Terraform, Helm, and related Infrastructure-as-Code tools. * Manage Kubernetes environments and deployment processes. * Implement and maintain monitoring, logging, and alerting using tools such as Prometheus, Grafana, ELK, and Google Cloud Monitoring. * Improve infrastructure reliability, scalability, performance, and availability. * Implement security practices including access controls, secrets/key management, IAM, and environment isolation. * Establish DevOps and infrastructure standards across engineering teams. * Work directly with engineering leadership on infrastructure architecture and technical decisions. * Mentor engineers and help grow the Infrastructure/DevOps team. * Use modern AI development tools to improve engineering productivity and automation.
Requirements * 5+ years of experience in DevOps, SRE, Cloud, or Infrastructure Engineering. * At least 3 years in a technical leadership or infrastructure ownership role. * Strong commercial experience with Google Cloud Platform (GCP). * Strong experience with Kubernetes and containerized environments. * Hands-on experience with Terraform and/or Helm. * Strong knowledge of cloud-native and serverless architectures. * Experience building and maintaining CI/CD pipelines using tools such as GitHub Actions and ArgoCD. * Experience with monitoring and observability platforms such as Prometheus, Grafana, ELK, or equivalent. * Strong understanding of cloud networking, IAM, security, and access management. * Experience building infrastructure from scratch or in zero-to-one environments. * Comfortable working in a fast-moving environment with changing requirements. * Strong ownership mindset and ability to work independently.
Nice to Have * Experience working with AI/ML or Generative AI platforms. * Familiarity with agentic AI workflows. * Experience supporting RAG pipelines and vector databases. * Familiarity with frameworks or architectures similar to LangChain. * Experience supporting infrastructure for enterprise-scale AI applications. * Previous experience in an early-stage or rapidly scaling technology company.
What We Offer * Long-term opportunity with a fast-growing technology company. * Direct collaboration with senior engineering leadership. * High level of technical ownership and influence over infrastructure architecture. * Opportunity to build and improve infrastructure from an early stage. * Work with modern cloud, DevOps, and AI technologies. * Continuous learning and professional development opportunities.