Our client is a leading global data center provider delivering hyperscale and edge infrastructure solutions across the Americas, EMEA, and Asia-Pacific. With 80+ data centers in 20+ countries, they partner with industry leaders such as Google, Oracle, NVIDIA, and Microsoft Azure to power the world’s digital infrastructure. Recognized as a USA TODAY Top Workplace for four consecutive years, the company continues to expand its global footprint and customer ecosystem. Project Description The project focuses on building a cloud-based data platform for processing high-volume sensor and telemetry data. The specialist will help develop and operate real-time and analytical data capabilities that support dashboards, applications, and downstream data consumers, while contributing to the platform’s migration and evolution in AWS. Technologies * Apache Kafka, Apache Flink * Java, Python, SQL * Apache Iceberg / Delta Lake / Hudi * Parquet, Avro, Schema Registry * Amazon Athena / Spark SQL * ClickHouse / Druid / Pinot * SQL Server / PostgreSQL * Kubernetes (Amazon EKS) * Helm, Argo CD, GitOps * Terraform, GitHub Actions * AWS (Amazon MSK, EMR Serverless, MWAA, S3, Athena, IAM) * LLM, Embeddings, Vector Search
What You’ll Do * Design, build, and operate real-time streaming pipelines with Apache Kafka (Amazon MSK) and Apache Flink for high-throughput sensor and telemetry data; * Define and manage streaming data contracts, including Avro schemas and schema evolution through a schema registry; * Build and maintain analytical serving layers using ClickHouse or similar columnar OLAP databases, and develop REST APIs for dashboards, applications, and downstream teams; * Develop and operate batch and scheduled data workflows with Apache Airflow (Amazon MWAA); * Build and operate a data lakehouse based on Apache Iceberg and Amazon S3, using PySpark on EMR Serverless and Amazon Athena; * Deploy and operate containerized data workloads on Kubernetes (Amazon EKS) using Argo CD, Helm, and GitOps practices; * Manage cloud infrastructure as code with Terraform and support CI/CD automation with GitHub Actions; * Monitor production data platforms with Prometheus and Grafana, troubleshoot issues, and participate in incident response; * Contribute to cloud migration initiatives by porting data pipelines from existing platforms, including Azure-based data platforms, to AWS; * Develop production-grade Java, Python, and SQL code with automated testing; * Use AI-assisted development tools to accelerate analysis and implementation while maintaining code quality and architectural integrity; * Keep technical documentation and operational runbooks current;
Job Requirements * 7+ years of experience in data or software engineering, including production experience with streaming systems; * Ability to work independently in ambiguous and fast-changing environments; * Deep hands-on experience with Apache Kafka and Apache Flink or an equivalent stream-processing framework, including Java; * Strong Python and SQL skills across transactional and analytical databases; * Experience with Apache Iceberg, Delta Lake, or Hudi; Parquet, Avro, schema registries, Athena, or Spark SQL; * Production experience with ClickHouse, Druid, Pinot, or similar analytical databases, as well as SQL Server or PostgreSQL; * Experience with Kubernetes, Amazon EKS, Helm, Argo CD, and GitOps; * Experience with Terraform and GitHub Actions; * Strong knowledge of AWS services, including MSK, EMR Serverless, MWAA, S3, Athena, and IAM; * Knowledge of ML fundamentals, including feature engineering, model training and evaluation, and ML data requirements; * Familiarity with LLMs, prompt-based workflows, embeddings, vector search, anomaly detection, and forecasting; * Experience with observability, alerting, troubleshooting, and incident response; * Ability to communicate technical decisions clearly to technical and non-technical stakeholders;
Nice to Have * Working knowledge of Azure Event Hubs, Data Factory, and Synapse; * Experience with OT or industrial telemetry, including OPC UA, BMS/EPMS, or time-series sensor data; * Experience with cloud-to-cloud or on-premises-to-cloud migrations; * Experience with data center or other critical-infrastructure operations; * Familiarity with data quality frameworks, data contracts, and metadata management;
What Do We Offer The global benefits package includes: * Technical and non-technical training for professional and personal growth; * Internal conferences and meetups to learn from industry experts; * Support and mentorship from an experienced employee to help you professional grow and development; * Health insurance; * English courses; * Sports activities to promote a healthy lifestyle; * Flexible work options, including remote and hybrid opportunities; * Referral program for bringing in new talent; * Work anniversary program and additional vacation days. Please, note, that we will consider all the applications with due respect, but only shortlisted candidates will be contacted