We’re looking for an experienced Infrastructure Engineer to work on a project of our Customers. They are building a multi-node GPU platform for large-scale model training and inference using managed GPU providers. In this role, you’ll have to own the platform performance, reliability, and the technical relationship with the Client’s providers.
Our Customers provide SaaS solutions that help companies to optimize their business. These solutions include business planning to automate and optimize business, delivery, and workflow solutions. The platform leverages industry-leading Artificial Intelligence (AI) and Machine Learning (ML) for better prediction and prevention of disruptions across business. Responsibilities: * Design and operate distributed GPU training infrastructure * Validate cluster topology, RDMA/InfiniBand, and NCCL performance * Standardize Kubernetes or Slurm scheduling, GPU images, and software versions * Diagnose issues across training workloads, networking, storage, and GPU hosts * Build monitoring, benchmarks, runbooks, and reliability standards * Work with GPU providers to resolve incidents and define technical requirements * Support SFT, DPO, RL, and large-scale inference workloads
Requirements: * Production experience with multi-node GPU training infrastructure * Strong Linux, containers, CUDA, and NVIDIA-GPU-stack knowledge * Hands-on experience with NCCL and InfiniBand or RoCE/RDMA troubleshooting * Deep experience with Kubernetes or Slurm * Infrastructure automation and observability experience * Strong skills in incident leadership and provider-facing communication * English level — Upper-Intermediate or higher
Will be a plus: * Experience in an AI lab, HPC environment, or specialist GPU cloud * Knowledge of distributed-training frameworks such as PyTorch, Megatron, or DeepSpeed * Experience with parallel storage, checkpoint optimization, and multi-provider platforms
We offer: * Remote-first work model with flexible working hours (we provide all equipment) * Comfortable and fully equipped offices in Lviv and Rzeszów * Competitive compensation with regular performance reviews * Referral Bonus Program * 18 paid vacation days per year * 12 days of paid sick leave per year * Extra paid leave (blood donation, marriage, childbirth, etc.) * Health & wellness support: either a monthly budget for medical insurance and sports activities, or a full medical insurance plan, depending on your cooperation model * Mental Health Program — company-covered psychologist consultations, with an annual coverage limit * English, German, and Polish language courses and Speaking Clubs * Corporate merch gifts for your special dates * Corporate subscription to learning platforms, regular meetups and webinars * Friendly team that values accountability, innovation, teamwork, and customer satisfaction * Maternity leave * Inclusive environment where everyone feels valued and treated equally. We proudly partner with VeteranHub to support Ukrainian veterans * We are committed to supporting Ukraine and actively participate in charity initiatives