AI Platform Engineer
Pubblicato il 14-08-2026 - Reply in Italia
Are you an AI Platform Engineer expert in the platform that turns a researcher's idea into a running job on a cluster?
Join Reply!
WHO WE ARE
Reply specialises in the design and implementation of solutions based on new communication channels and digital media. As a network of highly specialised companies, Reply supports major industrial groups in the telecom and media; industry and services; banking and insurance and public sectors in defining and developing business models enabled by the new paradigms of AI, cloud computing, digital media and the internet of things. Reply's services include: consulting, system integration and digital services.
WHAT WILL YOU DO?
- Core activities. You will keep the core of the Model Factory platform standing, and you will make it grow. You will orchestrate GPU workloads and distributed training that hold up under load, build versioning and governance across data, models and evaluation runs so results are still reproducible months later, and shorten the distance between a researcher's idea and a running job on the cluster. Every model we train, distil or optimise — for our own portfolio and for client engagements — runs on what you build.
- Tech & Tools Stack. You will work Sky Pilot, Dagster, NVIDIA NeMo, Kubernetes, MLflow and Unity Catalog, the core of the platform. Around them: the NVIDIA stack (CUDA, NCCL), job scheduling across national (Italian) and European HPC clusters and cloud (such as Scaleway or Nebius), infrastructure as code and Git Ops CI/CD, vLLM for inference serving, and Prometheus and Grafana observability that tells us what a run cost and why it failed.
- Team work. You will work alongside the AI research engineers who train the models, the model design engineers who specify them, the project managers who commit the dates, and cloud and security specialists — plus client platform teams when the environment lands on their side. Your first customers are internal: if the platform is slow, opaque or fragile,
they feel it long before any client does.
WE'LL TOTALLY LOVE YOU IF YOU HAVE…
- Academic background. Degree in Computer Engineering, Computer Science or a related technical field. Solid fundamentals in distributed systems, networking matter more to us than any specific coursework.
- Technical & strategic skills. Kubernetes and distributed training that genuinely work in your hands: you can debug a training job that hangs, size a GPU allocation, and tell a scheduling problem from a networking one. Strong Python, and infrastructure as code as a default rather than an afterthought. Having run ML platforms at scale is required — we care that the fundamentals are real and that you design for reproducibility, multi-tenancy and cost attribution from day one. The position is open to varying levels of expertise and seniority.
- Nice to have. Experience with Sky Pilot, Dagster or NVIDIA NeMo in production; HPC schedulers (Slurm) and Infini Band/RDMA networking; LLM inference optimisation (quantisation, continuous batching, KV-cache); Unity Catalog or Databricks governance; open-source contributions to the ML infrastructure ecosystem; sovereign or on-premise AI environments.
- Soft skills. You should be pragmatic and ownership-driven, comfortable being the person everyone else depends on. Clear written communication, a bias for automation over heroics, and the patience to make someone else's workflow actually work.
WHAT WE OFFER
- An offer tailored to your experience. This position is open to people with varying levels of expertise and seniority.
The compensation package will be determined based on your professional background, technical skills, expertise, and the level of responsibility associated with the role. The collective labour agreement (CCNL) applied is the Italian Metalworking Industry Agreement. The job classification will be assessed starting from level B2, and the gross annual salary (RAL) will range from €50,000 to €75,000. Previous experience with cutting-edge technologies such as GPU cluster orchestration, distributed training on national and European HPC infrastructure, Sky Pilot, Dagster and NVIDIA NeMo, LLM inference optimisation, and data, model and evaluation governance in MLflow and Unity Catalog will be considered a strong asset.
- A structured career path. Our Career Path offers opportunities to grow both as a technical specialist and as a future leader. Your ambitions, the skills you develop, and the results you achieve will shape your professional journey.
- Continuous learning. Technology evolves rapidly - and so do we. You will join an environment that encourages curiosity, continuous learning, and the exploration of new ideas and emerging technologies.
- The benefits of being a Replyer. You will also have access to the benefits and initiatives dedicated to our community.
WHAT ARE THE NEXT STEPS
The first step of our recruiting process will be the meetings with the technical referents and then a face to face interview with the HR team. We care about an equal recruiting process.
Feel interested?
Reply is committed to embracing diversity and creating an inclusive work environment by valuing the uniqueness of people regardless of age, gender, sexual orientation, religion, nationality, or disabilities as protected by Italian Law (L.68/99).
Furthermore, Reply is committed to ensuring a fair and accessible selection process: to help you during the recruitment process, please let us know of any kind of support you may need.
