22 set - Piemonte
AI4I Foundation
The Italian Institute of Artificial Intelligence (AI4I) is seeking a senior, hands-on Cloud / DevOps Engineer to design, build, and operate the cloud foundation of its HPC/AI infrastructure.You will own the Cloud layer that enables AI and HPC workloads to run reliably at scale, with a strong focus on OpenStack-based private cloud infrastructure and the evolution toward container-based environments such as Kubernetes, Rancher, and Harvester.This role is central to AI4I’s industrial deployment mission: you will create and operate the infrastructure that powers real AI projects for enterprises, public institutions, and strategic partners.Location: AI4I, OGR – Turin, ItalyQualsiasi informazione aggiuntiva su questo lavoro è disponibile nel testo sottostante. Si assicuri di leggere attentamente, quindi invii la sua candidatura.Hybrid work: Flexible arrangements may be negotiatedThe position will remain open until filled and multiple candidates may be hired.About the RoleAs Cloud / DevOps Engineer at AI4I, you will take ownership of the end-to-end cloud infrastructure, from architecture and automation to day-to-day operations and continuous improvement.Beyond compute orchestration, this role includes responsibility for the design and operational management of distributed storage systems that support AI and HPC workloads. You will ensure that storage performance, durability, and scalability meet the needs of GPU-intensive training, fine-tuning, and inference environments.This is a cross-unit role supporting multiple AI4I teams and projects, working closely with engineering, deployment, and compute specialists.You will work closely with:AI4I Deployment and Engineering teams, enabling production-grade AI servicesHPC / AI engineers,
integrating cloud and compute environmentsExternal technology partners and vendorsInternal stakeholders delivering industrial AI projectsThis is a strongly execution-oriented role, combining Cloud engineering, DevOps practices, storage architecture, and operational responsibility. You will help build a robust, scalable, and secure infrastructure that supports both current deployments and future growth.Key ResponsibilitiesDesign, deploy, and operate AI4I’s private cloud infrastructure, with strong ownership of OpenStack environmentsLead the evolution toward container-based environments leveraging tools such as Kubernetes, Rancher, and Harvester, enabling Container-as-a-Service capabilitiesDesign, deploy, and manage distributed and software-defined storage systems supporting HPC and AI workloads, ensuring high-performance block, object, and file services integrated with GPU clustersOptimize storage performance, data durability, replication strategies, and overall resource utilization for compute-intensive workloadsImplement infrastructure automation and CI/CD practices (Infrastructure-as-Code) for reliable and reproducible operationsDefine and enforce operational standards, including monitoring, alerting, backup, disaster recovery, and incident responseSupport internal engineering teams by providing reliable infrastructure, documentation,
and best practicesContribute to infrastructure architecture decisions and long-term evolution of the AI4I cloudRequired QualificationsStrong hands-on experience operating and troubleshooting OpenStack production environmentsProven experience managing container orchestration environments (e.G., Kubernetes) in production settingsSolid hands-on experience with distributed and software-defined storage systems in HPC or cloud environmentsStrong Linux system administration and networking fundamentalsExperience automating infrastructure using Infrastructure-as-Code and CI/CD practicesDemonstrated experience operating mission-critical production services with uptime, reliability, and incident response responsibilityAdditional Strengths (from candidate profile)Experience with Rancher or HarvesterExperience integrating cloud environments with GPU/HPC workloadsExperience xysqume operating multi-tenant cloud environmentsExperience with monitoring and observability stacks (Prometheus, Grafana, ELK, etc.)Security hardening and identity management in private cloud environmentsExperience supporting internal engineering teamsKey Performance MetricsInfrastructure availability and reliabilityMean time to detect and resolve incidentsTime required to onboard new internal projects or usersResource utilization efficiency of the infrastructureWhat We OfferA collaborative environment with engineers and researchers working on real industrial AI deploymentsDirect impact: your infrastructure will run daily AI workloads and production systemsAn office at the epicenter of tech: OGR Torino technology hubCompetitive compensation and access to advanced computing infrastructureHow to ApplySubmit your application exclusively through the online form:Cover letter (max. 1 page) describing how your profile fits this specific positionCV and optional links to technical projects or operational experienceAI4I The Italian Institute of Artificial Intelligence (AI4I)Corso Castelfidardo 22, 10129 TorinoCodice fiscale 97904430010Partita IVA: 13130030011#J-18808-Ljbffr
23 set - Bologna
Gruppo Hera
23 set - Milano
Hitachi Energy
23 set - Treviso
Sinelec
23 set - Milano
Avanade