Senior Infrastructure Engineer — OpenStack & Cloud Systems
Pubblicato il 05-09-2026 - Azienda Anonima in Palermo
pbSENIOR INFRASTRUCTURE ENGINEER - OPENSTACK CLOUD SYSTEMS /b /ppbr/ppWe are looking for a bSenior Infrastructure Engineer /b to serve as a technical leader within our Infrastructure Team, driving the design, evolution, and operation of the software-defined infrastructure that powers some of the largest HPC and AI clusters in Europe. OpenStack is the primary production platform today; the role owns the broader Linux, Kubernetes, and bare-metal systems stack around it. /ppWe hire first for bdepth of systems engineering /b — the ability to understand why something does not work, not just how to restart it: reading an strace, reading the source of a service, isolating a kernel or network issue, and driving a fix upstream when needed. A strong engineer with this foundation who knows OpenStack — or an equivalent large-scale production platform — is exactly who we are looking for; the specific stack can be learned, the way of reasoning cannot. /ppThe successful candidate will be responsible for the architecture and operations of cloud control planes in large-scale production environments, ensuring reliability, scalability, and operational continuity. They will lead critical technical initiatives, plan and execute infrastructure upgrades and migrations, collaborate with leading technology vendors to manage escalations, and provide technical leadership to the team through mentoring, knowledge sharing, and best practices. /ppbr/ppbKEY RESPONSIBILITIES: /b /pulliOwn the architecture and day-2 operations of OpenStack control planes for clusters of hundreds of nodes today, with a growth path to 1000+ nodes at upcoming public and private AI Factories: uptime, capacity, performance, and security KPIs. /liliDiagnose and resolve complex, cross-layer production issues down to the root cause — kernel, systemd, storage, and network stack — using tools such as strace, perf, and packet capture, reading service source code where needed and contributing fixes upstream. /liliDesign and operate the advanced networking underpinning high-throughput HPC and AI workloads (Neutron OVN / OVS, SR-IOV, VF-LAG, DPDK, BGP-EVPN), and drive technology and architecture decisions with senior team members and stakeholders. /liliCoordinate deployments and upgrades across geographically distributed sites, including cross-site data replication, federated identity, and disaster-recovery posture. /liliDevelop and maintain advanced Infrastructure-as-Code pipelines (bAnsible, OpenTofu / Terraform, Helm /b) and enforce gitops-style review for production change management. /liliProduce and maintain technical documentation,
including operational runbooks — step-by-step procedures with explicit go / no-go decision gates and rollback plans — for control-plane upgrades, security patching, and migrations, as well as RCA reports. /liliMentor mid-level and junior engineers on Linux and OpenStack internals, lifecycle operations, and production best practices; collaborate with the Presales team in designing systems end-to-end from hardware configuration to software stack. /li /ulpbr/ppbEDUCATION: /b /ppMaster’s degree or Ph.D. in Computer Science, Telecommunications Engineering, Network Engineering, or a related STEM field — or equivalent practical experience, including: /pullib5+ years /b of hands-on engineering experience in Linux systems, cloud architecture, network engineering, and complex production systems. /lilib3+ years /b of OpenStack production experience at scale (multi-site or strict SLA) — or equivalent experience operating a large-scale production platform, with the depth to become productive on OpenStack within the first months. /lilib2+ years /b of direct vendor escalation accountability in enterprise or hyperscaler environments. /li /ulpbr/ppbSKILLS AND COMPETENCES - CORE TECHNICAL SKILLS /b /ppbLinux systems engineering (deep) /b: /pullibExpert Linux sysadmin /b (Rocky / RHEL, Ubuntu, SLES); /lilibAdvanced troubleshooting /b strace, perf, ftrace / eBPF, gdb; comfortable reading the source of the services operated and contributing fixes upstream; /lilibContainer platforms /b (Docker, Podman, Singularity) and their runtime internals. /li /ulpbOpenStack cloud virtualization /b /pulliStrong, hands-on proficiency on Neutron, Nova, Ironic, Cinder, Keystone, Manila. /liliDirect hands-on experience with Kayobe + Kolla-Ansible is a strong plus. /li /ulpbAdvanced networking /b /pulliWorking knowledge of BGP / OSPF / EVPN / VXLAN, Spine-Leaf datacenter fabrics, SDN, OVN / OVS. /liliSR-IOV + VF-LAG, hardware offload, DPDK; Mellanox / low-latency networking. /li /ulpbInfrastructure as Code Automation /b /pulliProduction-quality automation in Bash, Python, or Go. /liliHands-on Ansible (modules / roles / collections); Terraform / OpenTofu modules; Helm. /li /ulpbContainer orchestration (a plus) /b /pulliProduction-grade Kubernetes (bare metal and over OpenStack), with focus on GPU-accelerated workloads; Cluster API, ArgoCD. /li /ulpbHPC ecosystem (a plus) /b /pulliFamiliarity with Slurm, MPI, GPU stacks,
and parallel computing workloads on cloud-managed infrastructure. /li /ulpbVersion Control GitOps /b /pulliSolid Git operations (branching, rebasing, conflict resolution); GitLab CI or equivalent — design and review of pipelines for production change management. /li /ulpbCommunication /b /pulliFluent English (working language with Mellanox / NVIDIA / Dell engineering); clear written and verbal communication for vendor, customer, and internal audiences. /li /ulpbr/ppbSKILLS AND COMPETENCES - SOFT SKILLS /b /pullibTechnical Leadership /b — Ability to guide technical decisions and mentor engineers. /lilibProblem Solving /b — Strong analytical skills to resolve complex, cross-layer infrastructure issues. /lilibCollaboration /b— Ability to work effectively with internal teams and external partners. /lilibOwnership /b— Strong sense of responsibility and accountability for systems and outcomes. /lilibDecision Making /b — Ability to act decisively in high-pressure, production-critical environments. /lilibAdaptability /b— Comfortable working in fast-evolving, complex infrastructure environments. /lilibKnowledge Sharing /b — Commitment to mentoring and promoting best practices within the team. /li /ulpbr/ppbWORKING ENVIRONMENT: /b /pullibRemote work — /bFlexibility to work from anywhere in Italy, with no hybrid mandate; on-site presence required for periodic group meetings and a few company events. /lilibObjective-based work — /bClear, measurable objectives reviewed and updated throughout the year, with a structured growth path and access to leading-edge HPC and AI infrastructure in production. /lilibOn-site missions — /bThe role involves periodic on-site engagements at customer and partner facilities in Italy and across the EU (cluster bring-up, vendor PoC, co-deployment) and attendance at national and international events, with travel covered by the company. /lilibRAL /b: 40.000,00 - 50.000,00 € /li /ulpbr/ppbGROWTH DEVELOPMENT /b /ppWhile this position is targeted at senior engineers, bwe also welcome applications from mid-level engineers /b with a solid systems-engineering foundation — strong Linux fundamentals and hands-on experience with OpenStack, Kubernetes, or large-scale distributed systems — looking to grow into full ownership of production cloud infrastructure. What matters most is depth of reasoning and a genuine curiosity for how systems work under the hood; the specific stack is something we help you master. /ppDepending on experience level, successful candidates will grow into increased responsibility through hands-on work on large-scale AI and HPC infrastructures, supported by structured mentorship and knowledge sharing within the team. /ppbr/p
