Senior DevOps Engineer
Pubblicato il 05-08-2026 - EnerSys in Arezzo
ppWe are looking for a senior, hands-on platform engineer to own the reliability, security and operational maturity of our Azure and Kubernetes platform. You will work within our platform team, in partnership with the DevOps Team Lead, architecture, engineering, cybersecurity and data/ML teams, to build secure-by-default cloud infrastructure and support the safe operation of emerging AI workloads. /ppIn this role, you will directly influence uptime, performance, cloud security posture and customer experience on Azure. You will also help define how the organization safely runs AI-enabled workloads in production, while evolving the engineering culture toward reliability-first, security-by-default, cloud-native delivery. /ph3Essential Duties and Responsibilities /h3p1. Azure, AKS and platform reliability /pulliDesign, build and operate secure, highly available Azure cloud architecture for microservices, event-driven systems and AI-related workloads. /liliDefine and improve SLIs, SLOs, error budgets, resilience patterns and operational standards for production services. /li /ulp2. Kubernetes platform engineering and automation /pulliOperate production and non-production AKS clusters, including node pools, upgrades, autoscaling, RBAC, network policies, resource quotas and isolation patterns. /liliAutomate provisioning and lifecycle management using Infrastructure as Code, GitOps, Helm and Kustomize. /li /ulp3. Platform security and DevSecOps /pulliOwn Azure and AKS security posture across identity, networking, secrets, policy, vulnerability management and software supply chain controls. /liliEmbed security into CI/CD pipelines through scanning, policy-as-code, deployment gates, audit logging and automated response pattern. /li /ulp4. API gateways, ingress and service traffic /pulliDesign and operate Kubernetes API gateway and ingress layers, ideally using Apache APISIX or a similar platform such as Kong. /liliImplement routing, load balancing, rate limiting, OAuth2/OIDC, JWT, mTLS, traffic controls and API lifecycle standards. /li /ulp5. Observability,
incidents and operational excellence /pulliBuild meaningful metrics, logs, traces, dashboards and alerts using Azure Monitor, Application Insights, Log Analytics, Prometheus, Grafana and OpenTelemetry. /liliLead or support incident response, root cause analysis, runbooks, reliability improvements, chaos testing and on-call readiness. /li /ulp6. Safe AI/LLM workload operations on Kubernetes /pulliSupport secure deployment and governance of AI, LLM or agentic workloads using controlled identities, network isolation, auditing, resource limits and approval gates. /liliPartner with data/ML and security teams to reduce risks such as prompt injection, tool abuse, data exfiltration and uncontrolled compute spend. /li /ulp7. Technical leadership and enablement /pulliPartner with engineering teams to define platform standards, coach teams on secure-by-default practices and improve developer self-service. /liliCreate and maintain documentation, runbooks, threat models and reusable automation that reduce operational toil. /li /ulh3Required skills and experience /h3h3Must-haves /h3ulli5+ years in DevOps, SRE, Platform, Cloud or Infrastructure Engineering roles with significant production ownership. /liliDeep hands-on Azure experience, especially AKS, Azure Networking, Entra ID, Azure Policy, Azure Monitor, Application Insights, Log Analytics and Azure DevOps. /liliStrong AKS and Kubernetes experience, including cluster and node pool design, CNI/networking, private clusters, DNS, ingress, RBAC, Managed Identities and autoscaling. /liliStrong infrastructure security background on Azure: Defender for Cloud, Microsoft Sentinel or SIEM/SOAR, Zero Trust, least privilege, network hardening,
secrets management and supply chain security. /liliInfrastructure as Code and configuration management experience with Terraform, Bicep or Pulumi, plus Helm and/or Kustomize. /liliCI/CD and DevSecOps experience, ideally with Azure DevOps Pipelines, including SAST/DAST, dependency or image scanning, IaC scanning, secret detection and deployment gates. /liliObservability and incident management experience across metrics, logs, traces, alerting, SLOs/SLIs, root cause analysis and durable remediation. /liliStrong scripting or programming ability in Python, Bash, PowerShell and/or Go. /liliPractical experience with Kubernetes API gateways or ingress platforms, such as Apache APISIX, Kong or similar. /li /ulh3Nice-to-haves /h3ulliExperience deploying or governing AI, LLM or agentic workloads on Kubernetes, including inference serving, tool gateways or MCP servers, guardrails, sandboxing, agent identity and behavioral observability. /liliRelevant certifications such as CKA, CKS, AZ-104, AZ-400, AZ-500, AZ-305, SC-100 or equivalent. /liliService mesh experience with Istio, Linkerd or Cilium for traffic management, observability and mTLS. /liliGitOps experience with Argo CD or Flux. /liliExperience with event-driven systems such as Kafka, Azure Event Hubs, Azure Service Bus or RabbitMQ. /liliExperience with multi-cluster or hybrid-cloud Kubernetes, chaos engineering, resilience testing or FinOps/cost governance on Azure. /liliPerformance tuning experience for high-throughput APIs, event-processing platforms or Kubernetes workloads. /li /ulh3What will help you succeed /h3ulliA calm, structured approach during high-pressure incidents, including security events. /liliA security-first mindset with the ability to balance reliability, delivery speed and risk. /liliClear communication skills and the ability to simplify complex technical topics for different audiences. /liliA collaborative, ownership-driven style: you work well in a platform team, share knowledge openly and can take a problem from architecture through production operation. /li /ul /p #J-18808-Ljbffr
