06 ago - Ro
Ntt Data Europe & Latam
ph3Who We Are /h3 pWe are looking for a hands‑on bSite Reliability Engineer (SRE) /b to help improve, scale, and operationalize an internal platform that enables engineering teams to ship faster and safer through reliable, automated, and resilient engineering workflows. /p pYou will work on platform reliability, observability, automation, incident management, CI/CD workflows, GitHub‑based engineering automation, and developer tooling that helps teams consistently adopt operational best practices across repositories and services. /p h3What You’ll Be Doing /h3 ul liDesign, implement and maintain monitoring, alerting, and observability solutions to ensure platform reliability and performance /li liDevelop and improve automation for infrastructure provisioning, deployment pipelines, and operational processes /li liAdminister and optimize GitHub Enterprise environments, including repository management, access controls, branch protection policies, and enterprise‑wide standards /li liPartner with engineering teams to define and measure Service Level Objectives (SLOs), Service Level Indicators (SLIs), and reliability metrics /li liInvestigate production incidents, perform root cause analysis,
and drive post‑incident improvements to prevent recurrence /li liImprove system resilience, scalability, and availability through proactive reliability engineering practices /li liBuild and maintain GitHub‑based automation, CI/CD pipelines, and developer self‑service capabilities /li /ul h3What You’ll Bring Along /h3 ul liBSc/MSc in Computer Science or related field /li liMinimum 6+ years as a SRE /li liStrong experience with GitHub Enterprise (repos, orgs, actions, integrations) /li liAdvanced hands‑on knowledge of Terraform (IaC) /li liExperience building and running CI/CD pipelines (GitHub Actions ideally) /li liSolid understanding of cloud platforms (Azure/AWS/GCP) and integrations /li liExperience with identity and access management (RBAC, token/app auth models) /li liKnowledge of backup, recovery, and cyber resilience principles (Cybervault) /li liExperience with security best practices and software supply chain risks (due to recent events) /li liFocus on improving developer experience and self‑service capabilities /li liExcellent problem‑solving and root cause analysis skills /li /ul /p #J-18808-Ljbffr
08 ago - Italia
Altro
08 ago - Italia
WAICO
08 ago - Italia
Altro
08 ago - Italia
Altro