11 ago - Italia
ToolsGroup
ph3About Us /h3 pWe are a dynamic, rapidly growing global company and the innovators of service-driven supply chain planning software. We help companies make better, faster supply chain decisions that reduce inventory, improve customer satisfaction, and deliver powerful financial results amid increasing complexity, product proliferation, and uncertainty. /p pOur solutions have been recognized by customers globally and analyst firms, such as Gartner, for our ability to support service and inventory trade-offs, while dramatically improving planner productivity. ToolsGroup has been successfully deployed worldwide in more than 44 countries, and we have one of the highest customer retention rates in our industry. /p h3About the Role /h3 pWe are looking for an experienced Site Reliability Engineer who can also lead IT service operations. You will lead a small IT/Ops team,remainthe senior technical escalation point for production services, and act as the operational interface between IT/Ops, Engineering, Product, Security, Support, and business teams. /p pThis is not a coordination-only service management role or a generalist infrastructure position. You will diagnose distributed-system failures using logs, metrics, traces, commands, and platform tooling; make safe recovery decisions; automate recurring work; and engineer lasting reliability improvements. /p h3Main Responsibilities /h3 ul liLead major incidents from impact assessment and containment through recovery, stakeholder communication, root-cause analysis, and corrective actions. /li liTroubleshoot complex issues across Windows and Linux systems, Kubernetes and container workloads, hybrid networking and DNS, cloud infrastructure, identity, authentication, databases, storage, APIs, and service dependencies. /li liOperate and improve Azure, OCI, or comparable cloud environments, including monitoring, access controls, backup and recovery, reliability, and cost-aware scaling. /li liDefine and improve service-level indicators andobjectives, observability, alert quality, capacity, resilience, dependency mapping, and recovery readiness for critical services. /li liAutomate operational tasks and controls using PowerShell, Python, infrastructure as code, or CI/CD pipelines, with validation, logging, secure credential handling, and rollback. /li liApply incident, change, and problem management pragmatically, protecting service availability without introducing unnecessaryprocess. /li liConnect technical and business teams: clarify service ownership and dependencies, translate business needs into reliability and infrastructure requirements, frame risk and tradeoffs,
align priorities, and ensure decisions have accountable owners and realistic commitments. /li /ul h3What We Are Looking For /h3 ul liWe care more aboutdemonstratedproduction engineering experience than a checklist of certifications. Strong candidates will bringall ofthe following. /li liA strong SRE or Production Engineering background, typically5+yearsoperating business-critical, customer-facing, or high-availability services. Recent work must include direct technical ownership, not only coordination or people management. /li liRecent ownership of high-severity incidents, including technical triage, recovery decisions, clear communications, and measurable follow-through. /li liStrong systems and network troubleshooting fundamentals: Windows and Linux, TCP/IP, DNS, routing, firewalls, proxies or load balancers, and the ability to isolate faults across service layers. /li liHands-on cloud operations experience in Azure, OCI, or a similar platform, includingcompute, storage, networking, IAM, observability, backup, and recovery. /li liProduction experience with containers and Kubernetes, including workload health, scheduling, networking, persistent storage, secrets, deployment and rollback, scaling, and backup or recovery considerations. /li liDeep observability and reliability engineering practice: metrics, logs, distributed tracing, actionable alerting, SLI/SLO design, capacity and saturation analysis, failure-mode thinking, and post-incident engineering. /li liAbility to diagnose database-backed and API-driven services across application, query, connection-pool, storage, certificate, network, and downstream dependency layers. /li liPractical identity and access management experience with Active Directory and Microsoft Entra ID or equivalent, including hybrid identity, privileged access, MFA, service identities orgMSAs, lifecycle controls, dependency mapping, and controlled recovery from identity failures. /li liEvidence of safe automation and infrastructure-as-code work using PowerShell, Python, Terraform or comparable tooling and CI/CD. You should be able to explain testing, idempotency, error handling, credential security, rollout, rollback, and measurable impact.
/li liExperience leading, mentoring, or acting as the senior escalation point for other technical professionals. /li liStrong business-facing and cross-functional leadership. You can translate technical complexity into business impact and options, challenge unsafe or unrealistic requests constructively, negotiate priorities, and communicate decisions clearly to engineers, executives, customers, and non-technical stakeholders. /li /ul h3Additional Relevant Experience /h3 ul liMicrosoft 365, endpoint management, EDR, device compliance, and hybrid workplace operations. /li liFormal ITIL, cloud, security, or infrastructure certifications. /li /ul h3What success looks like /h3 ul liIncidents are diagnosed and resolved with greater speed, structure, and confidence. /li /ul ul liMonitoring, runbooks, automation, recovery controls, and change practices reduce repeat failures and operational toil. /li /ul ul liThe team becomes more capable and accountable without relying on a single hero. /li /ul ul liTechnical and business teams share clear service ownership, priorities, risk decisions, and delivery expectations. /li /ul h3Our hiring process /h3 pThe process includes a scenario-based SRE technical discussion. We will ask you to think aloud through realistic production incidents involving cloud, Kubernetes, identity, networking, databases, storage, APIs, and service dependencies. You will be expected to describe the logs, metrics, traces, commands, tools, tradeoffs, and recovery criteria you would use. We will also assess how you align technical and business stakeholders when priorities, risk, and customer commitments conflict. /p h3Our Vision, Purpose, and Values /h3 pOur VISION: Unparalleled control over demand and supply to deliver certainty.br/Our PURPOSE: Problem Solvers Welcome.br/Our VALUES: Deliver the Goods - Have Deep Care - Find the Right Answer, Not the First Answer - Creativity That Endures - Brilliant But Not Loud. /p h3Salary range: /h3 p55-68k/year, plus 10% premio based on personal and company objectives. /p h3Equal Opportunity Employer /h3 pToolsGroup provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.br/ToolsGroup is an E-Verify employer, to learn more please visit E-Verify.gov /p /p #J-18808-Ljbffr
11 ago - Forlì
Visibilia Group
11 ago - Truccazzano
electio
11 ago - Roma
Italia civile
11 ago - Olbia
My Work