31 lug - Bardi
Altro
Experteer Overview As principal architect, you define the vision for compute infrastructure that underpins large-scale AI/ML systems. You guide the architecture across compute, networking, storage, and orchestration to deliver scalable, cost-aware production systems. You work with cross-functional teams, validate designs with prototypes and benchmarks, and mentor others while driving strategic roadmaps. Your deep hyperscaler expertise informs decisions that balance performance, cost, and business value, shaping the firm of AI infrastructure trajectory.
Benefits Set and communicate the overarching compute infrastructure strategy for AI/ML systems
Make authoritative architecture decisions across compute, networking, storage, orchestration, and model serving
Architect and prototype large-scale, cost-optimized compute and distributed training systems with reference implementations and benchmarks
Define reference architectures, standards, and patterns and implement foundational tooling and automation
Lead enterprise-scale architecture assessments and design reviews with hands?on validation
Shape the AI infrastructure roadmap and plan capacity and technology evolution
Identify and pilot emerging technologies with real-condition testing
Drive performance and cost optimization of GPU/compute workloads to meet SLAs
Serve as the principal authority on hyperscaler cloud platforms with hands?on AI/ML expertise
Lead root?cause analysis for complex issues across hardware, network, software, and models
Foster relationships with infrastructure partners for early access and credibility
Provide executive? and client?level advisory translating trade?offs into business outcomes
Define monitoring, observability, reliability strategies and implement SLAs, SLOs and governance for production AI/ML systems
Ensure security, compliance, and regulatory alignment across AI/ML infrastructure
Mentor and develop the architect community to deliver impact
Champion cost?efficiency and value realization across the stack
Responsibilities Significant experience coding, building, monitoring and troubleshooting AI/ML applications and deploying them on premises or public cloud
Strong understanding of AI/ML concepts
Strong understanding of computing infrastructure; knowledge of AI infrastructure preferred
Proficiency in Python, Java, or C++
Experience with data pipelines/workflow tools (e.g., Apache Airflow, Kubeflow)
Strong problem?solving ability and fast?paced adaptability
Excellent communication and collaboration skills
Extensive experience in AI/ML infrastructure engineering on hyperscaler platforms for large?scale deployments
Proven leadership and management of AI projects and teams
Strong project management skills with multi?project capability
Experience evaluating and selecting AI technologies and frameworks
Ability to collaborate with cross?functional teams and drive project alignment
#J-*****-Ljbffr
07 ago - Roma
Akkodis
07 ago - Milano
ORIENTA
07 ago - Veneto
Adecco
07 ago - Premana
Adecco