Mlops & Platform Engineer

04 set - Cologno Monzese
TXT GROUP

ppbTXT Group /b, a company within the bTXT Group /b, is seeking an bMLOps Platform Engineer /b to join its Industrial Business Unit.
The ideal candidate will be responsible to design, build and operate the company's application and AI platform, ensuring secure, scalable and highly available environments for both enterprise applications and AI/ML workloads.
At least three years' relevant experience in platform and infrastructure engineering is required, with production ownership of: Kubernetes, CI/CD pipelines and cloud infrastructure.
/ph3Main responsibilities /h3ulliDesign and manage cloud-native platforms, Kubernetes clusters and containerised applications; /liliBuild and maintain CI/CD pipelines, and automate infrastructure provisioning and application deployment; /liliDesign, deploy and manage infrastructure for AI systems: LLM and embedding model serving, vector databases, application databases and caching systems; /liliDefine and maintain CI/CD pipelines for AI applications and data pipelines, with reproducible environments and secure release strategies; /liliManage versioning of models, prompts and configurations, and support fine-tuning and retraining pipelines where required; /liliCollaborate with software developers and AI engineers to streamline delivery and operations; /liliImplement monitoring, logging, tracing and alerting for LLM applications and data pipelines, covering latency, error rate, cost per call, drift and production quality metrics; /liliEnsure scalability, high availability and operational continuity of AI services, including knowledge base ingestion and update pipelines; /liliOptimise inference and embedding costs through resource sizing, quantisation, batching and infrastructure-level caching; /liliEnsure overall platform security, observability, performance and operational reliability; /liliManage secrets, API keys, access control and environment isolation for AI services; /liliSupport offline evaluation and A/B testing activities by providing the necessary infrastructure and telemetry data.




/li /ulh3Indispensable technical skills /h3ulliSolid experience with containerisation and orchestration technologies such as Docker, Kubernetes, Helm and Docker Compose; /liliProficiency with CI/CD and DevOps tooling, including Git, GitLab and GitLab CI/CD (or equivalents such as GitHub Actions), GitOps workflows, container registries and release management practices, with automation skills in Python and Bash; /liliStrong background in infrastructure automation on Linux, using Terraform and Ansible to implement Infrastructure as Code and manage virtualised environments; /liliSolid understanding of networking and security fundamentals (TCP/IP, HTTP/HTTPS, DNS), reverse proxies and ingress controllers (Nginx, Traefik), TLS/SSL, identity and access management (Keycloak, OAuth2/OpenID Connect), secrets management and IAM; /liliExperience with observability stacks such as Prometheus, Grafana, Loki, OpenTelemetry, OpenSearch and Alertmanager (or an equivalent ELK-based stack), applied to production ML and LLM systems; /liliHands-on experience with AI/MLOps tooling, including MLflow and experiment tracking platforms such as Weights Biases, model registries, and model serving frameworks such as vLLM, Ollama, Triton Inference Server, SageMaker or Vertex AI, applied to GPU-based inference workloads; /liliExperience deploying and operating vector databases (Qdrant, pgvector), embedding models and RAG pipelines as part of production AI inference services; /liliProficiency in backend and data technologies, including Python, FastAPI and REST APIs, together with operational experience running PostgreSQL,



Microsoft SQL Server and Redis in production; /liliPractical experience operating production databases across SQL/NoSQL and vector stores, covering provisioning, backup and scaling; /liliFamiliarity with LLMOps practices, including prompt and experiment tracking, inference cost monitoring and model version management; /liliUnderstanding of inference optimisation techniques such as quantisation, batching, caching and GPU-level optimisation (e.g. TensorRT); /liliExperience managing the ML/LLM model lifecycle, including fine-tuning and retraining pipelines, offline evaluation and A/B testing; /liliWorking knowledge of at least one major cloud platform (Microsoft Azure, AWS or Google Cloud Platform) and its services for ML/AI workloads.
/li /ulpFamiliarity with HashiCorp Nomad is a PLUS.
/ph3Education /h3pbEducation /b: Bachelor's or Master's degree in Computer Science, Computer Engineering or a similar field.
/ppThe perfect candidate will also possess problem-solving skills, curiosity, the ability to work independently, a proactive approach, a sense of responsibility, a collaborative spirit, the ability to draft technical documentation, and the ability to work effectively within cross-functional teams.
/ph3Why choose TXT Group /h3ulliHybrid working mode; /liliCareer opportunities in a fast-growing, international and dynamic environment; /liliContinuous training on technical and project-related topics; /liliCorporate benefits including health insurance, welfare services, meal vouchers, and employee discounts; /liliThe contract offered for this position is a permanent one, and the salary range for this position is between EUR ****** € and EUR ****** € gross per year.
/liliThe grading/level will be defined during the selection process based on the candidate's profile and in accordance with the National Collective Labor Agreement (Metalworking Industry).
/li /ulpThe company promotes equal opportunities and values diversity in all its forms.br/br/Position open to candidates without distinction of gender, pursuant to Legislative Decree ********.br/br/ /p /p #J-*****-Ljbffr

Autista Ce cqc

06 set - Nola
Dv Service & Trade

Macellaio

06 set - Roma
Ro.ma.carni

Ricevi nuove offerte di lavoro

Crea una Job Alert gratuita per mlops & platform engineer / cologno monzese

Macellaio

06 set - Roma
Ro.ma.carni

Badante

06 set - Torino
Bruno Lombardi