Junior data engineer - Italia

23 set - Milano
Cato

Your mission

Bring a tender from the source portal into Cato: scraping, parsing, merging, enrichment. You'll start by owning a handful of sources end to end - the scraper, the job behind it, and the data that comes out - and take on more as you go. Not tickets handed to you: sources you're responsible for.

Se le sue competenze, la sua esperienza e le sue qualifiche corrispondono a quelle descritte in questa panoramica, non ritardi l'invio della sua candidatura.
What you'll actually do

Build and maintain scrapers for national tender portals, where reading the source in its original language is part of the job. Keep them alive: portals change their HTML, move endpoints, break pagination, throttle you. You find out before the customer does. Turn messy sources into clean records: broken HTML, inconsistent XML, APIs that lie about their own schema. Write and maintain orchestrator flows: retries, backfills, alerting, and a clear answer to "Did today's run actually land?" Work on merge and dedup - the same tender arrives three times, in three shapes, and only one version can reach the customer. Ship AI enrichment steps: batch LLM extraction of requirements, embeddings, OCR on attachments. Guard data quality with tests and checks that fail loudly before a customer finds the gap.

Ideal profile

Python that holds up: typed, tested, and readable six months later. You've scraped something real: HTTP, HTML and XML parsing, pagination, sessions,



rate limits - and you know why a scraper that worked yesterday is broken this morning. SQL you're comfortable in: joins, aggregations, window functions. You'll read from the database every day; tuning and running it isn't your job. Builder by default: you see a manual process and your first instinct is to automate it. Comfortable with messy sources: broken HTML, inconsistent XML, PDFs that were scans of scans. You close your own loop: you check that what you shipped actually ran, before someone else has to ask. Experience 1-2 years writing Python in production: scrapers, ETL scripts, automation - anything that had to run unattended and be fixed when it didn't. Exposure to an orchestrator (Prefect, Airflow, Dagster) is a plus, not a requirement: you'll learn ours properly. Exposure to LLM-based extraction is welcome; curiosity about it is mandatory.

What you won't find here

No micromanagement: we trust you to own your part of the stack. No "standard" 9-to-5 mentality: we care about outcomes and we are looking for people who are willing to go the extra mile. No "we've always done it this way" excuses: we're here to disrupt, not to follow old patterns. xysqume

Our Tech Stack

Data & Infra: Python, PostgreSQL, Prefect, AWS AI: batch LLM extraction, embeddings, OCR

Compensation

RAL €35,000 - €45,000 equity, depending on profile.

Hiring Manager

Lorenzo Rossetto

#J-18808-Ljbffr

Agente di commercio

25 set - Pescara
NOVAMEDICA

Operatore cnc

25 set - Fermo
Adecco Italia

Ricevi nuove offerte di lavoro

Crea una Job Alert gratuita per junior data engineer - italia / milano

Sales Development Representative

25 set - Trentino-Alto Adige
Chino.Io

Progettista Elettrico – Cabine Media / Bassa Tensione - Padova - €40.000 - €50.000 All'Anno

25 set - Padova
Modulogroup