17 set - Lazio
Cato
pstrongYour mission /strong /ppBring a tender from the source portal into Cato: scraping, parsing, merging, enrichment.
You'll start by owning a handful of sources end to end - the scraper, the job behind it, and the data that comes out - and take on more as you go.
Not tickets handed to you: sources you're responsible for.
/ppbr / /ppstrongWhat you'll actually do /strong /pulliBuild and maintain scrapers for national tender portals, where reading the source in its original language is part of the job.
/liliKeep them alive: portals change their HTML, move endpoints, break pagination, throttle you.
You find out before the customer does.
/liliTurn messy sources into clean records: broken HTML, inconsistent XML, APIs that lie about their own schema.
/liliWrite and maintain orchestrator flows: retries, backfills, alerting, and a clear answer to "Did today's run actually land?" /liliWork on merge and dedup - the same tender arrives three times, in three shapes, and only one version can reach the customer.
/liliShip AI enrichment steps: batch LLM extraction of requirements, embeddings, OCR on attachments.
/liliGuard data quality with tests and checks that fail loudly before a customer finds the gap.
/li /ulpbr / /ppstrongIdeal profile /strong /pullistrongPython that holds up: /strong typed, tested, and readable six months later.
/lilistrongYou've scraped something real: /strong HTTP, HTML and XML parsing, pagination, sessions, rate limits - and you know why a scraper that worked yesterday is broken this morning.
/lilistrongSQL you're comfortable in: /strong joins,
aggregations, window functions.
You'll read from the database every day; br / tuning and running it isn't your job.
/lilistrongBuilder by default: /strong you see a manual process and your first instinct is to automate it.
/lilistrongComfortable with messy sources: /strong broken HTML, inconsistent XML, PDFs that were scans of scans.
/lilistrongYou close your own loop: /strong you check that what you shipped actually ran, before someone else has to ask.
/li /ulpbr / /ppstrongExperience /strong /pulli1-2 years writing Python in production: scrapers, ETL scripts, automation - anything that had to run unattended and be fixed when it didn't.
/liliExposure to an orchestrator (Prefect, Airflow, Dagster) is a plus, not a requirement: you'll learn ours properly.
/liliExposure to LLM-based extraction is welcome; br / curiosity about it is mandatory.
/li /ulpbr / /ppstrongWhat you won't find here /strong /pulliNo micromanagement: we trust you to own your part of the stack.
/liliNo "standard" 9-to-5 mentality: we care about outcomes and we are looking for people who are willing to go the extra mile.
/liliNo "we've always done it this way" excuses: we're here to disrupt, not to follow old patterns.
/li /ulpbr / /ppstrongOur Tech Stack /strong /pulliData Infra: Python, PostgreSQL, Prefect, AWS /liliAI: batch LLM extraction, embeddings, OCR /li /ulpbr / /ppstrongCompensation /strong /ppRAL €35,000 - €45,000 + equity, depending on profile.
/ppbr / /ppstrongHiring Manager /strong /ppLorenzo Rossetto /p
19 set - Sona
Studio medico
19 set - Trieste
Euroservis
19 set - Marina di Ravenna
MedTug
19 set - Caserta
Bufalè srls