25 set - Torino
Icaro Foundation
pstrongAI Safety Researcher /strong /ppstrongIcaro Foundation · Rome preferred · Flexible arrangements /strong /ppbr / /ppstrongThe work /strong /ppIcaro Foundation is an independent non-profit AI safety lab based in Rome. We study advanced AI systems: what they can do, how they fail, and how those findings can support developers and institutions responsible for their governance. /ppbr / /ppstrongWe see AI safety as one of the defining scientific and societal challenges of our time. /strong As AI systems become more capable, autonomous, and widely deployed, understanding and reducing their risks is increasingly urgent. We are looking for people who are deeply interested in these questions and motivated to contribute through rigorous research. /ppYou will stronghelp produce new research and develop the lab’s shared codebase and knowledge base /strong, working closely with our researchers across the research process: reviewing literature, refining questions, implementing experiments, analysing results, and contributing to papers and technical reports. /ppOur research focuses particularly on strongagentic, multi-agent, and compositional safety /strong: how risks emerge across extended interactions, tool use, and systems involving multiple AI agents. We also study strongtesting awareness and evaluation validity /strong, including whether models behave differently when they recognise that they are being evaluated. /ppAlongside our research, we evaluate frontier models for international model providers as independent third-party evaluators, using public and proprietary benchmarks and red-teaming environments. /ppOur public work includes: /pulliBoiling the Frog, on multi-turn agentic safety; br / /liliAdversarial Humanities Benchmark, on the robustness of safety behaviour under stylistic reformulations; br / /liliresearch on LLM-to-LLM risks, multi-agent collusion, and interaction-level safety. /li /ulpYou can explore our research programme and papers to learn more. /ppbr / /ppstrongWhat you would do /strong /ppYour work will combine three closely connected areas. /ppstrongContribute to research /strong /pulliReview relevant literature, compare methods, and identify questions worth investigating. /liliHelp turn research questions into experimental protocols, including baselines, controls, and clear evaluation criteria. /liliImplement and run experiments with frontier and open-weight models, including agentic and multi-agent environments. /liliAnalyse results and model traces, investigate unexpected behaviour, and assess confounders and alternative explanations. /liliContribute to research papers, benchmarks, technical reports, and presentations. /li /ulpstrongDevelop the research codebase /strong /pulliWrite and improve Python code for experiments, evaluations, data processing, and analysis. /liliExtend existing tools and environments, fix bugs, and participate in code review. /liliAdd tests, documentation, and reproducible configurations so other researchers can inspect, rerun, and build on your work. /li /ulpstrongBuild the lab’s knowledge base /strong /pulliProduce concise, source-grounded notes on papers, methods, benchmarks, and research questions.
/liliDocument experimental setups, findings, limitations, and negative results. /liliOrganise and connect references, datasets, code, and research notes so the team can find relevant evidence and reuse previous work. /li /ulpYou may bring stronger skills in research or engineering. The role involves both writing code and reasoning carefully about evidence. /ppbr / /ppstrongWho should apply /strong /ppWe welcome applications from strongmaster’s students, PhD students, recent graduates, and researchers at the beginning of their careers /strong, including those who have recently completed a PhD. /ppRelevant experience may come from a thesis, academic research, independent experiments, open-source contributions, internships, or previous employment. We also welcome applicants from non-traditional backgrounds who can demonstrate strong research or engineering ability. /ppA completed PhD, previous AI safety employment, and published papers are not required. We care about the quality of your work, your contribution to it, and your ability to learn. /ppIf you are currently studying, please tell us about your availability and how you would combine the role with your academic commitments. /ppstrongWhat we are looking for /strong /pulliExperience with LLM evaluations, red-teaming, or benchmark development; br / /liliA solid technical or quantitative background, developed through university study, independent projects, or relevant work. /liliGood Python skills and familiarity with Git, debugging, and working with an existing codebase. /liliPractical experience with machine learning or LLMs through at least one substantive project or research contribution. /liliAn understanding of basic experimental reasoning and statistics: comparing conditions, interpreting results, and recognising uncertainty and possible confounders. /liliThe ability to read technical papers critically and explain methods, findings, and limitations clearly in English. /lilistrongA strong interest in AI safety and a sense of urgency about understanding and reducing the risks posed by increasingly capable AI systems. /strong /liliIntellectual curiosity, openness to criticism, and a willingness to revise your views in response to evidence. /liliMotivation to contribute to a shared research effort, including the code, documentation, and accumulated knowledge that make good research possible. /li /ulpstrongUseful, not required /strong /ppExperience with: /pulliagentic or multi-agent systems; br / /lilistatistical analysis or experimental replication; br / /lilisoftware testing, containers, or reproducible research workflows; br / /lililiterature reviews, research documentation, or open-source contributions; br / /liliInspect AI,
theopen-source evaluation framework developed by the UK AI Security Institute and Meridian Labs, or comparable tools. /li /ulpFor an example of our research software, see the Adversarial Humanities Benchmark codebase, also listed in Inspect Evals as an externally maintained evaluation. /ppYou do not need experience in all of these areas. /ppbr / /ppstrongHow we work /strong /ppWe are a small research team. You will work closely with experienced researchers and receive feedback on experimental design, code, analysis, and writing. /ppYou will begin with clearly scoped contributions to ongoing projects and take on greater responsibility as your skills and familiarity with the work develop. We encourage everyone to ask questions, challenge assumptions, and propose ideas. /ppExisting evaluation infrastructure, technical support, and API budget are available. Contributions may become public papers, benchmarks, datasets, or tools where compatible with confidentiality obligations. Authorship and acknowledgement will reflect contributions. /ppWe value work that others can understand and build on: clear reasoning, reliable code, well-documented experiments, and honest reporting of uncertainty. /ppstrongPractical details /strong /pullistrongLocation: /strong Flexible, with a preference for working in person with the team in strongRome, Italy /strong. /lilistrongIn-person collaboration: /strong We particularly welcome applicants who are based in Rome or would be interested in relocating. We value regular in-person discussion, collaborative experimentation, and learning from one another. /lilistrongRemote arrangements: /strong May be considered for candidates based in Europe or China, with substantial overlap with European working hours. /lilistrongEngagement: /strong Contractor role. /lilistrongCompensation: /strong The specific compensation range will be shared during the first interview. /lilistrongWorking language: /strong English. /lilistrongStart date: /strong By November 2026. /li /ulpbr / /ppstrongHow to apply /strong /ppSend your application to with the subject strong“AI Safety Researcher — (Your name)” /strong. /ppPlease include: /pollistrongYour CV. /strong /lilistrongRelevant links /strong, such as GitHub, a personal website, or Google Scholar, where available. /lilistrongOne example of relevant work: /strong a thesis, paper, repository, notebook, technical report, experimental replication, or a substantial contribution to a shared project. /lilistrongA few sentences about why you want to work on AI safety and what interests you about our research /strong, together with your availability and whether you could work with us in Rome. A separate cover letter is not necessary. /li /olpAccompany your work sample with a short explanation, up to one page, covering: /pullithe question or problem you addressed; br / /liliwhat you personally contributed; br / /lilithe approach and main result; br / /lilian important limitation or something you would change. /li /ulpIf your work cannot be shared publicly, you may submit a description that excludes confidential information. /ppApplications are reviewed on a rolling basis. Please include all requested materials so we can assess your application. /p
25 set - Roma
DAVID NAMAN
25 set - Udine
OLTRE I CONFINI 2.0
25 set - Friuli-Venezia Giulia
Arethusa
25 set - Pescara
Alll