Frontend Code Evaluation Specialist
Pubblicato il 27-09-2026 - Obsidian in Milano
ph3About the work /h3 pWe're building a high-quality dataset of human preference judgments on AI-generated frontend code. Each task hands you a reference web page — crawled from the real internet, delivered as a full zipped site tree plus screenshots of its default view and, on some pages, additional states reached by hovering, clicking, or scrolling. Alongside it come two model attempts, A and B, each a zipped self-contained site tree. The models only ever saw the screenshots; they never had the source. /p pYou download all three, run them locally, view each at a 1920×1080 viewport, interact with them to reach every required state, then open the source of both attempts and grade them against each other — on visual fidelity per state, and on how the code is actually constructed. Structure and responsiveness are explicitly part of the rubric, not just the render. /p pThis is evaluation work, not authoring. The defining skill is not that you can build a page — it's that you can open someone else's page and tell how it was built and where it cheats. /p h3Please read before applying /h3 ul lipEach unit takes roughly 2–3 hours and is timed. This is not microtask work; if you can only offer scattered 15-minute windows, you will not be able to finish a unit. /p /li lipYou need a real local development environment. A tablet, a Chromebook, or a locked-down work machine that cannot run a local static server will not work for this project. /p /li lipYou need at least one completed Mercor engagement, delivered in full. We are not onboarding net-new experts to this project. /p /li /ul h3What you'll do /h3 ul lipRender a reference page and two candidate replications at 1920×1080 and judge which is the closer reproduction, state by state. /p /li lipDiff visual fidelity in detail: box model and spacing, typography (family, size, weight, line-height, letter-spacing), color and border treatment, image and asset handling, z-order and overflow. /p /li lipRead the source of both attempts and grade construction quality — distinguishing a replication that is genuinely correct from one that merely looks correct at one viewport. Hardcoded pixel offsets,
absolute positioning standing in for real layout, inline style soup, a single undifferentiated div tree, or a screenshot pasted in as an instead of a rebuilt section. /p /li lipTest responsiveness: a nav bar that looks right at 1920px but collapses at 1400px is a defect, and you should be able to say precisely why. /p /li lipWrite a specific, evidence-cited justification for every preference. We need "B nests the article body in a single absolutely-positioned div, so the text overlaps the footer below 1600px, while A uses normal document flow" — not "A looks closer." /p /li lipUse the "this task is broken" escape hatch with judgment: distinguish a task that genuinely cannot be completed from the screenshots provided from one that is merely hard. Over-flagging and under-flagging are both failure modes. /p /li /ul h3You're a fit if you have /h3 ul lip3–8 years of professional web development experience, shipping web interfaces for a living, primarily in frontend or full-stack work. /p /li lipFluency across web eras. The reference pages are real crawled sites — one is a small charter-fishing business built in the table-and-image-map tradition, another a corporate press-release page with stacked navigation rows and social share widgets. If your entire career happened inside a modern component framework and you have never authored raw CSS or seen a table used for layout, you will misjudge many of these pages. /p /li lipCommand of hand-written HTML and CSS: semantic markup, flexbox, grid, media queries, and legacy float- and table-based layouts you can read and reason about. You should be able to look at a rendered layout and predict what's holding it together before opening DevTools.
/p /li lipBrowser DevTools as muscle memory — setting an exact viewport, walking the element tree, checking computed styles, watching what a hover handler mutates. /p /li lipCommand-line comfort: unzipping an archive, standing up a static local server because the relative asset paths demand it, and untangling a broken image reference rather than giving up and grading from the screenshot. /p /li lipEnough JavaScript to read a page's scripts and understand what they do to the DOM, even if you don't write JS daily. /p /li lipProfessional written English. Every task ends in a free-text justification, and a rating without a specific rationale is worth very little. /p /li /ul h3Equipment /h3 ul lipA desktop or laptop that displays a 1920×1080 viewport. /p /li lipAdministrator rights on your own machine, so you can install and run a local server. /p /li /ul h3Nice to have /h3 ul lipPrior RLHF, preference-labeling, model-evaluation, or structured code-review work — the strongest single signal. Rubric-driven comparison at volume needs almost no ramp here. /p /li lipPixel-perfect design-to-code experience: agency work, design systems, template production. Anyone who has had a designer reject a build over four pixels has exactly the fidelity eye this needs. /p /li lipAccessibility expertise (ARIA, semantic landmarks, heading hierarchy) — you'll notice immediately when an attempt renders a heading as a styled .. /p /li lipFamiliarity with how LLMs fail at code generation. /p /li lipWeb scraping, archiving, or DOM-parsing background — comfort with messy crawled site trees. /p /li lipMore than one completed Mercor project, and availability in contiguous multi-hour blocks. /p /li /ul h3Note: /h3 pthis seat is for practicing web developers. Backend-only, ML/data-science-only, mobile-native-only, and DevOps-only engineers do not have the UI instincts this requires, however strong they are otherwise. Designers who do not code cannot grade the source axes at all. Framework-only engineers who have never authored CSS outside a component library will struggle with the legacy reference pages. /p /p #J-18808-Ljbffr
