Data Scientist Talent Network
$100 - $150 / hour
Verified Mercor opportunity
About the work — We're building a high-quality dataset of human preference judgments on AI-generated frontend code. You'll be shown a reference web page alongside two candidate replications produced by AI models, and you'll decide which replication is better — then explain why in writing that a model can learn from. — This is evaluation work, not authoring. You won't be building sites from scratch. You'll be reading someone else's HTML and CSS, running it locally, comparing it pixel-by-pixel against a target, and articulating exactly where and why it falls short. — What you'll do • Render a reference page and two candidate replications side by side at desktop width and judge which is the closer reproduction. • Diff layout fidelity in detail: box model and spacing, typography (family, size, weight, line-height, letter-spacing), color and border treatment, image and asset handling, z-order and overflow. • Inspect the underlying markup with browser devtools to distinguish a replication that is genuinely correct from one that merely looks correct at one viewport — hardcoded pixel offsets, absolute positioning standing in for real layout, and inline styles that will not survive a resize. • Evaluate responsive behavior and semantic quality: whether flexbox and grid are used where they belong, whether legacy float or table layouts in the reference were reproduced faithfully, whether headings and landmarks carry real semantic meaning. • Write a structured rationale for every judgment — the specific defects you found, ranked by how much they matter, in language precise enough to be actionable. • Flag ties, ambiguous cases, and broken task items rather than forcing a preference. — You're a fit if you have • 3+ years of professional web development experience, primarily in frontend or full-stack work. • Fluency in hand-written HTML and CSS: semantic markup, flexbox, grid, media queries, and older float- and table-based layouts you can still read and reason about. • Working command of browser devtools — element inspection, computed styles, the box model, and the network panel. • Enough JavaScript to read a page's scripts and understand what they do to the DOM, even if you don't write JS daily. • Comfort in a terminal: cloning a folder and serving it over a local static server without help. • Strong written English and the discipline to justify a judgment rather than assert it. — Equipment • A desktop or laptop with a browser window that opens to at least 1920px wide. • Administrator rights on your own machine, so you can install and run a local server. — Nice to have • Prior RLHF, preference labeling, or model evaluation work. • A code review or technical assessment background. • Pixel-perfect design-to-code experience — translating Figma or PSD comps into production markup. • Web accessibility expertise (WCAG, ARIA, screen reader testing). • Familiarity with how LLMs typically fail at codegen. • Web scraping or DOM parsing experience. — Note: this seat is for practicing web developers. Backend-only, mobile-native-only, data science, DevOps, and design-without-code backgrounds are out of scope for this project.