Senior Design Expert - Paid AI Design Research Study
$150 - $250 / hour
Verified partner opportunity
$65 - $75 / per-task
About the work
About the work
Mercor is assembling a panel of chemistry and chemical safety experts to red-team frontier AI models. The goal is to test whether a model can correctly judge the misuse potential of a technical request
answering legitimate questions fully while refusing genuinely dangerous ones.
You will: • Write challenging single-turn prompts in your domain, labelled across three levels: benign, dual-use, and adversarial. • Evaluate the responses that come back against a defined policy standard, and judge whether each was handled correctly. • Write the reference answer
what a correct response looks like, and the technical reasoning for why.
This domain lives on the dual-use line more than most. The same method that detects a hazardous compound at trace level defines the threshold someone would try to stay beneath; the same library entry that identifies an unknown tells you what a screen will and won't see. A model that refuses a routine method-validation question is failing as badly as one that answers a question it shouldn't. Drawing that line reliably is why we need practitioners rather than generalists.
Background we're looking for
We're looking for analysts who have built and validated methods for hazardous compounds
people who know what a detection limit really means and what it implies. What matters most is being able to tell a routine professional question from one that's fishing for something else: the same detection method that finds a compound describes how to evade the screen. You will be writing prompts that sit exactly on that line, then judging whether the model held it.
Examples of relevant backgrounds (ideally with red-teaming experience): • Detection and identification of hazardous compounds
GC-MS, LC-MS/MS, IR, NMR • Method development and validation for trace-level analysis • Environmental, occupational or emergency-response sampling and analysis • Reference standards, spectral libraries and unknown identification • Laboratory accreditation, QA/QC and proficiency testing
What this role demands beyond technical depth
This is writing-intensive work. Every judgment you make needs a written rationale that a non-specialist can follow. Prior technical writing, published research, or expert witness experience is a strong signal; please include a sample or link. You will also be reading and writing about misuse scenarios in your field for sustained periods. We brief experts on this in advance, and you can pause or step away at any point without penalty.
Before you apply
Your work here will not involve, and must not draw on, classified or export-controlled information, or anything covered by an NDA or prepublication review obligation. If you hold such obligations you may still be a good fit; tell us in your application and we will scope the work accordingly.
This opportunity may suit professionals with relevant experience in AI evaluation, Research. Review the official description and requirements before applying.
The listing states $65 - $75 / per-task. Confirm the final rate, workload, and payment terms during the official application process.
Use your CV to find relevant roles on HumanitApp. This is separate from a partner application.
Match my CV →Optional role alerts from HumanitApp. Subscribe only if you want weekly emails.
Get weekly role alerts →Before you apply
This listing is not a promise of immediate work. Mercor may send separate offers to prequalified candidates. Read how Instant Work Offers work.
The listing states $65 - $75 / per-task. Confirm the final rate, workload, and payment terms during the official application process.
Yes. The button opens the exact verified Mercorlisting using the referral URL published for this role.
No. HumanitApp independently curates the opportunity. The partner platform manages applications and hiring decisions.
No. Choose the primary Apply action to continue directly without giving HumanitApp your name or email address.