Role directory

AI evaluation jobs

346 active, referral-verified opportunities.

Finance / Remote

Corporate Treasury Expert

About the work We're building a high-quality library of corporate treasury work products. You'll complete self-contained treasury exercises from mock files (bank statements, exposure schedules, credit agreements) and produce the deliverable for each, graded against a rubric. — What you'll do • Build rolling 13-week direct cash forecasts with receipts, disbursements and variance explanation. • Identify and net FX exposures, analyze fixed and floating rate risk, and design hedges against a corporate policy. • Produce hedge designation documentation with periodic effectiveness testing. • Analyze working capital performance (DSO/DPO/DIO), bank fees against services, trapped cash and repatriation options, and forecast leverage and covenant headroom. • Draft or refresh treasury policy covering liquidity, FX, counterparties and investment. — You're a fit if you have • 4+ years in a corporate treasury role at an operating company. • Hands-on with cash forecasting, hedging, or bank and liquidity management. — Nice to have • CTP; ACT, AMCT or MCT; treasury management systems. — Assessment A 13-week cash forecast from mock bank and receivables files; an FX exposure netting and hedge recommendation; a covenant headroom forecast. — Note: this is corporate treasury. Bank asset-liability management, funds transfer pricing and regulatory liquidity reporting are out of scope.

$2,000 / per-taskOpen / Referral verified
Finance / Remote

Corporate Tax Expert

About the work We're building a high-quality library of corporate tax work products. You'll complete self-contained tax exercises from mock files (trial balances, intercompany agreements, deal documents) and produce the deliverable for each, graded against a rubric. — What you'll do • Build quarterly and annual income tax provisions under ASC 740 with current and deferred calculations and rate reconciliation. • Prepare returns and supporting workpapers across federal, state apportionment and international regimes, and forecast cash taxes by jurisdiction. • Produce technical computations and studies: GILTI, FDII, BEAT, Pillar Two, 163(j) limitation, and net operating loss and attribute analysis. • Build transfer pricing master file and local file documentation with intercompany agreement and policy review. • Work transaction tax: diligence of exposures and attributes, structuring memos, and documented tax positions including FIN 48 reserves. — You're a fit if you have • 4+ years in corporate income tax, in-house or Big 4. • Depth in at least one of: provision and compliance, international and transfer pricing, or M&A tax. — Nice to have • CPA; JD or LLM in Taxation; EA; CTA. Tax equity and HLBV modeling experience. — Assessment An ASC 740 provision from a mock trial balance; a transfer pricing benchmarking memo; a deal tax structuring recommendation. — Note: this is corporate income tax. Individual return preparation, payroll tax, and sales-tax-only backgrounds are out of scope.

$2,000 / per-taskOpen / Referral verified
Code / Remote

Agent Engineer

We are looking for engineers who build and operate LLM agents in production, and who have real visibility into how agents are actually used inside a company. — You have probably: • Shipped an agent that real users depended on, and been on the hook when it broke. • Figured out how to tell whether an agent got better or worse after a change. • Run agents that outlive a single request: scheduled jobs, long-running work, cloud sandboxes. • Watched your org build an internal assistant, and seen who adopted it and who quietly did not. — We are especially interested in the layers most people do not talk about: internal monoagents wired into company data, shared company memory, reusable skills and playbooks, the tool and MCP surfaces agents call, and how anyone sees what agents did and what they cost. — Applying starts with a short conversational AI interview. No coding, no take-home. We want to hear how you actually think about agent reliability, evaluation, and adoption, and the tradeoffs you have made in real systems. Bring war stories. The messier and more specific, the better. — If that screen stands out, we will invite you to a live 30 minute conversation with our team. We pay $100 to $500 for that conversation, paid on completion of the call, with the amount depending on depth of experience. — If that sounds like you, apply and complete the screen. We review every submission.

$100 - $500 / per-taskOpen / Referral verified
Code / Remote

CUDA Engineering Expert

1. Role Overview — Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You’ll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures. — 2. Key Responsibilities • Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization • Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements • Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms • Write, modify, and reason about C++17, Python, and GPU programming code • Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes • Document optimization decisions clearly, including when specific profiler metrics are or are not useful — 3. Ideal Qualifications • Available to work at least 20 hrs/wk • Fluent in core C++ features through C++17 • Working knowledge of Python and Git • Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming • At least 1 year of professional or graduate-level research experience working with GPUs • Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels • Ability to optimize GPU kernels without needing deep prior context on every algorithm • Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus • Experience optimizing kernels for NVIDIA Blackwell hardware is a plus • Familiarity with NSight Compute is a plus • Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus • Open-source contributions related to GPU kernel optimization are a plus — 4. Application Process • Submit your resume or relevant technical background to get started • Qualified applicants may be asked to complete a brief technical assessment or submit additional information

$500 / per-taskOpen / Referral verified
Business / Remote

US-Based Business Owners Using Google Chat

Core Requirement: Profile Ownership — Applicants must be active owners or administrators of a verified Google Business Profile (GBP). We are looking for individuals who use Google Chat regularly within their business. — Eligibility • US-based, English only • Minimum 500 internal messages across all DMs and Spaces • At least 10 employees actively using Chat • Less than 10% of messages are from the past month — Compensation — One-time payment, up to $10,000

$250 / hourOpen / Referral verified
Medical / Remote — US-based or deep US-market experience

Disease-Area Clinician — Trial Endpoints & Prescribing

Mercor is partnering with a biotech and pharma research team on expert human-evaluation work — creating and critiquing rubrics that analyze commercial drugs and drug-development programs. We're seeking a disease-area clinician for a paid pilot, with the potential to continue as the work scales. — Responsibilities • Create and critique rubrics that analyze commercial drugs and drug-development programs • Interpret clinical trial endpoints and judge whether results would change prescribing • Assess what that means for real-world uptake in your therapeutic area • Apply this across the pilot's set of drugs and programs — Requirements • US-based disease-area clinician (practising physician, MD/DO) • Familiarity with biotech/pharma development programs • Deep expertise in a specific therapeutic area • Board-certified and actively practising in your specialty — Engagement Details • Duration: Approximately 10–20 hours over a 1–2 week pilot • Remote — US-based — Why Participate • Help improve expert AI evaluation of drugs and drug-development programs • Recurring engagement for strong contributors

$150 - $230 / hourOpen / Referral verified
STEM / Remote

Senior Design Expert - Paid AI Design Research Study

Mercor is partnering with a major technology company building AI design tooling on a paid research study with senior designers. — The client wants to capture how expert designers think: how you approach a problem, what separates strong craft from weak, and how you decide when something is ready to ship. This is captured in structured, recorded remote sessions. Depending on fit, you'd either be the expert being interviewed or the one running the interview. — What to expect • Roughly 4–5 hours per task/interview, including prep beforehand and a short review afterward • $250/hour — Scope is still being finalized with the client, so structure, timing, volume, and rate may change. — What we're looking for • Senior-level design experience (product, UX, brand, or adjacent craft disciplines) • A clear point of view on design quality, critique, and shipping standards • A portfolio you can share — a website link is ideal — Portfolios are used solely to assess fit for this study and will not be used for model training of any kind. Please only submit work you're free to share publicly — no confidential, unreleased, or NDA-covered material.

$150 - $250 / hourOpen / Referral verified
Finance / Remote

Investment Banking Expert

1. Role Overview — Mercor is partnering with a leading AI research organization to engage experienced investment banking professionals for a project focused on evaluating how well AI systems perform real-world banking work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications. — 2. Key Responsibilities • Design precise, task-specific grading criteria for real-world banking deliverables (financial models, valuation analyses, pitch decks and CIMs, deal memos, board and committee materials) • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score • Apply consistent, evidence-based judgment so that scores are reproducible and defensible • Incorporate structured feedback from senior reviewers and iterate quickly on your work — 3. Ideal Qualifications • 2+ years of professional investment banking experience (associate level or above) • Background at leading global investment banks or elite advisory boutiques, across M&A, coverage, or capital markets • Deep fluency in the day-to-day craft: three-statement and transaction modeling (DCF, LBO, comparables, merger math), valuation judgment, deal-process materials, and the formatting and accuracy standards of bank-quality work product • Exceptionally strong written communication, with the ability to explain exactly why a piece of work does or does not meet the bar • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers • Prior experience with AI training, evaluation, or human-data projects is a strong plus — 4. Application Process • Submit your resume or relevant professional background to get started • Qualified applicants may be asked to complete a brief assessment or submit additional information

$150 - $220 / hourOpen / Referral verified
Medical / Remote — US-based or deep US-market experience

Pharma Commercial Forecasting Expert — Launch Curves

Mercor is partnering with a biotech and pharma research team on expert human-evaluation work — creating and critiquing rubrics that analyze commercial drugs and drug-development programs. We're seeking a pharma commercial / forecasting expert for a paid pilot, with the potential to continue as the work scales. — Responsibilities • Create and critique rubrics that analyze commercial drugs and drug-development programs • Build or critique a launch curve (uptake, peak share, trajectory) • Judge whether a commercial forecast's assumptions (analogs, ramp, peak, loss of exclusivity) are defensible • Apply this across the pilot's set of drug launches — Requirements • US-based pharma commercial / forecasting professional (people from pharma sales teams can work here too) • Familiarity with biotech/pharma development programs • Experience building or critiquing drug launch / uptake models • 5+ years in pharma commercial forecasting, analytics, or strategy — Engagement Details • Duration: Approximately 10–20 hours over a 1–2 week pilot • Remote — US-based — Why Participate • Help improve expert AI evaluation of drugs and drug-development programs • Recurring engagement for strong contributors

$130 - $210 / hourOpen / Referral verified
Medical / Remote — US-based or deep US-market experience

Payer & Market Access Expert — Gross-to-Net & Formulary Strategy

Mercor is partnering with a biotech and pharma research team on expert human-evaluation work — creating and critiquing rubrics that analyze commercial drugs and drug-development programs. We're seeking a payer / market access expert for a paid pilot, with the potential to continue as the work scales. — Responsibilities • Create and critique rubrics that analyze commercial drugs and drug-development programs • Set and pressure-test gross-to-net and formulary-tier assumptions by drug class • Judge whether coverage, rebating, and tiering assumptions reflect how payers actually behave • Apply this across the pilot's set of drugs and drug classes — Requirements • US-based payer / market access professional • Familiarity with biotech/pharma development programs • Hands-on with gross-to-net, rebates, and formulary tiering by drug class • 5+ years in market access, pricing & reimbursement, or managed markets — Engagement Details • Duration: Approximately 10–20 hours over a 1–2 week pilot • Remote — US-based — Why Participate • Help improve expert AI evaluation of drugs and drug-development programs • Recurring engagement for strong contributors

$175 - $200 / hourOpen / Referral verified
Finance / Remote — US-based or deep US-market experience

Biotech Investment Analyst — Drug Asset & Company Assessment

Mercor is partnering with a biotech and pharma research team on expert human-evaluation work — creating and critiquing rubrics that analyze commercial drugs and drug-development programs. We're seeking a biotech investment analyst (specialist-fund background) for a paid pilot, with the potential to continue as the work scales. — Responsibilities • Create and critique rubrics that analyze commercial drugs and drug-development programs • Provide overall company and asset assessment — mechanism, clinical data, competitive position, and value drivers • Judge the quality of an investment-style thesis and flag weak, overconfident, or missing assumptions • Apply your assessment across the pilot's set of companies (~5–10) — Requirements • US-based biotech / healthcare investment analyst with a specialist-fund background • Familiarity with biotech/pharma development programs • Able to assess a drug asset and the company behind it end-to-end • 5+ years evaluating biotech/pharma companies or assets — Engagement Details • Duration: Approximately 10–20 hours over a 1–2 week pilot • Remote — US-based — Why Participate • Help improve expert AI evaluation of drugs and drug-development programs • Recurring engagement for strong contributors

$120 - $200 / hourOpen / Referral verified
Code / Remote

MCP & Plug In Connectors Expert

About the Opportunity — A leading AI research organization is seeking advanced LLM power users with strong experience using MCP and (more importantly) plugins/connectors for real-world personal life tasks. — This project focuses on evaluating how well AI systems handle personalized, multi-step life tasks that require context, judgment, planning, and use of connected tools such as Google Drive, Expedia, Notion, and other plugins/connectors. — This role is ideal for people who use AI heavily in their personal lives and can do a better job replicating what AI could do if it weren’t available — What You’ll Do — You will help evaluate AI systems on complex personal workflows, including tasks across: • Personal health • Travel • Activity planning, including food and dining • Services, such as home repair • Career search • Other life organization workflows — Responsibilities may include: • Creating realistic prompts for complex personal-life tasks • Executing tasks and actions while recording your screen (required) • Using your personal plugins/connectors while you complete actions • Writing clear explanations of AI successes and failures • Judging whether AI outputs are practical, personalized, and well-reasoned • Identifying where models miss context, overreach, fail to use tools correctly, or produce unrealistic results • Creating and applying detailed rubrics to assess model performance — Who We’re Looking For — Strong candidates will have: • US-based only • Strong MCP experience and plug in / connector usage • Experience using LLM plugins/connectors such as Google Drive, Expedia, Notion, and similar tools, multiple times a week • Heavy personal usage of LLM products • An active, rich LLM account with regular usage and approximately 6+ months of history • Willingness to sign a data-share consent form via DocuSign • Experience using AI for multi-step planning, research, decision-making, or personal workflows • Strong written judgment and attention to detail • Ability to explain what makes an AI output good, bad, incomplete, unsafe, or unrealistic • Experience writing and evaluating against rubrics — Extensive rubric experience is especially valuable, including 100+ hours on prior rubric projects involving rubric design, evaluation, and quality assessment. — Ideal Candidate Profile — The strongest candidates are LLM power users who are already using plug-in tools in their personal lives for high-context tasks such as trip planning, health research, home services, food and dining decisions, career planning, personal organization, or similar workflows. — Candidates who want to be more competitive for this and future opportunities are encouraged to proactively spend time learning MCP and using LLM plugins/connectors before applying. — Why This Work Matters — LLMs are quickly becoming personal assistants for everyday decisions, but truly useful AI needs to do more than produce generic advice. It needs to understand context, preferences, constraints, tradeoffs, and what success looks like in real life. — Your evaluations will help improve how AI systems support people with practical, high-context tasks across food, health, travel, productivity, careers, and life organization. This work directly contributes to making AI assistants more personalized, trustworthy, and useful for real-world personal workflows. — Engagement Details • Expected commitment: 20+ hours/week • Ramp-up: 1–2 days required • Turnaround expectation: Ability to complete tasks within 24 hours • Equipment: Desktop or laptop required; Chromebooks are not supported • Experts added to the project will begin in a trial period to assess project fit, quality, and consistency before being considered for ongoing tasking. • Please note: This project is still in its early stages, so there may be an initial delay before tasking begins.

$50 - $190 / hourOpen / Referral verified
STEM / Remote

Professional Design Experts

Overview — Mercor is seeking Professional Design Experts to work on a research project for one of the world’s top AI companies. This project involves leveraging your professional experience to make decisions about product design and taste preferences. — Note: Candidates must have native or near-native proficiency in English — Ideal Applicants Will Have • Worked in roles that require good taste in visual presentation • Designed graphics, documents, or other files for professional purposes • Experience with design tools such as Figma, Sketch, or Adobe Creative Suite — Role Specifics • Must be able to commit a minimum of 15 hours per week • Should be proficient in Slides, Sheets, Docs, and PDFs — Eligibility • Based in the United States, United Kingdom, or Canada • * * — Equal Opportunity — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request. • * *

$80 - $180 / hourOpen / Referral verified
Legal / Remote

Clinical Law Professor / Clinic Director

Mercor is seeking experienced Clinical Law Professors and Clinic Directors to evaluate AI-generated legal reasoning involving civil legal services. Experts will combine doctrinal knowledge, legal pedagogy, and practical supervision experience to assess AI-generated legal analyses. • * * — Responsibilities • Evaluate AI-generated legal responses. • Review issue spotting, legal reasoning, and procedural analysis. • Assess scenarios involving housing, family, consumer, and debt law. • Provide structured feedback to improve AI systems. • Apply experience supervising law students and live-client clinics. • * * — Required Qualifications • Juris Doctor (JD). • Current or previous U.S. bar admission. • Experience directing or supervising a law school clinic. • Expertise in one or more civil legal practice areas. • Strong legal writing and evaluation skills. • * * — Preferred Qualifications • Clinical Professor of Law or Clinic Director. • Experience with housing, family, consumer, elder, or disability law clinics. • Experience developing legal curriculum or assessment rubrics. • Published scholarship or educational materials. • * * — Why Join Mercor? • Help build more accurate AI for legal education and practice. • Bring your clinical teaching expertise to emerging legal technologies. • Short-term, intellectually engaging project. • Competitive compensation. • * * — Engagement Details — Location: Remote — Commitment: Approximately 15 hours/week — Project Duration: Estimated 2–3 weeks

$180 / hourOpen / Referral verified
Code / Remote — Global

CAM Programming Expert (Fusion 360)

About the role We're building a high-quality evaluation dataset for CNC manufacturing and are looking for experienced CAM programmers to author grading rubrics for Fusion 360 CAM programs. You'll create text-only, objective, verifiable rubrics that determine whether a given CAM solution would be approved — or rejected — by an expert machinist for a production run, and provide the reasoning behind each criterion. You may also review existing CAM programs and explain why they pass or fail. — What you'll do • Author 10-criterion rubrics for grading Fusion 360 CAM programs across provided CAD models and machining context (stock, machine, tooling). • Write clear verification rationale for each criterion. • Review sample CAM solutions and explain pass/fail outcomes. — Minimum requirements • Hands-on Fusion 360 CAM experience, with access to a Fusion 360 license. • Ability to program 4- and/or 5-axis CNC toolpaths. • 5+ years of CAM programming experience. • Strong production judgment — able to tell what would and wouldn't run safely on a real machine. • Written English fluency (all deliverables are text-based). — Preferred • CAM instructor or teaching experience.

$120 - $175 / hourOpen / Referral verified
Language / Remote — Global

CNC Machining Expert

About the role We're building a high-quality evaluation dataset for CNC manufacturing and are looking for experienced CNC machinists to help author and validate grading rubrics for CNC machining work. You'll bring real production-floor judgment to determine whether a machining approach would be approved or rejected for a production run, and clearly explain your reasoning. — What you'll do • Contribute to text-only, objective, verifiable rubrics for grading CNC machining solutions across provided CAD models and machining context (stock, machine, tooling). • Write clear verification rationale for each criterion. • Review sample solutions and explain pass/fail outcomes. — Minimum requirements • 5+ years of hands-on CNC machining experience. • Strong production judgment — able to tell what would and wouldn't run safely on a real machine. • Written English fluency (all deliverables are text-based). — Preferred • Experience with hobbyist machines (Makera Carvera) and Haas machines (VF-2).

$120 - $175 / hourOpen / Referral verified
Medical / Remote — US-based or deep US-market experience

Epidemiologist — Patient Population Sizing

Mercor is partnering with a biotech and pharma research team on expert human-evaluation work — creating and critiquing rubrics that analyze commercial drugs and drug-development programs. We're seeking an epidemiologist for a paid pilot, with the potential to continue as the work scales. — Responsibilities • Create and critique rubrics that analyze commercial drugs and drug-development programs • Size the patient population for a given drug or indication (prevalence, incidence, diagnosed → treated → addressable) • Judge whether a patient-population estimate is sound and well-sourced, and flag over- or under-counting • Apply this across the pilot's set of drugs and indications — Requirements • US-based epidemiologist • Familiarity with biotech/pharma development programs • Experience turning real-world data (claims, EHR, registries) into patient funnels • 5+ years practising epidemiology (graduate degree in epidemiology, biostatistics, or public health) — Engagement Details • Duration: Approximately 10–20 hours over a 1–2 week pilot • Remote — US-based — Why Participate • Help improve expert AI evaluation of drugs and drug-development programs • Recurring engagement for strong contributors

$150 - $175 / hourOpen / Referral verified
STEM / Remote

Pro Bono Counsel (Access to Justice Expert)

Mercor is seeking experienced Pro Bono Counsel to evaluate AI-generated legal analyses involving civil legal services and access-to-justice issues. Experts will leverage their broad civil practice experience to assess legal reasoning and provide structured feedback to improve AI quality. • * * — Responsibilities • Review AI-generated legal analyses. • Evaluate housing, family, consumer, foreclosure, and debt scenarios. • Assess legal reasoning and procedural guidance. • Identify legal inaccuracies and access-to-justice considerations. • Deliver structured written evaluations. • * * — Required Qualifications • Juris Doctor (JD). • Active U.S. bar admission. • Experience managing or participating in pro bono legal initiatives. • Strong background in civil legal practice. • Excellent legal writing and analytical skills. • * * — Preferred Qualifications • Experience as Pro Bono Counsel or Pro Bono Partner. • Experience coordinating legal clinics or volunteer attorneys. • Expertise across multiple civil legal practice areas. • Experience with legal technology or AI. • * * — Why Join Mercor? • Help improve AI systems supporting legal reasoning and access to justice. • Apply broad civil legal expertise to impactful AI research. • Collaborate with leading AI organisations. • Competitive hourly compensation.

$170 / hourOpen / Referral verified
Math / Remote

Data Science Expert

1. Role Overview — Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications. — 2. Key Responsibilities • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations) • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score • Apply consistent, evidence-based judgment so that scores are reproducible and defensible • Incorporate structured feedback from senior reviewers and iterate quickly on your work — 3. Ideal Qualifications • 5+ years of professional data science experience in industry • Background in business operations, product, or growth data science at top-tier technology companies • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders • Exceptionally strong written communication • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers • Prior experience with AI training, evaluation, or human-data projects is a strong plus — 4. Application Process • Submit your resume or relevant technical background to get started • Qualified applicants may be asked to complete a brief technical assessment or submit additional information

$120 - $170 / hourOpen / Referral verified
Finance / Remote

Revenue-cycle Executive (VP/Sr. Director Revenue Cycle, or RCM-focused Finance Leader)

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Revenue Cycle Executives — including VP of Revenue Cycle, Sr. Director of Revenue Cycle, and RCM-focused finance leaders — to evaluate AI tools designed to transform end-to-end revenue cycle performance. Your strategic expertise in revenue cycle operations, payer contracting, financial performance management, and team leadership will directly shape AI systems that drive sustainable financial improvement across healthcare organizations. — Responsibilities • Provide executive-level oversight and strategic direction for end-to-end revenue cycle operations, including patient access, coding, billing, denials, and collections. • Evaluate AI-generated revenue cycle performance analyses, strategic recommendations, and operational improvement plans for accuracy and feasibility. • Assess AI tools designed to automate revenue cycle workflows, identify revenue leakage, and optimise payer reimbursement. • Define and monitor enterprise revenue cycle KPIs, including net collection rate, days in A/R, denial rate, cost-to-collect, and cash collections performance. • Develop and implement revenue cycle improvement strategies aligned with organisational financial goals. • Oversee managed care contracting strategy and payer relationship management. • Collaborate with clinical, finance, compliance, and IT leadership to align revenue cycle initiatives with organisational priorities. • Present revenue cycle performance to executive leadership, boards, and investors. • Annotate AI outputs and provide structured strategic feedback to support AI training datasets. — Requirements • 10+ years of progressive revenue cycle leadership experience, with at least 5 years in a VP, Sr. Director, or equivalent executive role. • Comprehensive knowledge of end-to-end revenue cycle operations, including front-end, mid-cycle, and back-end functions. • Strong financial acumen with experience managing P&L, budget oversight, and financial performance reporting. • Track record of leading large-scale revenue cycle transformation initiatives in health system, hospital, or physician group settings. • Deep familiarity with managed care contracting, payer relations, and reimbursement optimisation. • Exceptional written and verbal English communication skills, including executive-level presentation skills. • Comfortable working independently in a fully remote environment. • Ability to synthesise complex revenue cycle performance data and AI-generated outputs into strategic recommendations. — Preferred Qualifications • HFMA Fellow (FHFMA), CHFP, or CRCR certification. • Experience with RCM technology platforms and large-scale EHR implementations. • Background in PE-backed healthcare organizations, health systems, or large physician enterprise environments. • Familiarity with AI tools and comfort evaluating AI-generated revenue cycle strategy content. • Experience leading multi-site or multi-entity revenue cycle consolidation or integration initiatives. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Flexible, remote-first work environment. • Competitive compensation at $162/hr. • Gain exposure to cutting-edge AI workflows across the full revenue cycle enterprise. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$162 / hourOpen / Referral verified
Finance / Remote

Accounting Expert

1. Role Overview — Mercor is partnering with a leading AI research organization to engage experienced accountants for a project focused on evaluating how well AI systems perform real-world accounting work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications. — 2. Key Responsibilities • Design precise, task-specific grading criteria for real-world accounting deliverables (reconciliations, close packages, financial statements and disclosures, workpapers, technical accounting memos) • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score • Apply consistent, evidence-based judgment so that scores are reproducible and defensible • Incorporate structured feedback from senior reviewers and iterate quickly on your work — 3. Ideal Qualifications • 5+ years of professional accounting experience • Background in audit, assurance, or advisory at leading global accounting firms, or in senior accounting roles at large public companies • CPA or equivalent professional certification strongly preferred • Deep fluency in the day-to-day craft: US GAAP application, month-end close, account reconciliations, financial statement preparation and disclosures, and audit-ready workpapers • Exceptionally strong written communication, with the ability to explain exactly why a piece of work does or does not meet professional standards • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers • Prior experience with AI training, evaluation, or human-data projects is a strong plus — 4. Application Process • Submit your resume or relevant professional background to get started • Qualified applicants may be asked to complete a brief technical assessment or submit additional information

$130 - $160 / hourOpen / Referral verified
Legal / France (remote)

In-House Counsel (France)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of French law. We are hiring experienced in-house counsel to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • 5+ years as in-house legal counsel practising in France • Strong command of French law, especially commercial and contract matters • Sharp attention to detail and precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$150 - $160 / hourOpen / Referral verified
Language / United Kingdom (remote)

In-House Counsel (United Kingdom)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of UK law. We are hiring experienced in-house counsel to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • 5+ years as in-house legal counsel practising in the United Kingdom • Strong command of UK law, especially commercial and contract matters • Sharp attention to detail and precise written English • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$150 - $160 / hourOpen / Referral verified
Legal / France (remote)

Litigation Lawyer (France)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of French law. We are hiring experienced litigation & disputes lawyers to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • Primarily law-firm experience — at least 2 years (3+ preferred) at a law firm in litigation, disputes, or contentious practice. If you are currently in-house, that is fine only if you moved in-house within the last 2 years and your prior experience was at a law firm. • Qualified to practise in France (avocat/e), with a strong command of French litigation and procedural law • Native or fluent French, with precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$150 - $160 / hourOpen / Referral verified
Legal / France (remote)

Corporate/M&A Lawyer (France)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of French law. We are hiring experienced corporate/M&A lawyers to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • Primarily law-firm experience — at least 2 years (3+ preferred) at a law firm in corporate/M&A or transactional practice. If you are currently in-house, that is fine only if you moved in-house within the last 2 years and your prior experience was at a law firm. • Qualified to practise in France, with a strong command of French corporate and commercial law • Native or fluent French, with precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$150 - $160 / hourOpen / Referral verified
Legal / United Kingdom (remote)

Corporate/M&A Lawyer (United Kingdom)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of UK law. We are hiring experienced corporate/M&A lawyers to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • Primarily law-firm experience — at least 2 years (3+ preferred) at a law firm in corporate/M&A or transactional practice. If you are currently in-house, that is fine only if you moved in-house within the last 2 years and your prior experience was at a law firm. • Qualified to practise in the United Kingdom, with a strong command of UK corporate and commercial law • Native or fluent English, with precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$150 - $160 / hourOpen / Referral verified
Business / Remote

Agency Brand Design Expert

1. Role Overview — Mercor is partnering with a leading AI research organization to engage experienced brand designers for a project focused on evaluating how well AI systems perform real-world brand and design work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications. — 2. Key Responsibilities • Design precise, task-specific grading criteria for real-world brand design deliverables (brand identity systems and guidelines, logos and visual identity, campaign and marketing visuals, packaging concepts, presentation and layout design, digital and social assets, and brand messaging and positioning) • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score • Apply consistent, evidence-based judgment so that scores are reproducible and defensible • Incorporate structured feedback from senior reviewers and iterate quickly on your work — 3. Ideal Qualifications • 5+ years of professional experience in brand, visual, or graphic design • Background specifically at world-class independent brand and design agencies (e.g., Pentagram, Wolff Olins, Landor, Collins, IDEO) — agency experience, not in-house or enterprise design teams • Experience in one of the core brand disciplines: Creative Director, Brand Designer (logo, visual identity, supporting graphics), Web Designer (UI/UX, web architecture), or Graphic Designer • Deep fluency in the day-to-day craft: typography, layout and visual hierarchy, color, composition, logo and brand identity systems, and articulating critique the way a creative director would in a review • Exceptionally strong written communication, with the ability to explain exactly why a piece of work does or does not meet the bar • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers • Experience with copywriting and brand messaging — including audience research, competitor research, and positioning — is a strong plus • Prior experience with AI training, evaluation, or human-data projects is a strong plus — 4. Application Process • Submit your resume, portfolio, background to get started • Qualified applicants may be assessment or submit additional information

$80 - $150 / hourOpen / Referral verified
Code / Remote

Sales Engineering Expert

1. Role Overview — Mercor is partnering with a leading AI research organization to engage experienced sales engineers for a project focused on evaluating how well AI systems perform real-world technical sales work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications. — 2. Key Responsibilities • Design precise, task-specific grading criteria for real-world pre-sales and solutions-engineering deliverables (technical discovery plans, tailored product demonstrations, proof-of-concept and technical evaluation plans, security questionnaires and RFP responses, integration/API/architecture explanations, and product-fit and technical-risk assessments) • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score • Apply consistent, evidence-based judgment so scores are reproducible and defensible • Incorporate structured feedback from senior reviewers and iterate quickly on your work — 3. Ideal Qualifications • 4+ years of relevant experience as a Sales Engineer, Solutions Engineer, Solutions Consultant, Pre-Sales Consultant, or Technical Sales Consultant • Deep fluency in the pre-sales craft: leading technical discovery, translating customer requirements into product solutions, delivering tailored demos, handling technical objections, supporting proofs of concept, and evaluating product fit and technical risk • Hands-on experience with technically complex B2B software — SaaS, cloud infrastructure, cybersecurity, data platforms, developer tools, or enterprise applications • Exceptionally strong written communication, with the ability to explain exactly why a piece of technical work would or would not land with both technical and business stakeholders • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers • Prior experience with AI training, evaluation, or human-data projects is a strong plus — 4. Application Process • Submit your resume or relevant professional background to get started • Qualified applicants may be asked to complete a brief assessment or submit additional information

$100 - $150 / hourOpen / Referral verified
Code / Remote

UI / UX Design Expert

1. Role Overview — Mercor is partnering with a leading AI research organization to engage experienced UI/UX and product designers for a project focused on evaluating how well AI systems perform real-world digital product design work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications. — 2. Key Responsibilities • Design precise, task-specific grading criteria for real-world digital product and UX deliverables (design systems and component libraries, responsive UI layouts and dashboards, end-to-end product flows such as onboarding, checkout, and booking, logged-in customer portals, UX writing and product content, and design-to-engineering/production handoff) • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score • Apply consistent, evidence-based judgment so that scores are reproducible and defensible • Incorporate structured feedback from senior reviewers and iterate quickly on your work — 3. Ideal Qualifications • 5+ years of professional experience in product design, UI/UX, or design systems for real digital products (not brand or marketing sites) • Background at leading digital-product and UX studios, in-house product-design teams at major product-led companies, or e-commerce and marketplace companies with sophisticated customer portals and dashboards • Experience across one or more relevant tracks: Product / UI / Design Systems Designer, UX Lead, UX Writer / Content Designer, or digital Product / Business Analyst • Deep fluency in the day-to-day craft: interaction and visual design, design systems, responsive layouts, end-to-end product flows, prototyping and production handoff in Figma, and articulating critique the way a design or UX lead would in a review • Exceptionally strong written communication, with the ability to explain exactly why a piece of work does or does not meet the bar for a real product and its users • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers • Prior experience with AI training, evaluation, or human-data projects is a strong plus — 4. Application Process • Submit your resume, portfolio, or relevant professional background to get started • Qualified applicants may be asked to complete a brief assessment or submit additional information

$80 - $150 / hourOpen / Referral verified
Legal / Japan (remote)

In-House Counsel (Japan)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of Japanese law. We are hiring experienced in-house counsel to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • 5+ years as in-house legal counsel practising in Japan • Strong command of Japanese law, especially commercial and contract matters • Sharp attention to detail and precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$140 - $150 / hourOpen / Referral verified
Legal / Spain (remote)

In-House Counsel (Spain)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of Spanish law. We are hiring experienced in-house counsel to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • 5+ years as in-house legal counsel practising in Spain • Strong command of Spanish law, especially commercial and contract matters • Sharp attention to detail and precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$140 - $150 / hourOpen / Referral verified
Legal / Germany (remote)

In-House Counsel (Germany)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of German law. We are hiring experienced in-house counsel to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • 5+ years as in-house legal counsel practising in Germany • Strong command of German law, especially commercial and contract matters • Sharp attention to detail and precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$140 - $150 / hourOpen / Referral verified
Code / Remote

Software Engineer — Agentic Search Systems

We're looking for engineers who have shipped production search systems (especially agentic ones now that we're in the era of LLMs and agents) and are thinking hard about the requirements and optimizations of those systems. — You've probably: • Owned relevance or retrieval on a system real users depended on • Decided how to quantify impact and improvements with these systems. • Built out systems related to agentic search. • Scaled data infrastructure to power modern search systems. — This is a 25-minute conversational interview. No coding, no take-home. We want to hear how you actually think about search quality, evaluation, and the tradeoffs you've made in real systems. Bring war stories — the messier and more specific, the better. — If your interview stands out, we'll follow up with a paid 30-minute live conversation with our team — $200 for your time, paid on completion of the call. — If that sounds like you, apply and complete the interview. We review every submission.

$80 - $150 / per-taskOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Swiss German)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Switzerland-based native Swiss German speakers to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native Swiss German speaker currently based in Switzerland • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $150 / hourOpen / Referral verified
Code / Remote

Data Scientist Talent Network

Mercor is building a network of experienced data scientists for potential future projects with leading AI research organizations. These projects may focus on evaluating how effectively AI systems perform real-world data science work. — There is no immediate project opening, but qualified applicants may be contacted as relevant opportunities become available. — 2. Potential Responsibilities — Future projects may involve: • Designing precise, task-specific grading criteria for data science deliverables, including exploratory data analyses, statistical modeling work, machine learning pipelines, experimentation and A/B test write-ups, feature engineering, and technical reports or notebooks • Evaluating AI-generated or human-created work against established criteria • Providing detailed written justifications for evaluations and scores • Applying consistent, evidence-based judgment so that assessments are reproducible and defensible • Incorporating structured feedback from senior reviewers and iterating on submitted work — Specific responsibilities will vary depending on the project. — 3. Ideal Qualifications • 1+ years of professional data science experience • Experience at a leading technology, research, or quantitative firm (such as top FAANG, AI labs, top-tier quant funds, or equivalent) • Strong command of Python, SQL, statistical modeling, machine learning, experimentation and causal inference, and translating messy real-world data into rigorous analyses • Exceptional written communication skills, including the ability to convey technical findings clearly • A detail-oriented and consistent approach to evaluating complex work • Comfort receiving feedback and calibrating judgment against established standards

$100 - $150 / hourOpen / Referral verified
Business / Remote

Management Consultant Talent Network

Mercor is building a network of experienced management consultants for potential future projects with leading AI research organizations. These projects may focus on evaluating how effectively AI systems perform real-world consulting work. — There is no immediate project opening, but qualified applicants may be contacted as relevant opportunities become available. — 2. Potential Responsibilities — Future projects may involve: • Designing precise, task-specific grading criteria for consulting deliverables, including market analyses, strategy recommendations, financial models, client-ready presentations, and implementation roadmaps • Evaluating AI-generated or human-created work against established criteria • Providing detailed written justifications for evaluations and scores • Applying consistent, evidence-based judgment so that assessments are reproducible and defensible • Incorporating structured feedback from senior reviewers and iterating on submitted work — Specific responsibilities will vary depending on the project. — 3. Ideal Qualifications • 1+ years of professional management consulting experience • Experience at a leading strategy or management consulting firm (such as McKinsey, Bain, BCG, or equivalent) • Strong command of structured problem-solving, market sizing, financial modeling, business-case development, and executive-level presentation development • Exceptional written communication skills • A detail-oriented and consistent approach to evaluating complex work • Comfort receiving feedback and calibrating judgment against established standards

$100 - $150 / hourOpen / Referral verified
Legal / Germany (remote)

Corporate/M&A Lawyer (Germany)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of German law. We are hiring experienced corporate/M&A lawyers to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • Primarily law-firm experience — at least 2 years (3+ preferred) at a law firm in corporate/M&A or transactional practice. If you are currently in-house, that is fine only if you moved in-house within the last 2 years and your prior experience was at a law firm. • Qualified to practise in Germany, with a strong command of German corporate and commercial law • Native or fluent German, with precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$140 - $150 / hourOpen / Referral verified
Legal / Spain (remote)

Corporate/M&A Lawyer (Spain)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of Spanish law. We are hiring experienced corporate/M&A lawyers to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • Primarily law-firm experience — at least 2 years (3+ preferred) at a law firm in corporate/M&A or transactional practice. If you are currently in-house, that is fine only if you moved in-house within the last 2 years and your prior experience was at a law firm. • Qualified to practise in Spain, with a strong command of Spanish corporate and commercial law • Native or fluent Spanish, with precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$140 - $150 / hourOpen / Referral verified
Finance / Bay Area, CA

Investment Banking & M&A - Finance Domain Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — We are hiring a senior finance domain expert to work directly with a leading AI lab's research and program management teams, improving how frontier AI models reason about real financial work. — Your finance expertise is the substance of this role. You will review the quality of finance knowledge work tasks, write the instruction specs and golden solutions that define what "correct" looks like, and build the benchmarks that show whether the model is genuinely improving. We are looking for a practising specialist rather than a generalist. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. You will be provisioned with client-issued accounts and equipment, and will work inside the client's own tools alongside their research teams. — Location: This is a hybrid role based in the Bay Area, California. You must live in the Bay Area and work on-site with the client's team multiple days each week, when required. This is not a remote role. If you do not currently live in the Bay Area, you must be willing to relocate there at your own cost before the engagement starts — relocation assistance is not provided. — 2. Key Responsibilities • Data QA and reviews: Vet the quality of finance knowledge work tasks and model outputs — spotting missing behaviors, thin reasoning, flawed assumptions, and answers that read well but would not survive professional scrutiny. • Instruction specs and golden datasets: Write high-quality instruction specs, produce golden solutions to financial problems, and define new finance tasks that reflect how the work is actually done in practice. • Benchmarks and domain depth: Design challenging finance tasks and evaluation sets, and help build finance-specific skills and tools together with the research team. • Calibration: Work with client researchers and specialists in adjacent fields to keep standards consistent, translating tacit financial judgment into explicit, teachable criteria. — 3. Core Qualifications • Experience: 5+ years of substantive, dedicated professional finance experience at a recognized institution — for example an investment bank, asset manager, private equity or credit fund, Big Four firm, a large corporate finance function, or a financial regulator. Generalist roles that only touch finance peripherally do not count. • Domain depth: Genuine specialization in at least one core finance discipline, for example corporate finance and FP&A, investment banking and M&A, asset or wealth management, private equity or private credit, quantitative finance and risk management, treasury, or accounting and audit. • Seniority: Clear progression to a senior individual-contributor or leadership level — for example Vice President, Director, Principal, Managing Director, Portfolio Manager, Controller, or CFO — with real ownership of analysis and decisions. • Education and credentials: An advanced degree from a strong program (MBA, MS, or PhD in finance, economics, accounting, or a quantitative field) and/or a recognized professional credential such as CFA, CPA, FRM, or an actuarial designation. Strongly preferred. • AI fluency: Hands-on working use of large language models in your professional work, and the judgment to tell a well-reasoned answer from a plausible-sounding wrong one. • Availability: Able to commit reliably to 40 hours per week for an initial engagement of 6 months. • Location: Living in the Bay Area, California, and able to work on-site with the client's team multiple days each week, when required. Candidates not currently based in the Bay Area must be willing to relocate there at their own cost; relocation assistance is not provided. • Excellent written communication, and the ability to give precise, well-structured written feedback. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$100 - $150 / hourOpen / Referral verified
Legal / Bay Area, CA

Partner & General Counsel - Legal Domain Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — We are hiring a senior legal domain expert to work directly with a leading AI lab's research and program management teams, improving how frontier AI models reason about real legal work. — Your legal expertise is the substance of this role. You will review the quality of legal knowledge work tasks, write the instruction specs and golden solutions that define what "correct" looks like, and build the benchmarks that show whether the model is genuinely improving. We are looking for a practising specialist rather than a generalist. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. You will be provisioned with client-issued accounts and equipment, and will work inside the client's own tools alongside their research teams. — Location: This is a hybrid role based in the Bay Area, California. You must live in the Bay Area and work on-site with the client's team multiple days each week, when required. This is not a remote role. If you do not currently live in the Bay Area, you must be willing to relocate there at your own cost before the engagement starts — relocation assistance is not provided. — 2. Key Responsibilities • Data QA and reviews: Vet the quality of legal knowledge work tasks and model outputs — spotting missing behaviors, thin reasoning, and answers that read well but would not survive professional scrutiny. • Instruction specs and golden datasets: Write high-quality instruction specs, produce golden solutions to legal problems, and define new legal tasks that reflect how the work is actually done in practice. • Benchmarks and domain depth: Design challenging legal tasks and evaluation sets, and help build legal-specific skills and tools together with the research team. • Calibration: Work with client researchers and specialists in adjacent fields to keep standards consistent, translating tacit legal judgment into explicit, teachable criteria. — 3. Core Qualifications • Education: Juris Doctor (JD) from an accredited law school; a highly ranked school is strongly preferred. • Experience: 5+ years of substantive post-qualification legal practice at a reputable institution — an established law firm, a corporate legal department, a regulatory body, or a court. Internships and clerkships alone do not count. • Domain depth: Genuine specialization in at least one substantive practice area, for example corporate and transactional, litigation and dispute resolution, regulatory and compliance, intellectual property, employment and labor, or tax. • Seniority: Clear progression to a senior level — partner, of counsel, counsel, senior associate, senior in-house counsel, or general counsel — with real ownership of matters. • Licensure: Active admission to at least one U.S. state bar, in good standing. • AI fluency: Hands-on working use of large language models in your professional work, and the judgment to tell a well-reasoned answer from a plausible-sounding wrong one. • Availability: Able to commit reliably to 40 hours per week for an initial engagement of 6 months. • Location: Living in the Bay Area, California, and able to work on-site with the client's team multiple days each week, when required. Candidates not currently based in the Bay Area must be willing to relocate there at their own cost; relocation assistance is not provided. • Excellent written communication, and the ability to give precise, well-structured written feedback. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$100 - $150 / hourOpen / Referral verified
Finance / Bay Area, CA

Private Equity & Venture Capital - Finance Domain Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — We are hiring a senior finance domain expert to work directly with a leading AI lab's research and program management teams, improving how frontier AI models reason about real financial work. — Your finance expertise is the substance of this role. You will review the quality of finance knowledge work tasks, write the instruction specs and golden solutions that define what "correct" looks like, and build the benchmarks that show whether the model is genuinely improving. We are looking for a practising specialist rather than a generalist. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. You will be provisioned with client-issued accounts and equipment, and will work inside the client's own tools alongside their research teams. — Location: This is a hybrid role based in the Bay Area, California. You must live in the Bay Area and work on-site with the client's team multiple days each week, when required. This is not a remote role. If you do not currently live in the Bay Area, you must be willing to relocate there at your own cost before the engagement starts — relocation assistance is not provided. — 2. Key Responsibilities • Data QA and reviews: Vet the quality of finance knowledge work tasks and model outputs — spotting missing behaviors, thin reasoning, flawed assumptions, and answers that read well but would not survive professional scrutiny. • Instruction specs and golden datasets: Write high-quality instruction specs, produce golden solutions to financial problems, and define new finance tasks that reflect how the work is actually done in practice. • Benchmarks and domain depth: Design challenging finance tasks and evaluation sets, and help build finance-specific skills and tools together with the research team. • Calibration: Work with client researchers and specialists in adjacent fields to keep standards consistent, translating tacit financial judgment into explicit, teachable criteria. — 3. Core Qualifications • Experience: 5+ years of substantive, dedicated professional finance experience at a recognized institution — for example an investment bank, asset manager, private equity or credit fund, Big Four firm, a large corporate finance function, or a financial regulator. Generalist roles that only touch finance peripherally do not count. • Domain depth: Genuine specialization in at least one core finance discipline, for example corporate finance and FP&A, investment banking and M&A, asset or wealth management, private equity or private credit, quantitative finance and risk management, treasury, or accounting and audit. • Seniority: Clear progression to a senior individual-contributor or leadership level — for example Vice President, Director, Principal, Managing Director, Portfolio Manager, Controller, or CFO — with real ownership of analysis and decisions. • Education and credentials: An advanced degree from a strong program (MBA, MS, or PhD in finance, economics, accounting, or a quantitative field) and/or a recognized professional credential such as CFA, CPA, FRM, or an actuarial designation. Strongly preferred. • AI fluency: Hands-on working use of large language models in your professional work, and the judgment to tell a well-reasoned answer from a plausible-sounding wrong one. • Availability: Able to commit reliably to 40 hours per week for an initial engagement of 6 months. • Location: Living in the Bay Area, California, and able to work on-site with the client's team multiple days each week, when required. Candidates not currently based in the Bay Area must be willing to relocate there at their own cost; relocation assistance is not provided. • Excellent written communication, and the ability to give precise, well-structured written feedback. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$110 - $150 / hourOpen / Referral verified
Legal / Japan (remote)

Corporate/M&A Lawyer (Japan)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of Japanese law. We are hiring experienced corporate/M&A lawyers to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • Primarily law-firm experience — at least 2 years (3+ preferred) at a law firm in corporate/M&A or transactional practice. If you are currently in-house, that is fine only if you moved in-house within the last 2 years and your prior experience was at a law firm. • Qualified to practise in Japan, with a strong command of Japanese corporate and commercial law • Native or fluent Japanese, with precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$140 - $150 / hourOpen / Referral verified
Medical / Remote

Child & Adolescent Mental Health Clinical Advisor (AI Safety Benchmark Project)

We are partnering with a leading AI research organisation to develop a clinician-informed benchmark for evaluating how AI companion chatbots respond to adolescents experiencing mental health challenges. As AI companions become increasingly common sources of emotional support, there is a growing need for rigorous, clinically grounded evaluation of how these systems handle sensitive situations such as suicide risk, therapeutic guidance, and emotional dependency. We're seeking experienced Child & Adolescent Psychiatrists, Clinical Psychologists, and Mental Health Researchers to contribute their expertise in designing realistic clinical scenarios and evaluation standards for this important AI safety initiative. • * * — Responsibilities • Author fictionalised adolescent mental health case scenarios based on real-world clinical expertise. • Review scenarios for clinical realism, developmental appropriateness, and ethical considerations. • Help develop and refine clinician-informed evaluation rubrics for AI chatbot responses. • Review a subset of AI-generated conversations and provide expert judgment to calibrate automated evaluation systems. • Advise on best practices for assessing AI behaviour across key mental health domains, including: • Suicide & Self-Harm (risk recognition, safety planning, crisis referral) • Medical & Therapeutic Impersonation (diagnosis, treatment advice, discouraging professional care) • Parasocial Attachment & Anthropomorphism (claims of consciousness, emotional dependency, discouraging trusted relationships) • * * — Requirements — We are looking for clinicians and researchers with substantial expertise in child and adolescent mental health. — Required Qualifications • MD, DO, PhD, PsyD, or equivalent qualification in Psychiatry, Psychology, or a related mental health field. • Clinical or research experience focused on children and adolescents. • Expertise in one or more of the following: • Suicide prevention • Self-harm assessment • Child & adolescent psychiatry • Clinical child psychology • Digital mental health • Evidence-based therapies such as DBT or CBT • Strong written communication skills and ability to provide structured clinical feedback. • Comfortable reviewing fictional clinical scenarios and evaluating AI-generated conversations. — Preferred Qualifications • Experience conducting suicide risk assessments or crisis intervention. • Academic or published research in adolescent mental health, suicide prevention, or digital therapeutics. • Experience developing clinical guidelines, assessment frameworks, or educational materials. • Prior work involving AI, digital health technologies, or mental health product evaluation. • Experience supervising trainees or participating in multidisciplinary clinical teams. • * * — Role Details • Commitment: Approximately 10–15 hours per week. • Compensation: $80 - $150 USD per hour • * * — Why Join? • Help shape the future of safe AI systems for adolescents. • Apply your clinical expertise to one of the most important emerging questions in AI safety and digital mental health. • Collaborate with researchers developing rigorous evaluation standards for AI companion chatbots. • Contribute to work that may influence future AI research, independent benchmarking, and responsible deployment of conversational AI.

$80 - $150 / hourOpen / Referral verified
Legal / Japan (remote)

Litigation Lawyer (Japan)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of Japanese law. We are hiring experienced litigation & disputes lawyers to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • Primarily law-firm experience — at least 2 years (3+ preferred) at a law firm in litigation, disputes, or contentious practice. If you are currently in-house, that is fine only if you moved in-house within the last 2 years and your prior experience was at a law firm. • Qualified to practise in Japan (bengoshi, or registered foreign lawyer / gaikokuho jimu bengoshi), with a strong command of Japanese litigation and procedural law • Native or fluent Japanese, with precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$140 - $150 / hourOpen / Referral verified
Finance / Remote

Utilisation Management / Case Management leader (RN/Physician-advisor)

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Utilisation Management and Case Management leaders — including RN UM directors, case management managers, and physician advisors — to evaluate AI tools designed to enhance clinical review, medical-necessity determination, and care-coordination workflows. Your clinical and operational expertise will directly shape AI systems that improve utilisation performance, reduce unnecessary care, and support value-based care objectives. — Responsibilities • Lead utilisation management and/or case management operations, including concurrent review, retrospective review, discharge planning, and care coordination. • Evaluate AI-generated medical necessity determinations, clinical review outputs, and UM decision-support recommendations for accuracy and clinical appropriateness. • Conduct and oversee clinical reviews against InterQual, MCG, or Milliman care guidelines to support admission, continued stay, and level-of-care determinations. • Manage the physician advisor program to support complex UM cases, peer-to-peer review requests, and denial appeals. • Coordinate with clinical teams, payers, and post-acute providers to facilitate appropriate care transitions and discharge planning. • Monitor UM/CM KPIs, including avoidable days, denial rates, observation vs. inpatient conversion rates, and readmission rates. • Ensure compliance with CMS Conditions of Participation, Two-Midnight Rule, and payer-specific UM requirements. • Collaborate with revenue cycle, compliance, and clinical leadership to align UM performance with financial and quality objectives. • Annotate AI outputs and provide structured clinical feedback to support AI training datasets. — Requirements • 5+ years of experience in utilisation management, case management, or clinical review, with at least 2 years in a leadership role. • Active clinical licensure (RN required; physician advisor/MD preferred) with strong medical necessity review expertise. • Deep knowledge of InterQual, MCG, or Milliman clinical review criteria. • Expertise in CMS Two-Midnight Rule, observation status regulations, and inpatient criteria. • Experience managing physician advisor programs and peer-to-peer review processes with payers. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to critically evaluate clinical documentation and AI-generated UM outputs. — Preferred Qualifications • CPUR (Certified Professional in Utilisation Review), ACM (Accredited Case Manager), or CCM (Certified Case Manager) credential. • MD or DO with UM/physician advisor experience preferred for physician advisor roles. • Experience with UM software platforms (Allscripts, Utilisation Management software, or equivalent). • Familiarity with AI tools and comfort evaluating AI-generated clinical review content. • Background in health system, ACO, or value-based care organisation UM programs. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in utilisation management and clinical decision-support. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$100 - $150 / hourOpen / Referral verified
Medical / Remote

Psychiatry Expert

Psychiatrist — Mercor is seeking board-certified psychiatrists for a project with one of the world's top AI labs. — We welcome applications from psychiatrists across all practice settings, with particular interest in subspecialty expertise. This role evaluates how large language models reason through advanced psychiatric cases. You will design expert-level diagnostic and reasoning challenges that current models fail, then provide authoritative, gold-standard answers. — Required • Completed psychiatry residency: ACGME, AOA, or national equivalent. In-training does not count. • Medical degree: MD or DO. DO and international degrees qualify. • Active board certification in psychiatry: ABPN, FRCPC, MRCPsych, FRANZCP, or recognized national equivalent. Board-eligible and lapsed do not count. • Active, unrestricted medical license. No suspension, revocation, or open disciplinary action. • 3+ years of independent practice after residency, currently or within the last 2 years. Training years excluded. • Practice in an accredited or government-licensed setting: hospital, academic center, community clinic, private practice, or telehealth. Recognized means accredited, not self-described. • Based in the US, UK, Europe, or Australia. US preferred. — Preferred (scored, not gates) • 10+ direct patient-contact hours per week, current. Clinical care only. • Psychiatry-Specific Subspecialty depth in Addiction, Child/Adolescent, Forensic, Geriatric, or Consultation-Liaison. Demonstrated through fellowship, subspecialty board, or 2+ years of practice. • Must be a psychiatry focus • OSCE development or examiner experience. Sitting one as a candidate does not count. • Faculty appointment in a department of psychiatry, current or within 2 years. — Role details: • Develop gold-standard answers with clear clinical reasoning • Emphasize edge cases, nuanced differentials, and risk assessment • Collaborate directly with research teams at a leading AI lab • Provide feedback on common failure modes in model reasoning — Screening Process: • Complete a short interview (approximately 20–30 minutes) • The interview explores your clinical background and subspecialty expertise

$150 / hourOpen / Referral verified
Legal / Remote

Law Experts

Role Overview • Mercor is seeking senior legal professionals to build evaluation tasks for AI systems operating in Fortune 500 legal contexts. • The workflows are calibrated to the regulatory complexity, transaction scale, and risk posture of Fortune 500 and large public companies. • Contributors design realistic enterprise legal scenarios, draft model-grade reference work, and write rubrics that distinguish strong legal reasoning from surface-level output. — Key Responsibilities • Construct enterprise legal scenarios across in-house counsel, regulatory, transactional, and litigation work for F500-scale organizations. • Draft and review materials grounded in enterprise compliance regimes (FCPA, SOX, GDPR, HIPAA, sector-specific regulation in banking, healthcare, pharma, and insurance). • Build tasks around enterprise contracting at material scale (MSAs, SOWs, DPAs, licensing) and M&A diligence for F500 transactions. • Develop ERM and COSO-aligned risk scenarios, board-level reporting, and regulated-industry investigations. • Author detailed, criterion-referenced rubrics that capture the judgment a senior F500 lawyer would apply. — Ideal Qualifications • 2+ years as in-house counsel at a Fortune 500 or large public company, or as a partner / senior associate at an AmLaw 100 firm representing F500 clients. • Deep exposure to one or more enterprise practice areas: M&A, securities, regulatory, employment, IP, privacy, or commercial contracting at F500 scale. • Familiarity with CLM platforms (Ironclad, Agiloft, Icertis) and enterprise GRC tooling. • JD with active bar admission. Prior rubric, exam, or training-content authorship is a plus. — Compensation Note • Hourly Pay: $110 to $150 per hour, set by Mercor based on demonstrated expertise. • Minimum Commitment: 20 hours per week. • Onboarding via the Mercor Rubric Academy, a paid program that calibrates contributors to the quality bar before live work. • Advancement: strong contributors move into reviewer, lead, and domain SME roles with elevated rates.

$110 - $150 / hourOpen / Referral verified
Finance / Remote

Investment Banking Expert

Mercor is recruiting U.S./UK/Canada/Europe/Australia-based Investment Banking Experts for a research project with a leading foundational model AI lab. — You are a good fit if you: • Have at least 2 years of experience working at top firms in investment banking and experience in at least one of the following • Financial Modeling • Pitch Decks • Investment/Analysis Summaries and Memos • Company/Industry Analysis — Here are more details about the role: • You must be able to commit at least 10 hours per week for this role • This is a minimum four week engagement, with potential for significant extension or rotation to similar, future projects • Successful contributions increase the odds that you are selected on future projects with Mercor • This role will pay between $100-$130/hour with potential for increases for top performers

$100 - $130 / hourOpen / Referral verified
Business / Remote

Marketing Expert

1. Role Overview — Mercor is partnering with a leading AI research organization to engage experienced marketing professionals for a project focused on evaluating how well AI systems perform real-world marketing work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications. — 2. Key Responsibilities • Design precise, task-specific grading criteria for real-world marketing deliverables (creative briefs, campaign concepts and decks, brand positioning and messaging frameworks, ad copy and scripts, integrated campaign plans, client-ready presentations) • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score • Apply consistent, evidence-based judgment so that scores are reproducible and defensible • Incorporate structured feedback from senior reviewers and iterate quickly on your work — 3. Ideal Qualifications • 5+ years of professional experience in marketing or advertising • Background as a creative director, account lead, or brand strategist at top-tier global creative agencies, or in senior brand marketing roles at leading consumer companies • Deep fluency in the day-to-day craft: writing and evaluating creative briefs, campaign development across channels, brand positioning and messaging, and judging creative work against a brief and a business objective • Exceptionally strong written communication, with the ability to explain exactly why a piece of work does or does not meet the bar • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers • Prior experience with AI training, evaluation, or human-data projects is a strong plus • No specialized media-planning or measurement background is required for this role — 4. Application Process • Submit your resume or relevant professional background to get started • Qualified applicants may be asked to complete a brief assessment or submit additional information

$80 - $150 / hourOpen / Referral verified
Business / Remote

B2B Sales Expert

1. Role Overview — Mercor is partnering with a leading AI research organization to engage experienced B2B sales professionals for a project focused on evaluating how well AI systems perform real-world sales work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications. — 2. Key Responsibilities • Design precise, task-specific grading criteria for real-world sales deliverables (prospecting sequences and outreach emails, discovery and demo call plans, account and territory plans, proposals, mutual action plans, forecast and pipeline reviews) • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score • Apply consistent, evidence-based judgment so that scores are reproducible and defensible • Incorporate structured feedback from senior reviewers and iterate quickly on your work — 3. Ideal Qualifications • 5+ years of professional B2B sales experience, including quota-carrying roles • Background as an Account Executive or Sales Development Representative selling into large enterprises • Deep fluency in the day-to-day craft: outbound prospecting and sequencing, discovery and qualification, multi-stakeholder deal management, pipeline hygiene and forecasting, and structured sales methodologies • Exceptionally strong written communication, with the ability to explain exactly why a piece of work would or would not land with a real buyer • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers • Prior experience with AI training, evaluation, or human-data projects is a strong plus — 4. Application Process • Submit your resume or relevant professional background to get started • Qualified applicants may be asked to complete a brief assessment or submit additional information

$100 - $150 / hourOpen / Referral verified
Finance / Remote

Patient Financial Clearance Leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Patient Financial Counselling and Financial Clearance leaders to support the evaluation of AI tools designed to enhance financial clearance operations and patient financial assistance workflows. Your expertise in charity care, self-pay collections, financial screening, and patient financial advocacy will help shape AI systems that improve patient access while optimising revenue capture. — Responsibilities • Lead patient financial counselling and financial clearance operations including charity care screening, financial assistance applications, and self-pay resolution. • Evaluate AI-generated financial counselling recommendations, eligibility screening outputs, and patient communication drafts for accuracy and appropriateness. • Screen patients for Medicaid eligibility, charity care qualification, and financial assistance program enrollment. • Counsel patients on financial obligations, payment plan options, and available assistance programs. • Coordinate with social work, case management, and billing teams to address complex patient financial situations. • Develop and maintain SOPs for financial clearance workflows, including pre-service financial screening and point-of-service collections. • Monitor KPIs including charity care conversion rates, financial assistance enrollment, and point-of-service collection performance. • Ensure compliance with regulatory requirements for financial assistance programs (501(r) regulations, EMTALA). • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in patient financial counselling, financial clearance, or self-pay revenue cycle management, with at least 2 years in a leadership role. • Deep knowledge of charity care programs, financial assistance eligibility, Medicaid screening, and self-pay collections. • Familiarity with 501(r) regulatory requirements and hospital financial assistance policy compliance. • Experience with point-of-service collections, payment plan administration, and patient financial advocacy. • Proficiency with EHR platforms (Epic, Cerner) and financial assistance management tools. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to evaluate financial documentation and AI-generated outputs. — Preferred Qualifications • Certified Revenue Cycle Representative (CRCR) or similar revenue cycle certification. • Experience with presumptive eligibility screening tools and Medicaid enrolment facilitation. • Background in hospital, health system, or federally qualified health centre (FQHC) settings. • Familiarity with AI tools and comfort evaluating AI-generated patient financial content. • Experience developing patient financial education materials and staff training programs. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in healthcare revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$135 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (USA)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking USA-based voice actors with native American English accents to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native American English speaker currently based in the United States • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the clients CX AI Agent so please only apply if you are okay with voice cloning — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $150 / hourOpen / Referral verified
Finance / Remote

Netherlands Domestic Tax Specialist

We are seeking experienced Netherlands Domestic Tax Specialists to support the development and evaluation of AI systems focused on Dutch taxation. You will leverage your expertise in Dutch domestic tax law to answer complex tax questions, review AI-generated responses, identify inaccuracies, and provide detailed feedback to improve model performance. This is an excellent opportunity for tax professionals who enjoy solving complex tax issues while contributing to the next generation of AI-powered tax solutions. — What You'll Do • Answer complex questions related to Dutch domestic tax law using relevant legislation, case law, and official guidance. • Review AI-generated tax analyses for legal accuracy, completeness, and clarity. • Correct incorrect or incomplete responses and provide structured explanations. • Evaluate reasoning across personal and corporate tax scenarios. • Flag ambiguous, conflicting, or insufficient legal guidance where applicable. • Consistently apply project rubrics and quality standards. • Collaborate with calibration and feedback sessions to improve evaluation consistency. — Qualifications • 5+ years of experience advising on Dutch domestic taxation. • Experience in a Big Four firm, tax advisory firm, law firm, corporate tax department, or the Dutch Tax Administration (Belastingdienst) preferred. • Strong knowledge of Dutch Income Tax, Corporate Tax, VAT, and related domestic legislation. • Ability to interpret Dutch tax statutes, regulations, and official guidance. Dutch tax professionals routinely advise on compliance, planning, disputes, and representation before the tax authorities. • Excellent analytical and written communication skills. • Comfortable reviewing detailed legal and tax reasoning. • Native or professional fluency in Dutch; strong English proficiency is preferred. — Preferred Qualifications • Registered Dutch tax adviser (e.g., NOB, RB, or equivalent professional qualification). • Master's degree in Tax Law, Fiscal Economics, Accounting, or a related discipline. • Experience handling complex tax controversies or tax planning matters. • Prior experience reviewing legal or tax content, conducting quality assurance, or contributing to AI/LLM evaluation projects. — Why Join? • Apply your Dutch tax expertise to cutting-edge AI systems. • Work on intellectually challenging tax scenarios across multiple domains. • Collaborate with a global network of legal and tax professionals.

$100 - $130 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (African American)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking USA-based African American voice actors to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • USA-based voice actor with a natural African American English delivery • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $150 / hourOpen / Referral verified
Legal / Remote

Public Interest Attorney (Civil Justice Expert)

Mercor is seeking experienced Public Interest Attorneys to evaluate AI-generated legal reasoning across a broad range of civil justice matters. You'll review realistic legal scenarios, assess legal analysis, identify reasoning gaps, and provide structured expert feedback to improve AI performance. This role is ideal for attorneys with backgrounds in nonprofit advocacy, civil litigation, impact litigation, or public interest law. • * * — Responsibilities • Review AI-generated legal analyses for legal accuracy and practical usefulness. • Evaluate scenarios involving housing, family law, consumer protection, debt collection, and bankruptcy. • Assess legal reasoning, issue spotting, procedural correctness, and litigation strategy. • Identify legal errors and recommend improvements. • Deliver structured written evaluations. • * * — Required Qualifications • Juris Doctor (JD). • Active U.S. bar admission. • 5+ years practicing public interest or civil litigation. • Strong experience in civil legal advocacy. • Excellent legal writing and analytical skills. • * * — Preferred Qualifications • Experience at nonprofit legal organizations or impact litigation firms. • Expertise in employment law, elder law, or disability rights. • Experience with policy advocacy or systemic reform. • Experience supervising attorneys or legal fellows. • Interest in legal AI or emerging legal technology. • * * — Why Join Mercor? • Help shape AI tools supporting access to justice. • Apply real-world litigation experience to cutting-edge AI evaluation. • Work alongside leading legal and AI experts. • Competitive expert compensation. • * * — Engagement Details — Location: Remote (United States) — Commitment: Approximately 15 hours/week — Project Duration: Estimated 2–3 weeks

$150 / hourOpen / Referral verified
Medical / Remote (United States)

Medicare Advantage Members (Devoted Health) – Insight Study

Mercor is conducting a paid research study in collaboration with a leading AI research lab focused on improving healthcare and member experiences. We are seeking current or former Devoted Health Medicare Advantage members/Age 64+ to participate in a short online survey about their experience with their health plan. Participants will share perspectives on plan enrollment, benefits, member support, and day-to-day healthcare experiences. Insights gathered will help inform the development of AI tools designed to improve how health plans serve their members. — Responsibilities — Participants will be asked to: • Complete a structured ~20-minute online survey • Provide a mix of multiple-choice answers and short voice-recorded responses (approximately 10 voice responses) • Share perspectives on their experience as a Medicare Advantage plan member • Submit all responses within the given timeline — Requirements • Based in the United States • Age 64+ / Medicare-eligible • Currently or previously enrolled in a Devoted Health Medicare Advantage plan • Comfortable providing voice-recorded responses • Access to a microphone and a quiet environment • Able to independently complete a ~20-minute online survey — Engagement Details • Format: Online survey (voice-recorded responses + multiple-choice questions) • Duration: Approximately 20 minutes • Compensation: One-time payment upon successful verification of the submission • Location: Remote (United States) — Why Participate • Contribute to research shaping the next generation of healthcare AI tools • Share your real-world experience as a Medicare Advantage plan member

$120 / hourOpen / Referral verified
Code / Remote

LLM Research Scientist (Pre-training & Computer Vision & Adversarial Robustness)

We're looking for experienced machine learning researchers with hands-on experience training and improving deep learning models end-to-end, across vision and language. You'll work on well-scoped empirical open-ended ML research problems. — Responsibilities • Train image classifiers and generative image models from scratch, and fine-tune open-weight language models. • Get the most out of limited data, compute, and model-size budgets. • Make models robust — to adversarial inputs and to adversarial conversations. • Compress models to meet hard size and latency constraints without sacrificing accuracy. • Diagnose and resolve training issues. — Requirements — We are looking for candidates with strong expertise in one or more of the following areas: — Adversarial Robustness — Experience with: • Adversarial training of image classifiers (e.g. PGD-based training, TRADES). • Evaluating robust accuracy under standard threat models (e.g. L∞ attacks, AutoAttack) and avoiding gradient-masking pitfalls. • Managing the robustness–accuracy trade-off and robust overfitting. — Efficient Computer Vision — Experience with: • Training image classifiers end-to-end, especially for fine-grained recognition (many visually similar classes, few examples per class). • Model compression: quantization, pruning, and knowledge distillation from large teachers into small students. • Deploying models under hard size or latency budgets (on-device, edge, or embedded settings). — Generative Image Modeling — Experience with: • Training image generative models from scratch: diffusion models, GANs, VAEs, or flow-based models. • Iterating against sample-quality metrics such as FID. • Training-efficiency tricks that produce good generators quickly and at small parameter counts. — LLM Post-Training & Behavioral Robustness — Hands-on experience with one or more of: • Supervised fine-tuning and preference optimisation (DPO, RLHF, RLAIF) of open-weight language models, including building your own datasets via synthetic generation, noisy or weak supervision, and rejection sampling. • Shaping conversational behaviour over multiple turns: resistance to persuasion and sycophancy, calibrated confidence, and knowing when to accept corrections. • Alignment-style fine-tuning that changes a specific behaviour while preserving general capability. — Multilingual Pre-training — Experience with: • Training multilingual or low-resource-language models from scratch. • Tokenizer design across scripts and typologically diverse languages. • Balancing highly unequal per-language data (sampling temperatures, cross-lingual transfer) in data-constrained regimes. — Additional Areas of Interest — Experience in any of the following is a plus: • Scaling laws and training-efficiency research. • Curriculum learning and data ordering. • Model evaluation: benchmark construction, contamination control, statistically sound comparisons. • Uncertainty estimation and model calibration. • Data augmentation and synthetic data for robustness. — General Qualifications • 3+ years of machine learning research experience (PhD research counts toward this requirement). • Strong experience with PyTorch, JAX, TensorFlow, or similar ML frameworks. • Degree from a top-100 university, experience at a FAANG or comparable AI company, or an equivalent research track record through publications or impactful open-source contributions. — Why Join • Work on cutting-edge machine learning research. • Collaborate with leading AI researchers on challenging, high-impact projects. • Flexible, project-based work with competitive compensation.

$100 - $120 / hourOpen / Referral verified
Language / Australia (remote)

In-House Counsel (Australia)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of Australian law. We are hiring experienced in-house counsel to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • 5+ years as in-house legal counsel practising in Australia • Strong command of Australian law, especially commercial and contract matters • Sharp attention to detail and precise written English • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$100 - $120 / hourOpen / Referral verified
Language / Remote

Pricing / ROI / revenue economics Evaluator

About the role — We are hiring expert Evaluators in Pricing / ROI / revenue economics to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Pricing / ROI / revenue economics. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Business / Remote

Product management / roadmap / PRD Evaluator

About the role — We are hiring expert Evaluators in Product management / roadmap / PRD to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Product management / roadmap / PRD. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Business / Remote

Public-sector procurement / RFI response Evaluator

About the role — We are hiring expert Evaluators in Public-sector procurement / RFI response to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Public-sector procurement / RFI response. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

User/customer research and feedback synthesis Evaluator

About the role — We are hiring expert Evaluators in User/customer research and feedback synthesis to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in User/customer research and feedback synthesis. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

BI dashboards / performance reporting Evaluator

About the role — We are hiring expert Evaluators in BI dashboards / performance reporting to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in BI dashboards / performance reporting. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Nonprofit / philanthropy / community programs Evaluator

About the role — We are hiring expert Evaluators in Nonprofit / philanthropy / community programs to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Nonprofit / philanthropy / community programs. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Legal / Australia (remote)

Litigation Lawyer (Australia)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of Australian law. We are hiring experienced litigation & disputes lawyers to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • Primarily law-firm / chambers experience — at least 2 years (3+ preferred) at a law firm or barristers' chambers in litigation, disputes, or contentious practice. If you are currently in-house, that is fine only if you moved in-house within the last 2 years and your prior experience was at a firm/chambers. • Qualified to practise in Australia, with a strong command of Australian litigation and procedural law • Native or fluent English, with precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$100 - $120 / hourOpen / Referral verified
Legal / Australia (remote)

Corporate/M&A Lawyer (Australia)

Mercor is partnering with a leading legal AI company to benchmark how well AI systems answer real questions of Australian law. We are hiring experienced corporate/M&A lawyers to serve as expert reviewers — running the same prompts through two AI platforms, comparing the answers, and scoring them against a standardized rubric. — What you'll do • Run the same set of prompts across two AI (LLM) platforms • Compare the outputs from both platforms side by side • Score each response against a standardized rubric we provide • Submit concise written feedback for every evaluation — Requirements • Primarily law-firm experience — at least 2 years (3+ preferred) at a law firm in corporate/M&A or transactional practice. If you are currently in-house, that is fine only if you moved in-house within the last 2 years and your prior experience was at a law firm. • Qualified to practise in Australia, with a strong command of Australian corporate and commercial law • Native or fluent English, with precise written communication • Able to work independently to a rubric and to deadlines — Logistics • Initial commitment: ~10 hours, with strong potential for more • Confidentiality: all work is covered by NDA

$100 - $120 / hourOpen / Referral verified
Legal / Bay Area, CA

Senior Counsel — Legal Domain Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — We are hiring an experienced counsel-level lawyer to work directly with a leading AI lab's research and program management teams, improving how frontier AI models reason about real legal work. — This listing is for practitioners at counsel or senior associate level, with roughly 8 to 15 years of post-qualification practice. If you are earlier in your career, or if you are at partner or General Counsel level, please apply to the corresponding Legal Domain Expert listing instead — we run separate listings by seniority so the rate matches the experience. — Your legal expertise is the substance of this role. You will review the quality of legal knowledge work tasks, write the instruction specs and golden solutions that define what "correct" looks like, and build the benchmarks that show whether the model is genuinely improving. We are looking for a practising specialist rather than a generalist. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. You will be provisioned with client-issued accounts and equipment, and will work inside the client's own tools alongside their research teams. — Location: This is a hybrid role based in the Bay Area, California. You must live in the Bay Area and work on-site with the client's team multiple days each week, when required. This is not a remote role. If you do not currently live in the Bay Area, you must be willing to relocate there at your own cost before the engagement starts — relocation assistance is not provided. — 2. Key Responsibilities • Data QA and reviews: Vet the quality of legal knowledge work tasks and model outputs — spotting missing behaviors, thin reasoning, and answers that read well but would not survive professional scrutiny. • Instruction specs and golden datasets: Write high-quality instruction specs, produce golden solutions to legal problems, and define new legal tasks that reflect how the work is actually done in practice. • Benchmarks and domain depth: Design challenging legal tasks and evaluation sets, and help build legal-specific skills and tools together with the research team. • Calibration: Work with client researchers and specialists in adjacent fields to keep standards consistent, translating tacit legal judgment into explicit, teachable criteria. — 3. Core Qualifications • Education: Juris Doctor (JD) from an accredited law school; a highly ranked school is strongly preferred. • Experience: 8 to 15 years of substantive post-qualification legal practice at a reputable institution — an established law firm, a corporate legal department, a regulatory body, or a court. Internships and clerkships alone do not count. • Seniority: Currently at or has reached counsel, senior associate, senior in-house counsel, or Assistant General Counsel level, with real ownership of matters. • Domain depth: Genuine specialization in at least one substantive practice area, for example corporate and transactional, litigation and dispute resolution, regulatory and compliance, intellectual property, employment and labor, or tax. • Licensure: Admission to at least one U.S. state bar, in good standing. • AI fluency: Hands-on working use of large language models in your professional work, and the judgment to tell a well-reasoned answer from a plausible-sounding wrong one. • Availability: Able to commit reliably to 40 hours per week for an initial engagement of 6 months. • Location: Living in the Bay Area, California, and able to work on-site with the client's team multiple days each week, when required. Candidates not currently based in the Bay Area must be willing to relocate there at their own cost; relocation assistance is not provided. • Excellent written communication, and the ability to give precise, well-structured written feedback. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$85 - $120 / hourOpen / Referral verified
Business / Remote

Operations / inventory / capacity planning Evaluator

About the role — We are hiring expert Evaluators in Operations / inventory / capacity planning to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Operations / inventory / capacity planning. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
STEM / Remote

Biology / environmental science Evaluator

About the role — We are hiring expert Evaluators in Biology / environmental science to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Biology / environmental science. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Business / Remote

Incident management / reliability / SRE Evaluator

About the role — We are hiring expert Evaluators in Incident management / reliability / SRE to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Incident management / reliability / SRE. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Medical / Remote

Healthcare / clinical Evaluator

About the role — We are hiring expert Evaluators in Healthcare / clinical to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Healthcare / clinical. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Medical / Remote

Applied Health & Medicine Benchmark Specialist

Role Overview — We are seeking expert medical and health science professionals to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core health and medicine domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of medical expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Health & Medicine Domains Covered — Clinical Medicine & Surgery, Medical Imaging & Diagnostics, Pharmacovigilance, Healthcare Management & Economics, Rehabilitation and Allied Health. — Key Responsibilities • Author original health and medicine questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise reasoning in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, clinical guidelines) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • MD, DO, PhD, or doctoral candidate in Medicine, Biomedical Sciences, Public Health, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level medical knowledge, clinical reasoning, and biomedical research methodology • Board certification, clinical experience, or research publications in health fields is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$94 - $119 / hourOpen / Referral verified
Finance / Remote

Tax Accountant / Specialist

Role Overview — Mercor is collaborating with a leading AI lab to engage experienced tax professionals. You'll translate real tax compliance, provision, and planning work into structured, high-quality training data that teaches AI to reason about how tax accountants actually work. — Focus Areas — Tax — across corporate, individual, partnership / pass-through, SALT, and international sub-specialties. — Key Responsibilities • Design realistic scenarios from your tax work — return preparation (1040 / 1120 / 1120-S / 1065), tax provision (ASC 740), responses to notices / IDRs and controversy support, tax planning and estimated payments • Review and compare AI-generated tax outputs for technical accuracy and defensible positions • Provide clear written feedback that improves how AI performs tax tasks • Collaborate asynchronously with the research team — Ideal Qualifications • CPA or EA (Enrolled Agent) • A clear tax sub-specialty — corporate, individual, partnership, SALT, or international / cross-border • Public accounting (firm) or in-house tax experience • Bachelor's degree in Accounting, Finance, or a related field • Strong written communication and attention to detail — Application Process • Submit a resume or a short summary of your tax experience • Complete a short form on your practice area, specialties, and certifications • Selected applicants may complete a brief sample task

$80 - $120 / hourOpen / Referral verified
Finance / Remote

Corporate / Controllership Accountant

Role Overview — Mercor is collaborating with a leading AI lab to engage experienced corporate / controllership accountants. You'll translate the day-to-day of running the books — financial reporting, the month-end close, and core accounting operations — into structured, high-quality training data that teaches AI to reason about real accounting work. — Focus Areas — Financial accounting & reporting · bookkeeping / client accounting services (CAS) · accounting operations (AP / AR / payroll). — Key Responsibilities • Design realistic scenarios from your controllership work — financial statement preparation, journal entries, account & bank reconciliations, month-end / period-end close, fixed-asset & depreciation schedules, accruals and prepaids • Review and compare AI-generated accounting outputs for accuracy, GAAP compliance, and sound professional judgment • Provide clear written feedback that improves how AI performs close and reporting tasks • Collaborate asynchronously with the research team — Ideal Qualifications • CPA, or 3–5+ years of corporate/in-house or firm accounting experience (no niche certification required) • Hands-on ownership of the general ledger and the period-end close • Bachelor's degree in Accounting, Finance, or a related field • Comfortable with common accounting tools (QuickBooks, NetSuite, SAP, Oracle, Excel) • Strong written communication and attention to detail — Application Process • Submit a resume or a short summary of your accounting experience • Complete a short form on your practice area, specialties, and certifications • Selected applicants may complete a brief sample task

$80 - $120 / hourOpen / Referral verified
Finance / Remote

Accounting Expert

Role Overview — Mercor is collaborating with a leading AI lab to engage experienced accounting professionals across all areas of practice — including audit, tax, financial reporting, bookkeeping, controllership, forensic, and accounting systems. Contributors help build AI systems that reason about real accounting work by translating everyday accounting workflows, judgments, and decision-making into structured, high-quality training data. — Key Responsibilities • Design realistic accounting scenarios and tasks drawn from your day-to-day work (e.g., financial statement preparation, reconciliations, journal entries, audit procedures, tax filings, month-end close, internal controls) • Review and compare AI-generated accounting outputs for accuracy, standards compliance (GAAP/IFRS), and sound professional judgment • Create structured examples that reflect how accountants actually reason through problems • Provide clear written feedback that improves how AI performs accounting tasks • Collaborate asynchronously with the research team — Ideal Qualifications • 3+ years of professional experience in accounting, audit, tax, bookkeeping, or finance operations (public accounting, corporate/in-house, advisory, or a firm) • CPA, CA, ACCA, CMA, EA, or an equivalent professional credential required • Bachelor's degree in Accounting, Finance, or a related field • Comfortable with common accounting tools (e.g., QuickBooks, NetSuite, SAP, Oracle, Excel) • Strong written communication and attention to detail — More About the Opportunity • Open to all accounting specialties — contribute where your expertise is strongest • Work spans task design, evaluation, and structured feedback on AI accounting outputs • Strong contributors advance into reviewer, lead, and domain-expert roles — Application Process • Submit a resume or a short summary of your accounting experience • Complete a short form on your practice area, specialties, and certifications • Selected applicants may complete a brief sample task • Follow-up typically provided within a few days

$80 - $120 / hourOpen / Referral verified
Math / Remote

Data analysis / quantitative readouts Evaluator

About the role — We are hiring expert Evaluators in Data analysis / quantitative readouts to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Data analysis / quantitative readouts. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Business / Remote

Healthcare operations Evaluator

About the role — We are hiring expert Evaluators in Healthcare operations to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Healthcare operations. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Code / United States Remote

Senior Software Engineer, Full Stack (Python, Java, Rust, C#, C++)

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Senior Full-Stack Software Engineers with deep, hands-on production expertise across modern application stacks — Python, Java, Rust, C#, C++, and TypeScript — to build, integrate, and stress-test real software on top of frontier models before they ship. Engineers work directly with the AI lab's engineering manager on fast-moving projects and are expected to onboard quickly and contribute working code early. This is a full-time commitment of 40 hours per week. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Build and ship full-stack applications, services, and internal tools — in Python, Java, Rust, C#, C++, or TypeScript — that exercise frontier model capabilities end to end. • Integrate pre-release model APIs into working software, including the surrounding scaffolding: tool interfaces, evaluation harnesses, and telemetry. • Diagnose and document model and integration failure modes surfaced while building, translating engineering observations into clear written feedback for the research team. • Prototype quickly against shifting requirements, delivering working increments with limited up-front specification. • Collaborate with the engineering manager and other engineers to maintain consistency in code quality, architecture, and technical documentation. — 3. Core Qualifications • 6+ years of dedicated professional experience building and shipping production software at a recognized, top-tier organization, including ownership of systems end to end. • Deep production experience in at least one of Python, Java, Rust, C#, or C++, plus demonstrated delivery in a second language across a different ecosystem — this cohort is intentionally staffed across multiple language ecosystems. • Hands-on ownership across the stack: backend service and API design, a modern front-end framework (React or equivalent), data modelling, and cloud deployment and operations. • Demonstrated ability to onboard onto an unfamiliar codebase and deliver working code quickly with minimal ramp-up. • Demonstrable career progression. • Ability to engage reliably for at least 40 hours/week during weekdays. • Strong written communication skills and the ability to explain complex technical decisions clearly. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$90 - $110 / hourOpen / Referral verified
Code / United States Remote

MLOps Engineer (JAX, PyTorch, Pallas/Triton)

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented MLOps Engineers with deep, hands-on expertise in modern ML frameworks — specifically JAX, PyTorch, and kernel-level programming (Pallas/Triton). This role involves AI model training and evaluation work, including writing and assessing MLOps tasks and solutions to generate high-quality training data for frontier AI systems. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. This is a 40-hour full-time engagement, with no conflicts/no other engagements. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps and improve AI model performance in MLOps, training infrastructure, and ML framework-level topics. • Design challenging, domain-relevant tasks, and write accurate and well-structured solutions to MLOps and ML systems problems. • Evaluate MLOps tasks and solutions and provide clear, written technical feedback. • Develop guidelines and detailed rubrics/evaluation frameworks to assess training pipeline design, distributed systems reasoning, and kernel-level optimization across tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 2+ years of dedicated professional experience in ML infrastructure, MLOps, or ML systems engineering at a recognized, top-tier organization. • Hands-on production experience with JAX and/or PyTorch at scale. • Experience writing or optimizing custom GPU kernels using Pallas (JAX) or Triton. • Demonstrable career progression. • Ability to engage reliably for at least 40 hours/week during weekdays. • Strong written communication skills and the ability to explain complex technical decisions clearly. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$70 - $110 / hourOpen / Referral verified
Code / United States Remote

Performance Engineer (C++, Python, Rust)

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — We're seeking talented Performance Engineers with deep expertise in low-level systems optimization — specifically C++, Python, and Rust — to bring hands-on technical excellence and elevate the quality of our AI training and inference infrastructure data. This role involves AI model training and evaluation work, including writing and assessing performance-engineering tasks and solutions to generate high-quality training data for frontier AI systems. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. This is a 40-hour full-time engagement, with no conflicts/no other engagements. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps and improve AI model performance in systems-level optimization, compiler engineering, and runtime performance topics. • Design challenging, domain-relevant tasks across multiple specializations, and write accurate and well-structured solutions to performance engineering problems. • Evaluate performance engineering tasks and solutions and provide clear, written technical feedback. • Develop guidelines and detailed rubrics/evaluation frameworks to assess systems design quality across AI workloads. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 2+ years of dedicated professional experience in performance engineering, systems programming, or low-level optimization. • Deep hands-on expertise in at least one of the following: C++, Python, or Rust — with working familiarity across the others being a strong plus. • Demonstrable track record of measurable performance improvements on production systems (e.g., latency reduction, throughput gains, memory footprint optimization). • Demonstrable career progression. • Ability to engage reliably for at least 40 hours/week during weekdays. • Strong written communication skills and the ability to explain complex technical decisions clearly. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$70 - $110 / hourOpen / Referral verified
Code / Remote

Risk-adjustment / HCC coding leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Risk Adjustment and HCC Coding leaders to evaluate AI tools designed to improve risk score accuracy and coding completeness in Medicare Advantage, Medicaid managed care, and ACA markets. Your expertise in hierarchical condition category (HCC) methodology, RADV audit preparation, and risk adjustment coding will directly shape AI systems that enhance documentation capture and risk score integrity. — Responsibilities • Lead risk adjustment and HCC coding operations across Medicare Advantage, Medicaid, and/or ACA risk adjustment programs. • Evaluate AI-generated HCC coding assignments and risk adjustment recommendations for clinical accuracy and regulatory compliance. • Review medical records to ensure complete and accurate capture of HCC-eligible conditions supported by clinical documentation. • Conduct and oversee retrospective and prospective chart reviews for risk score optimisation. • Manage RADV (Risk Adjustment Data Validation) audit preparation and response processes. • Monitor risk adjustment KPIs including HCC capture rates, risk score accuracy, and chart retrieval rates. • Collaborate with clinical, coding, and compliance teams to improve documentation and coding for risk adjustment purposes. • Ensure compliance with CMS risk adjustment guidelines (RAPS, EDGE submissions) and Official Coding Guidelines. • Annotate AI outputs and provide structured coding feedback to support AI training datasets. — Requirements • 5+ years of experience in risk adjustment coding, HCC coding, or Medicare Advantage coding operations, with at least 2 years in a leadership role. • Deep expertise in CMS-HCC, RxHCC, and/or ACA HHS-HCC risk adjustment methodologies. • Strong knowledge of ICD-10-CM coding guidelines as applied to HCC risk adjustment. • Experience with RADV audit preparation and CMS compliance requirements. • Familiarity with RAPS and EDGE submission processes. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify coding inaccuracies and documentation gaps in AI-generated outputs. — Preferred Qualifications • CRC (Certified Risk Coder), CCS, CPC, or RHIA credential. • Experience with risk adjustment analytics platforms and chart retrieval systems. • Background in health plan, Medicare Advantage organisation, or value-based care setting. • Familiarity with AI-assisted HCC coding tools and comfort evaluating AI-generated risk adjustment content. • Experience presenting risk adjustment performance to actuarial or executive teams. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in risk adjustment and value-based care. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$110 / hourOpen / Referral verified
Business / Remote

Security Expert

About the Role — Mercor is building realistic, high-fidelity simulated environments to evaluate and train AI models on real-world procurement workflows for a leading spend-management technology company. We're looking for security and vendor-risk professionals to author and validate security-review tasks inside these simulated environments. — Key Responsibilities • Review simulated vendor SOC 2 reports, security questionnaires, and pen-test evidence against a buyer's security standard • Catch scope mismatches and lapsed bridge letters that a surface-level review would miss • Author step-level rubrics and golden responses capturing how an experienced security reviewer would judge a request or renewal • Assess data-handling and sub-processor risk for vendors touching sensitive data — Ideal Qualifications • 8+ years of professional experience in security review, vendor risk management, or third-party risk (TPRM) • Hands-on experience evaluating SOC 2 reports, security questionnaires, and compliance evidence • Strong written communication skills; comfortable producing structured, rubric-style feedback — Nice to Have • Relevant security certification, such as CISSP or CISA • Prior task-writing, rubric-authoring, or AI-training data experience

$90 - $110 / hourOpen / Referral verified
Code / Remote

Document/deck production QA Evaluator

About the role — We are hiring expert Evaluators in Document/deck production QA to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Document/deck production QA. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Finance / Remote

Advisory & Transaction Services Expert (M&A / Valuation)

Role Overview — Mercor is collaborating with a leading AI lab to engage experienced transaction advisory and valuation professionals. You'll translate real M&A, due diligence, and valuation work into structured, high-quality training data that teaches AI to reason the way deal teams do. — Focus Areas — M&A / transaction advisory · financial due diligence (quality of earnings) · business valuation · purchase accounting (ASC 805) support. — Key Responsibilities • Design realistic scenarios from your work — quality-of-earnings (QoE) analyses, working-capital and net-debt analyses, financial due-diligence findings, valuation models (DCF, comparable companies, precedent transactions), purchase price allocation (ASC 805), and deal / LBO models • Review and compare AI-generated advisory outputs for analytical accuracy, defensible assumptions, and sound professional judgment • Provide clear written feedback that improves how AI performs transaction and valuation tasks • Collaborate asynchronously with the research team — Ideal Qualifications • Transaction advisory (Big 4 TAS / boutique) and/or corporate development, investment banking, or valuation background • CPA, and/or ASA / ABV / CFA a plus for valuation-leaning candidates • Hands-on ownership of QoE, valuation, or deal models • Bachelor's degree in Accounting, Finance, or a related field • Strong written communication and attention to detail — Application Process • Submit a resume or a short summary of your transaction advisory / valuation experience • Complete a short form on your practice area, specialties, and certifications • Selected applicants may complete a brief sample task

$80 - $120 / hourOpen / Referral verified
Finance / Remote

Accounting Domain Intake — Shape Upcoming Accounting Engagements

About This Survey — Mercor is opening a series of accounting-focused engagements with leading AI labs over the coming weeks and months — spanning audit reasoning, financial reporting, tax, GAAP/IFRS analysis, and adjacent workstreams. Before we finalize the scope of each, we want to build them around how accountants *actually* work, not how the field is described in textbooks or public sources. — This is a written intake designed to map out the real workflows, tools, decision points, and specialty cuts inside the profession — from the inside. — What You'll Do • Complete a ~30-minute written intake about your accounting practice • Describe your day-to-day workflows, how your time is split, and the workflows core to your role • Explain how experienced accountants actually divide the field from the inside — the cuts insiders use, not the textbook ones • Share the "invisible" parts of your work that outsiders (including project designers) tend to miss • No right or wrong answers — detail and specificity matter far more than polish — Ideal Respondents • Practicing or recently practicing accountants — public accounting, corporate finance, audit, tax, or advisory • CPA (or equivalent) preferred; deep domain experience without the credential also welcome • Willing to write detailed, specific responses about the mechanics of your work — Why This Matters • Your intake shapes the scope of the accounting engagements we launch • Thoughtful responses translate directly into more offer matches and priority routing as we open new engagements over the coming weeks and months • Typical accounting rates on Mercor range from $80–$120/hr and scale with seniority — About Mercor — Mercor is a talent marketplace that connects experts with leading AI labs and research organizations. Backed by Benchmark, General Catalyst, Adam D'Angelo, Larry Summers, and Jack Dorsey.

$80 - $120 / hourOpen / Referral verified
Finance / Remote

FP&A & Treasury Analyst

Role Overview — Mercor is collaborating with a leading AI lab to engage experienced FP&A and treasury professionals. You'll translate real planning, analysis, and cash / treasury work into structured, high-quality training data that teaches AI to reason about how finance teams actually work. — Focus Areas — Financial planning & analysis (FP&A) · treasury & cash management. — Key Responsibilities • Design realistic scenarios from your work — operating budgets and rolling forecasts, flux / variance analysis and commentary, 13-week cash flow forecasts, daily cash positioning & bank administration, FX revaluation, debt / covenant compliance schedules, and formula-driven financial models • Review and compare AI-generated FP&A / treasury outputs for accuracy and sound judgment on assumptions • Provide clear written feedback that improves how AI performs planning and cash-management tasks • Collaborate asynchronously with the research team — Ideal Qualifications • Finance / MBA background with strong financial-modeling skills and hands-on cash / treasury exposure • CTP (Certified Treasury Professional) a plus for treasury-leaning candidates • Corporate FP&A or treasury experience at a mid-size or larger company • Bachelor's degree in Finance, Accounting, or a related field • Strong written communication and attention to detail — Application Process • Submit a resume or a short summary of your FP&A / treasury experience • Complete a short form on your practice area, specialties, and certifications • Selected applicants may complete a brief sample task

$80 - $120 / hourOpen / Referral verified
Legal / Remote

Applied Legal Benchmark Specialist

Role Overview — We are seeking legal experts to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core law domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of legal expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Law Domains Covered — Intellectual Property, Privacy, and Technology, Regulatory and Government Affairs, Securities, Capital Markets, Financial Regulation & Compliance, Private Equity, M&A & Transaction Structuring, Antitrust, Merger Control & Competition, Healthcare, Life Sciences, & Pharmaceuticals, Environmental, Energy, & ESG/Climate. — Key Responsibilities • Author original law questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise reasoning in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, legal repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • JD, LLM, SJD, or doctoral candidate in Law or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of legal reasoning, statutory interpretation, and jurisprudential theory • Experience with bar exam writing, legal academia, or judicial clerkships is a strong plus • Excellent written English and ability to express complex legal ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$83 - $105 / hourOpen / Referral verified
STEM / Remote UK

Molecular Biology Experts

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Molecular Biology subject-matter experts (SMEs) with hands-on experience designing DNA and RNA sequences — primers, plasmids, guide RNAs, mRNA constructs, and repair templates — to bring deep domain expertise and elevate the quality of our AI training data. This is a part-time to full-time commitment of up to 40 hours per week. — This is a W-2 employment position with Cincinnatus LLC (or appropriate international entity), with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps and improve AI model performance in molecular biology, nucleic-acid sequence design, and construct engineering. • Design challenging, domain-relevant tasks and write accurate, well-documented solutions — spanning primers (PCR/qPCR/cloning), plasmids/expression vectors, gRNA/sgRNA for CRISPR editing, mRNA/RNA constructs, and HDR/repair templates — that serve as ground truth. • Evaluate molecular-biology tasks and AI model outputs against expert-quality solutions and provide clear, written technical feedback on correctness, rigor, and biological reasoning. • Develop guidelines and detailed rubrics/evaluation frameworks to assess sequence-design quality — guide/target selection, homology-arm length, primer melting temperature and specificity, codon optimization, and regulatory elements. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • Advanced degree (PhD strongly preferred) in molecular biology, genetics, biochemistry, synthetic biology, bioengineering, or a related life-science field. • Hands-on experience designing nucleic-acid constructs prior to experiments — across several of primers, plasmids/vectors, gRNA/sgRNA, mRNA, and HDR/repair templates — together with cloning strategy (Gibson, Golden Gate, Gateway, restriction) and codon optimization. • A strong peer-reviewed publication record, weighted toward first-author work in notable venues (e.g., Nature and the Nature family, Cell, eLife, Nucleic Acids Research, PNAS, EMBO Journal). Please include publication links with your application. • Demonstrable career progression. • Ability to engage reliably for at least 20 hours/week during weekdays. • Past experience in AI training, model evaluation, and data annotation is preferred. • Strong written communication skills and the ability to justify design choices clearly and precisely. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$70 - $105 / hourOpen / Referral verified
Medical / Remote

Prior Authorisation Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Prior Authorisation Managers with clinical and medical expertise to support the evaluation of AI tools designed to streamline prior authorisation workflows. Your clinical knowledge of medical-necessity criteria, payer review processes, and utilisation management will help train AI systems to reduce administrative burden and improve patient access to care. — Responsibilities • Manage end-to-end prior authorisation workflows for medical and clinical services across multiple payer types. • Review clinical documentation to assess medical necessity against InterQual, MCG, or payer-specific criteria. • Evaluate AI-generated prior authorisation recommendations and clinical justification drafts for accuracy and appropriateness. • Coordinate with clinical staff, physicians, and payers to obtain timely authorisation approvals. • Track authorisation status, denials, and appeal outcomes to identify workflow improvement opportunities. • Develop and maintain SOPs for prior authorisation submission, follow-up, and escalation processes. • Monitor KPIs including authorisation approval rates, turnaround times, and denial rates. • Ensure compliance with payer requirements, CMS guidelines, and clinical review criteria. • Annotate AI outputs and provide structured clinical feedback to support AI training datasets. — Requirements • 5+ years of experience in prior authorisation, utilisation management, or clinical review, with at least 2 years in a management role. • Strong clinical background with knowledge of medical necessity criteria (InterQual, MCG, or equivalent). • Deep familiarity with commercial, Medicare Advantage, and Medicaid prior authorisation requirements. • Experience managing authorisation workflows across multiple specialities and service types. • Proficiency with authorisation management systems and EHR platforms (Epic, Cerner, or equivalent). • Exceptional written and verbal English communication skills. • High attention to detail with the ability to critically evaluate clinical documentation and AI-generated outputs. — Preferred Qualifications • Clinical licensure (RN, LPN, or equivalent) or relevant certification (CPUR, CPUM). • Experience with appeals and peer-to-peer review processes. • Familiarity with CMS prior authorisation rules and the No Surprises Act. • Background in physician office, hospital, or health plan prior authorisation operations. • Familiarity with AI tools and comfort evaluating AI-generated clinical content. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in healthcare revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$105 / hourOpen / Referral verified
Medical / Remote

Insurance Verification & Benefit Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Insurance Verification and Eligibility & Benefits Managers to evaluate AI tools designed to automate front-end revenue cycle operations. Your deep expertise in payer systems, EDI transactions, and eligibility workflows will directly inform AI models that drive accuracy and efficiency in insurance verification processes. — Responsibilities • Oversee insurance verification, eligibility determination, and benefits investigation workflows across commercial, Medicare, Medicaid, and managed care payers. • Verify patient insurance coverage using payer portals, clearinghouses, and EDI 270/271 real-time eligibility transactions. • Identify and resolve coordination of benefits issues, coverage gaps, and eligibility discrepancies prior to service delivery. • Evaluate and annotate AI-generated eligibility verification outputs for accuracy, completeness, and payer compliance. • Develop and document SOPs for eligibility and benefits verification workflows. • Monitor KPIs including verification accuracy rates, front-end denial rates related to eligibility, and turnaround times. • Ensure compliance with payer-specific requirements, CMS guidelines, and HIPAA regulations. • Identify process gaps and recommend workflow improvements to reduce eligibility-related claim denials. • Provide structured feedback and annotations to support AI training datasets. — Requirements • 5+ years of experience in insurance verification, eligibility and benefits management, or front-end revenue cycle operations, with at least 2 years in a management role. • Expert knowledge of EDI 270/271 transactions, payer portal navigation, and real-time eligibility tools. • Strong familiarity with Medicare, Medicaid, and commercial payer eligibility requirements and benefit structures. • Proficiency with Epic, Cerner, Meditech, or equivalent EHR platforms. • Experience with clearinghouse platforms such as Availity, Change Healthcare, or similar. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify subtle discrepancies in coverage data. — Preferred Qualifications • CHAM, CHAA, or equivalent front-end revenue cycle certification. • Experience with automated eligibility verification tools and RPA solutions. • Background in denial root cause analysis related to eligibility and coverage errors. • Familiarity with AI tools and comfort evaluating AI-generated healthcare content. • Experience presenting eligibility performance data to senior leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in healthcare revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$105 / hourOpen / Referral verified
Legal / Remote

IP / trademark / copyright law Evaluator

About the role — We are hiring expert Evaluators in IP / trademark / copyright law to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in IP / trademark / copyright law. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Legal / Remote

Legal contracts / diligence / redlines Evaluator

About the role — We are hiring expert Evaluators in Legal contracts / diligence / redlines to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Legal contracts / diligence / redlines. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Finance / Remote

Personal finance / consumer planning Evaluator

About the role — We are hiring expert Evaluators in Personal finance / consumer planning to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Personal finance / consumer planning. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Finance / Remote

Compliance / regulatory response with financial-services AI Evaluator

About the role — We are hiring expert Evaluators in Compliance / regulatory response with financial-services AI to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Compliance / regulatory response with financial-services AI. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Business / Remote

Brand / creative direction / marketing collateral Evaluator

About the role — We are hiring expert Evaluators in Brand / creative direction / marketing collateral to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Brand / creative direction / marketing collateral. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Policy & Safety / Remote

Privacy / regulatory compliance Evaluator

About the role — We are hiring expert Evaluators in Privacy / regulatory compliance to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Privacy / regulatory compliance. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Process improvement / SOPs Evaluator

About the role — We are hiring expert Evaluators in Process improvement / SOPs to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Process improvement / SOPs. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Training / onboarding / L&D Evaluator

About the role — We are hiring expert Evaluators in Training / onboarding / L&D to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Training / onboarding / L&D. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Cybersecurity / IT GRC Evaluator

About the role — We are hiring expert Evaluators in Cybersecurity / IT GRC to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Cybersecurity / IT GRC. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Finance / Remote

Finance operations / audit support Evaluator

About the role — We are hiring expert Evaluators in Finance operations / audit support to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Finance operations / audit support. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Media / journalism / communications Evaluator

About the role — We are hiring expert Evaluators in Media / journalism / communications to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Media / journalism / communications. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Business / Remote

General business strategy / management Evaluator

About the role — We are hiring expert Evaluators in General business strategy / management to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in General business strategy / management. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Finance / Remote

FP&A / corporate finance Evaluator

About the role — We are hiring expert Evaluators in FP&A / corporate finance to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in FP&A / corporate finance. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Medical / Remote

Clinical / biomedical / pharma Evaluator

About the role — We are hiring expert Evaluators in Clinical / biomedical / pharma to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Clinical / biomedical / pharma. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Finance / Remote

Investment analysis / valuation / credit Evaluator

About the role — We are hiring expert Evaluators in Investment analysis / valuation / credit to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Investment analysis / valuation / credit. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
STEM / Remote

Computational Chemistry & Electronic Structure Expert

Computational Chemistry & Electronic Structure Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience in: — Computational Chemistry & Electronic Structure — working with PySCF for quantum chemistry calculations including Hartree-Fock, DFT, TDDFT, CASSCF, and post-HF methods. Ideal candidates can design problems around excited-state analysis, orbital diagnostics, choosing the right method for tricky electronic structures, and interpreting computational artifacts that come from method limitations. — _Experience with other specialized software in this domain will also be considered._ — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $100 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Southern USA)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking USA-based voice actors with native Southern American English accents (e.g. Texas, Georgia, Tennessee, the Carolinas, Louisiana, Alabama, Mississippi) to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native Southern American English speaker currently based in the United States, with an authentic Southern US accent • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5–10 hours per week throughout the project duration. Please note: This is a short-term project expected to last 1–2 weeks. • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $100 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Standard French)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking native or native-level French speakers with a standard, neutral, accent-free French delivery to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native or native-level French speaker with a standard, neutral French accent (accent-free) — no strong regional accent and no foreign-language influence • Location does not matter — we welcome qualified speakers from anywhere, as long as the delivery is neutral, standard French • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration. • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $100 / hourOpen / Referral verified
STEM / Remote

Computational Particle & Nuclear Physics Expert

Particle & Nuclear Physics Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience in: — Particle & Nuclear Physics — working with scikit-hep and related HEP Python tools for particle physics data analysis, cross-section computations, renormalization group calculations, and perturbative QCD. Experience with Monte Carlo event generation or collider phenomenology is a plus. — _Experience with other specialized software in this domain will also be considered._ — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $100 / hourOpen / Referral verified
STEM / Remote

Computational Astrophysics & Cosmology Expert

Computational Astrophysics & Cosmology Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience in: — Astrophysics & Cosmology — working with astropy and related tools for cosmological calculations, angular power spectra, galaxy survey analysis, and observational data reduction pipelines. — _Experience with other specialized software in this domain will also be considered._ — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $100 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning - Northern UK

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Northern UK-based female voice actors with authentic Northern English accents (e.g. Yorkshire, Manchester, Newcastle, Liverpool) to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Female voice actor with a natural, authentic Northern English accent, currently based in the northern UK • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $100 / hourOpen / Referral verified
Code / Remote

Bioinformatics & Computational Single-Cell Genomics Expert

Bioinformatics & Computational Single-Cell Genomics Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience in: — Bioinformatics & Single-Cell Genomics — working with tools like scanpy, scvelo, squidpy, and gudhi for single-cell RNA-seq analysis, trajectory inference, spatial transcriptomics, and topological data analysis. You should be comfortable designing problems around cell-type annotation, pseudotime ordering, multi-omic integration, spatial variable gene identification, and persistence-based analysis pipelines. This is our highest-throughput domain and where we're scaling first. — _Experience with other specialized software in this domain will also be considered._ — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $100 / hourOpen / Referral verified
Finance / Bay Area, CA

Finance Domain Expert — AI Training & Evaluation

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — We are hiring a senior finance domain expert to work directly with a leading AI lab's research and program management teams, improving how frontier AI models reason about real financial work. — Your finance expertise is the substance of this role. You will review the quality of finance knowledge work tasks, write the instruction specs and golden solutions that define what "correct" looks like, and build the benchmarks that show whether the model is genuinely improving. We are looking for a practising specialist rather than a generalist. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. You will be provisioned with client-issued accounts and equipment, and will work inside the client's own tools alongside their research teams. — Location: This is a hybrid role based in the Bay Area, California. You must live in the Bay Area and work on-site with the client's team multiple days each week, when required. This is not a remote role. If you do not currently live in the Bay Area, you must be willing to relocate there at your own cost before the engagement starts — relocation assistance is not provided. — 2. Key Responsibilities • Data QA and reviews: Vet the quality of finance knowledge work tasks and model outputs — spotting missing behaviors, thin reasoning, flawed assumptions, and answers that read well but would not survive professional scrutiny. • Instruction specs and golden datasets: Write high-quality instruction specs, produce golden solutions to financial problems, and define new finance tasks that reflect how the work is actually done in practice. • Benchmarks and domain depth: Design challenging finance tasks and evaluation sets, and help build finance-specific skills and tools together with the research team. • Calibration: Work with client researchers and specialists in adjacent fields to keep standards consistent, translating tacit financial judgment into explicit, teachable criteria. — 3. Core Qualifications • Experience: 5+ years of substantive, dedicated professional finance experience at a recognized institution — for example an investment bank, asset manager, private equity or credit fund, Big Four firm, a large corporate finance function, or a financial regulator. Generalist roles that only touch finance peripherally do not count. • Domain depth: Genuine specialization in at least one core finance discipline, for example corporate finance and FP&A, investment banking and M&A, asset or wealth management, private equity or private credit, quantitative finance and risk management, treasury, or accounting and audit. • Seniority: Clear progression to a senior individual-contributor or leadership level — for example Vice President, Director, Principal, Managing Director, Portfolio Manager, Controller, or CFO — with real ownership of analysis and decisions. • Education and credentials: An advanced degree from a strong program (MBA, MS, or PhD in finance, economics, accounting, or a quantitative field) and/or a recognized professional credential such as CFA, CPA, FRM, or an actuarial designation. Strongly preferred. • AI fluency: Hands-on working use of large language models in your professional work, and the judgment to tell a well-reasoned answer from a plausible-sounding wrong one. • Availability: Able to commit reliably to 40 hours per week for an initial engagement of 6 months. • Location: Living in the Bay Area, California, and able to work on-site with the client's team multiple days each week, when required. Candidates not currently based in the Bay Area must be willing to relocate there at their own cost; relocation assistance is not provided. • Excellent written communication, and the ability to give precise, well-structured written feedback. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $100 / hourOpen / Referral verified
Legal / Bay Area, CA

Legal Domain Expert — AI Training & Evaluation

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — We are hiring a senior legal domain expert to work directly with a leading AI lab's research and program management teams, improving how frontier AI models reason about real legal work. — Your legal expertise is the substance of this role. You will review the quality of legal knowledge work tasks, write the instruction specs and golden solutions that define what "correct" looks like, and build the benchmarks that show whether the model is genuinely improving. We are looking for a practising specialist rather than a generalist. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. You will be provisioned with client-issued accounts and equipment, and will work inside the client's own tools alongside their research teams. — Location: This is a hybrid role based in the Bay Area, California. You must live in the Bay Area and work on-site with the client's team multiple days each week, when required. This is not a remote role. If you do not currently live in the Bay Area, you must be willing to relocate there at your own cost before the engagement starts — relocation assistance is not provided. — 2. Key Responsibilities • Data QA and reviews: Vet the quality of legal knowledge work tasks and model outputs — spotting missing behaviors, thin reasoning, and answers that read well but would not survive professional scrutiny. • Instruction specs and golden datasets: Write high-quality instruction specs, produce golden solutions to legal problems, and define new legal tasks that reflect how the work is actually done in practice. • Benchmarks and domain depth: Design challenging legal tasks and evaluation sets, and help build legal-specific skills and tools together with the research team. • Calibration: Work with client researchers and specialists in adjacent fields to keep standards consistent, translating tacit legal judgment into explicit, teachable criteria. — 3. Core Qualifications • Education: Juris Doctor (JD) from an accredited law school; a highly ranked school is strongly preferred. • Experience: 5+ years of substantive post-qualification legal practice at a reputable institution — an established law firm, a corporate legal department, a regulatory body, or a court. Internships and clerkships alone do not count. • Domain depth: Genuine specialization in at least one substantive practice area, for example corporate and transactional, litigation and dispute resolution, regulatory and compliance, intellectual property, employment and labor, or tax. • Seniority: Clear progression to a senior level — partner, of counsel, counsel, senior associate, senior in-house counsel, or general counsel — with real ownership of matters. • Licensure: Active admission to at least one U.S. state bar, in good standing. • AI fluency: Hands-on working use of large language models in your professional work, and the judgment to tell a well-reasoned answer from a plausible-sounding wrong one. • Availability: Able to commit reliably to 40 hours per week for an initial engagement of 6 months. • Location: Living in the Bay Area, California, and able to work on-site with the client's team multiple days each week, when required. Candidates not currently based in the Bay Area must be willing to relocate there at their own cost; relocation assistance is not provided. • Excellent written communication, and the ability to give precise, well-structured written feedback. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $100 / hourOpen / Referral verified
Language / Remote

Voice Actor: CX Agent Voice Cloning (UK English)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking UK-based native English speakers with a neutral, standard UK English accent to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native English speaker with a neutral, standard UK English accent, currently based in the United Kingdom • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $100 / hourOpen / Referral verified
Language / Remote

Voice Actor: CX Agent Voice Cloning (Australian English)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Australia-based native English speakers with authentic Australian accents to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native English speaker with an authentic Australian accent, currently based in Australia • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $100 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (German)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Germany-based native German speakers to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native German speaker currently based in Germany • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $100 / hourOpen / Referral verified
Finance / Remote

Revenue-cycle analytics / decision-support / RCM reporting leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Revenue Cycle Analytics and Decision-Support leaders — including RCM reporting managers and analytics directors — to evaluate AI tools designed to transform revenue cycle intelligence and financial decision-making. Your expertise in RCM data analytics, financial reporting, and decision-support will directly shape AI systems that deliver actionable insights across the full revenue cycle. — Responsibilities • Lead revenue cycle analytics, reporting, and decision-support functions to drive data-driven performance improvement across RCM operations. • Evaluate AI-generated revenue cycle analytics outputs, KPI dashboards, and financial modeling recommendations for accuracy and completeness. • Develop and maintain revenue cycle reporting frameworks covering patient access, coding, billing, denials, A/R, and collections performance. • Build and interpret dashboards, scorecards, and trend analyses to support operational and executive decision-making. • Conduct root cause analyses of revenue cycle performance variances and develop data-driven improvement recommendations. • Collaborate with finance, IT, and operational revenue cycle teams to align analytics infrastructure with strategic priorities. • Manage RCM data governance, including data definitions, data quality standards, and reporting consistency. • Support revenue cycle forecasting, budget modelling, and net revenue realisation analysis. • Annotate AI outputs and provide structured analytical feedback to support AI training datasets. — Requirements • 5+ years of experience in revenue cycle analytics, RCM reporting, or healthcare financial decision-support, with at least 2 years in a leadership role. • Deep knowledge of revenue cycle KPIs and financial metrics across patient access, coding, billing, denials, and collections domains. • Proficiency with healthcare analytics platforms, BI tools (Tableau, Power BI, or equivalent), and SQL-based data analysis. • Experience with EHR-integrated analytics platforms and RCM reporting systems. • Strong financial modelling and data interpretation skills with the ability to translate complex data into actionable recommendations. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify data quality issues and analytical errors in AI-generated outputs. — Preferred Qualifications • HFMA CRCR, CHFP, or healthcare analytics certification. • Experience with predictive analytics, revenue cycle forecasting, and net revenue modelling. • Background in health system, large physician enterprise, or RCM outsourcing analytics operations. • Familiarity with AI tools and comfort evaluating AI-generated analytics content. • Experience implementing enterprise-wide revenue cycle analytics infrastructure or BI platform migrations. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in revenue cycle analytics and decision-support. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$100 / hourOpen / Referral verified
Language / Remote

Market research / competitive intelligence Evaluator

About the role — We are hiring expert Evaluators in Market research / competitive intelligence to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Market research / competitive intelligence. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Business / Remote

General Sales / GTM Evaluator

About the role — We are hiring expert Evaluators in General Sales / GTM to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in General Sales / GTM. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Humanities / arts / culture Evaluator

About the role — We are hiring expert Evaluators in Humanities / arts / culture to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Humanities / arts / culture. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Real estate / hospitality / events Evaluator

About the role — We are hiring expert Evaluators in Real estate / hospitality / events to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Real estate / hospitality / events. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Special education / IEP Evaluator

About the role — We are hiring expert Evaluators in Special education / IEP to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Special education / IEP. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Public health communications Evaluator

About the role — We are hiring expert Evaluators in Public health communications to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Public health communications. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Government / public administration Evaluator

About the role — We are hiring expert Evaluators in Government / public administration to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Government / public administration. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Business / Remote

Customer success / support operations Evaluator

About the role — We are hiring expert Evaluators in Customer success / support operations to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Customer success / support operations. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Finance / Remote

Applied Economics & Finance Benchmark Specialist

Role Overview — We are seeking expert economists and finance professionals to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core economics and finance domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Economics & Finance Domains Covered — Algorithmic Trading & Market Microstructure, Macroprudential Policy, Behavioral Finance & Experimental Economics, Urban Economics, Tokenomics & Decentralized Finance. — Key Responsibilities • Author original economics and finance questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Economics, Finance, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level economic theory, quantitative methods, and financial modeling • Experience with academic research, CFA/CPA certification, or financial industry expertise is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$77 - $98 / hourOpen / Referral verified
Business / Remote

Applied Business & Commerce Benchmark Specialist

Role Overview — We are seeking business and commerce experts to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core business domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of business expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Business & Commerce Domains Covered — Logistics / Supply Chain Management, Operations Research Analysis, Sustainability, Management Analysis, Product Management and Strategy, Digital Marketing. — Key Responsibilities • Author original business questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise reasoning in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD, DBA, or doctoral candidate in Business Administration, Management, Marketing, or a closely related field • MBA or Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level business strategy, organizational theory, and quantitative methods • Industry leadership experience or research publications in business fields is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$77 - $98 / hourOpen / Referral verified
Business / Remote

Data quality / CRM operations Evaluator

About the role — We are hiring expert Evaluators in Data quality / CRM operations to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Data quality / CRM operations. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Finance / Remote

General finance / accounting Evaluator

About the role — We are hiring expert Evaluators in General finance / accounting to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in General finance / accounting. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Policy & Safety / Remote

Legal / compliance Evaluator

About the role — We are hiring expert Evaluators in Legal / compliance to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Legal / compliance. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Product launch / experiment readiness Evaluator

About the role — We are hiring expert Evaluators in Product launch / experiment readiness to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Product launch / experiment readiness. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Business / Remote

Program management / implementation planning Evaluator

About the role — We are hiring expert Evaluators in Program management / implementation planning to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Program management / implementation planning. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Investor materials / fundraising / pitchbook Evaluator

About the role — We are hiring expert Evaluators in Investor materials / fundraising / pitchbook to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Investor materials / fundraising / pitchbook. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Language / Remote

Education / school Evaluator

About the role — We are hiring expert Evaluators in Education / school to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. — This is a remote, hourly engagement. — Requirements (must have) — 1. 5+ years of relevant professional experience in Education / school. 2. Native or professional fluency in English. 3. Highly proficient in Microsoft Office and Google Workspace, especially Slides (Google Slides / PowerPoint). — Preferred (nice to have) • Advanced degree (Master's or higher) from a reputable institution. — What you'll do • Evaluate AI-generated artifacts against domain-specific quality rubrics. • Identify factual, aesthetic, and presentation errors. • Provide clear, structured written feedback.

$80 - $120 / hourOpen / Referral verified
Finance / Remote

Patient Financial Services Leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Patient Collections and Patient Financial Services leaders to evaluate AI tools designed to improve self-pay revenue recovery and patient financial engagement. Your expertise in self-pay collections, payment plan administration, and patient financial advocacy will directly inform AI systems that optimise patient payment experiences while maximising revenue recovery. — Responsibilities • Lead patient collections and self-pay operations including early-out collections, bad debt management, and patient payment plan administration. • Evaluate AI-generated patient financial communication drafts, payment plan recommendations, and self-pay resolution strategies for accuracy and compliance. • Develop and implement self-pay collection strategies across the revenue cycle, including pre-service, point-of-service, and post-service collections. • Manage patient payment plan enrollment, monitoring, and compliance processes. • Coordinate with financial counselling, billing, and bad debt recovery teams to optimise self-pay revenue capture. • Monitor self-pay KPIs including self-pay collection rates, payment plan conversion rates, bad debt write-off rates, and patient satisfaction scores. • Ensure compliance with FDCPA, HIPAA, state collections laws, and internal patient financial assistance policies. • Oversee relationships with collection agencies and early-out vendors as applicable. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in patient collections, self-pay revenue cycle, or patient financial services, with at least 2 years in a leadership role. • Deep knowledge of self-pay collection workflows, FDCPA compliance, and patient financial engagement best practices. • Experience managing early-out and bad debt collection programs, including vendor oversight. • Familiarity with propensity-to-pay tools and patient payment technology platforms. • Proficiency with EHR systems and patient collections/billing platforms. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to evaluate collections strategies and identify issues in AI-generated patient communication content. — Preferred Qualifications • CRCR, CHAM, or similar revenue cycle certification. • Experience with digital patient payment platforms and text/email-based collections outreach. • Background in hospital, health system, or physician group patient financial services. • Familiarity with AI tools and comfort evaluating AI-generated patient financial communication content. • Experience implementing propensity-to-pay analytics to prioritise collection efforts. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in patient financial services and the revenue cycle. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$92 / hourOpen / Referral verified
Finance / Remote

Investment Banking Expert

Role Overview • Mercor is seeking senior investment banking professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise M&A and capital markets contexts. • The workflows are calibrated to the deal complexity, stakeholder sensitivity, and transaction stakes of Fortune 500 M&A, capital raises, and large-cap advisory mandates. • Contributors design enterprise investment banking scenarios, draft reference outputs, and write rubrics that capture how senior F500-focused bankers think. — Key Responsibilities • Construct enterprise investment banking scenarios spanning $1B+ M&A transactions, multi-stakeholder deal committees, and complex capital markets or regulatory review cycles. • Build tasks across M&A advisory, equity and debt capital markets, leveraged finance, valuation and financial modeling, and pitch/deal execution. • Develop deal execution scenarios involving tools such as Bloomberg, Capital IQ, FactSet, and enterprise financial modeling and deal-management platforms used in bulge-bracket workflows. • Apply enterprise investment banking methodologies (DCF/comparable company/precedent transaction analysis, LBO modeling, fairness opinion standards) and produce reference pitch books, valuation models, and executive/board-level deal narratives. • Author rubrics that distinguish authentic investment banking judgment from generic textbook or interview-prep-level recall. — Ideal Qualifications • 5+ years working in M&A advisory, capital markets, or leveraged finance at a bulge-bracket bank, elite boutique, or major financial institution (Goldman Sachs, Morgan Stanley, JPMorgan, Evercore, Lazard). • Direct ownership of F500-scale deal execution, capital raises, or advisory mandates. • Fluency in investment banking tooling, valuation methodologies, and deal-process mechanics, plus understanding of how F500 board approvals, regulatory review (SEC, antitrust), and legal/due diligence processes actually work. • Prior rubric, financial-modeling curriculum, or pitch book/deal documentation authorship is a plus.

$90 - $100 / hourOpen / Referral verified
Legal / Remote

Regulatory Law Expert

Role Overview • Mercor is seeking senior regulatory law professionals to build evaluation tasks for AI systems operating in regulatory compliance and government affairs contexts. • The workflows are calibrated to the regulatory complexity, enforcement stakes, and scope of major regulated-industry compliance programs. • This role builds worlds on two tracks: a US track (Administrative Procedure Act and agency-specific statutes such as SEC, FDA, and FTC rules) and an International track (EU regulatory frameworks and UK regulatory bodies). Experts qualified in either or both tracks are encouraged to apply. • Contributors design regulatory law scenarios, draft reference outputs, and write rubrics that capture how senior regulatory counsel think. — Key Responsibilities • Construct scenarios spanning regulatory compliance program design, agency investigations and enforcement actions, and rulemaking or comment processes. • Build tasks across financial services regulation, healthcare and FDA regulation, antitrust and competition law, data privacy and cybersecurity regulation, and environmental and energy regulation. • Develop scenarios involving tools such as regulatory tracking platforms, compliance management systems, and agency filing systems used across regulated industries. • Apply regulatory law methodologies (regulatory interpretation, compliance risk assessment, enforcement response strategy) to the standards track a world targets (US: Administrative Procedure Act, agency-specific statutes and regulations; International: EU regulations and directives, UK regulatory frameworks), and produce reference compliance memoranda, regulatory filings, and enforcement response documents. • Author rubrics that distinguish authentic regulatory judgment from generic law school or bar exam-level recall. — Ideal Qualifications • 5+ years working as a regulatory attorney or compliance counsel at a major law firm, regulated company, or government agency (Covington & Burling, Sidley Austin, WilmerHale, or in-house regulatory/compliance counsel at a bank, pharmaceutical company, or technology company). • Direct ownership of regulatory compliance programs, agency interactions, or enforcement matters. • Fluency in regulatory tooling, plus understanding of how administrative process and agency rulemaking actually work. • A recognized professional credential is strongly preferred (JD with bar admission, or an international equivalent); prior rubric, training, or compliance documentation authorship is a plus.

$90 - $100 / hourOpen / Referral verified
Business / Remote

Intellectual Property Expert

Role Overview • Mercor is seeking senior intellectual property professionals to build evaluation tasks for AI systems operating in patent prosecution, licensing, and IP enforcement contexts. • The workflows are calibrated to the technical complexity, commercial stakes, and procedural scope of major patent portfolios, licensing programs, and IP litigation. • This role builds worlds on two tracks: a US track (35 U.S.C., MPEP, USPTO procedure) and an International track (European Patent Convention, PCT, WIPO treaties). Experts qualified in either or both tracks are encouraged to apply. • Contributors design IP scenarios, draft reference outputs, and write rubrics that capture how senior IP counsel think. — Key Responsibilities • Construct IP scenarios spanning patent prosecution and portfolio strategy, trademark and copyright disputes, and complex licensing or technology transfer processes. • Build tasks across patent prosecution, IP litigation and enforcement, trademark and copyright practice, licensing and technology transfer, and freedom-to-operate analysis. • Develop scenarios involving tools such as patent search platforms (PatSnap, Innography, USPTO PatFT), IP docketing and portfolio management systems (Anaqua, CPA Global), and claim-charting tools used across major IP practices. • Apply IP methodologies (patentability analysis, claim construction, freedom-to-operate review) to the standards track a world targets (US: 35 U.S.C., MPEP, Federal Circuit precedent; International: EPC, PCT, WIPO treaties), and produce reference patent applications, office action responses, licensing agreements, and litigation memoranda. • Author rubrics that distinguish authentic IP judgment from generic law school or bar exam-level recall. — Ideal Qualifications • 5+ years working as an IP attorney or patent agent at a major IP firm or corporate IP department (Fish & Richardson, Finnegan, Wilson Sonsini, or in-house IP counsel at a technology or pharmaceutical company). • Direct ownership of patent prosecution portfolios, licensing negotiations, or IP litigation matters. • Fluency in IP tooling and methodologies, plus understanding of how USPTO/EPO procedure and international treaty systems actually work. • A recognized professional credential is strongly preferred (JD with bar admission and USPTO patent bar registration, or an international equivalent such as European Patent Attorney); a technical/STEM background plus prior rubric or training authorship is a plus.

$90 - $100 / hourOpen / Referral verified
Code / Remote

Frontend Engineer — Web Replication Preference Rater

About the work — We're building a high-quality dataset of human preference judgments on AI-generated frontend code. You'll be shown a reference web page alongside two candidate replications produced by AI models, and you'll decide which replication is better — then explain why in writing that a model can learn from. — This is evaluation work, not authoring. You won't be building sites from scratch. You'll be reading someone else's HTML and CSS, running it locally, comparing it pixel-by-pixel against a target, and articulating exactly where and why it falls short. — What you'll do • Render a reference page and two candidate replications side by side at desktop width and judge which is the closer reproduction. • Diff layout fidelity in detail: box model and spacing, typography (family, size, weight, line-height, letter-spacing), color and border treatment, image and asset handling, z-order and overflow. • Inspect the underlying markup with browser devtools to distinguish a replication that is genuinely correct from one that merely looks correct at one viewport — hardcoded pixel offsets, absolute positioning standing in for real layout, and inline styles that will not survive a resize. • Evaluate responsive behavior and semantic quality: whether flexbox and grid are used where they belong, whether legacy float or table layouts in the reference were reproduced faithfully, whether headings and landmarks carry real semantic meaning. • Write a structured rationale for every judgment — the specific defects you found, ranked by how much they matter, in language precise enough to be actionable. • Flag ties, ambiguous cases, and broken task items rather than forcing a preference. — You're a fit if you have • 3+ years of professional web development experience, primarily in frontend or full-stack work. • Fluency in hand-written HTML and CSS: semantic markup, flexbox, grid, media queries, and older float- and table-based layouts you can still read and reason about. • Working command of browser devtools — element inspection, computed styles, the box model, and the network panel. • Enough JavaScript to read a page's scripts and understand what they do to the DOM, even if you don't write JS daily. • Comfort in a terminal: cloning a folder and serving it over a local static server without help. • Strong written English and the discipline to justify a judgment rather than assert it. — Equipment • A desktop or laptop with a browser window that opens to at least 1920px wide. • Administrator rights on your own machine, so you can install and run a local server. — Nice to have • Prior RLHF, preference labeling, or model evaluation work. • A code review or technical assessment background. • Pixel-perfect design-to-code experience — translating Figma or PSD comps into production markup. • Web accessibility expertise (WCAG, ARIA, screen reader testing). • Familiarity with how LLMs typically fail at codegen. • Web scraping or DOM parsing experience. — Note: this seat is for practicing web developers. Backend-only, mobile-native-only, data science, DevOps, and design-without-code backgrounds are out of scope for this project.

$90 / hourOpen / Referral verified
Business / US Remote

Enterprise Sales Domain Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is teaching its AI to do real enterprise sales work — the kind of multi-step, judgment-heavy workflows that only someone who has actually carried a quota knows how to get right. To do that well, they need seasoned sellers to act as the ground truth: people who can show the model what excellent looks like, catch where it goes wrong, and set the bar its work is measured against. — That is where you come in. As a Sales Domain Expert, you will bring years of real selling experience to auditing sales workflows, building the "golden" reference examples the model learns from, and shaping the rubrics used to evaluate it. This is a role for accomplished enterprise sellers — not general data annotators. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States. — 2. What You'll Do • Set the standard for "good": Review multi-step enterprise sales workflows and judge whether the AI handled them the way a strong seller would, flagging what is off and why. • Build golden trajectories: Work real sales tasks end-to-end — prospecting, pipeline management, CRM hygiene, reporting, and customer-facing documents — to create the expert reference examples the model is trained on. • Shape the evaluation: Pressure-test and refine the rubrics used to score AI outputs so they capture what actually separates great selling from average. • Work in realistic sandboxes: Operate inside test environments that mirror a real sales tech stack, completing and simulating workflows across CRM, communication, and billing tools. • Produce senior-level artifacts: Create polished, customer-facing materials — pitches, slide decks, and summaries — at the quality a senior sales leader would put their name on. — 3. Who We're Looking For • Around 10 years of hands-on experience selling in enterprise environments, with a real feel for the full sales cycle. • A consistent track record of meeting or exceeding your sales quota, which you can speak to in specifics. • Sellers at every level are welcome — quota-carrying individual contributors, managers who lead selling teams, and VPs running regional or global sales organizations. • Day-to-day fluency with the standard enterprise stack (Salesforce, Jira, Confluence, Google Workspace, and Microsoft Office), and the comfort to pick up tools like Slack, HubSpot, Notion, Aircall, Stripe, Intercom, Zoho, and Microsoft Teams. • Genuine depth in CRM administration, sales process management, pipeline tracking, and data-driven reporting. • Strong writing and a good eye for polish — you can turn a complex situation into a clear, professional, customer-facing artifact. — Nice to Have • Experience working within or alongside a two-sided, AI-powered B2B SaaS creator marketplace — for example, running outcome-based campaigns against contracted performance targets (tracking views, engagement, installs, and sign-ups in real time) and troubleshooting across a connected stack such as Zoho CRM, Jira, Confluence, Freshdesk, Microsoft Teams, and Google Workspace. If this sounds like you, tell us — we are especially keen to talk. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / Remote

Computational Statistics and Applied Mathematics Expert (R, Python, and Matlab/Scilab)

Computational Statistics and Applied Mathematics Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized statistical, mathematical, or scientific software packages. Some will ask the AI to compute reproducible numerical answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We welcome statisticians and applied mathematicians working across a wide range of specializations. You do not need experience with every package listed below; strong expertise with one or more specialized computational packages is sufficient. — We're especially interested in experts with deep, hands-on experience using one or more specialized R or Python packages, including examples such as: • Bayesian statistics: rstan, cmdstanr, rjags, runjags, brms, rstanarm, nimble, bayesplot, posterior, loo • Item response theory and psychometrics: TAM, sirt, mirt, mirtCAT, eRm, ltm, lordif, psych • Structural equation and latent variable modelling: lavaan, semTools, OpenMx • Topological data analysis: TDAstats, TDApplied • Differential equations and dynamical systems: deSolve, pomp, FME • State-space and time-series modelling: KFAS, MARSS, forecast, vars, urca, rugarch, rmgarch, tseries, timeSeries • Survival and event-history analysis: survival, flexsurv, timereg, mets • Mixed, additive, and advanced regression models: lme4, nlme, mgcv, glmmTMB, TMB, quantreg, scam • Spatial statistics and geostatistics: spatstat, spatstat.geom, spatstat.linnet, spdep, gstat, geoR, spBayes, sf, stars, terra, lwgeom • Statistical learning and specialized modelling: mclust, kernlab, earth, pROC, multcomp, sandwich, effectsize, irr • Optimization and mathematical programming: lpSolve, linprog, nloptr, DEoptimR, SQUAREM • Numerical linear algebra and high-precision computation: RSpectra, Rmpfr, gmp, pracma • Computational geometry: geometry, deldir, polyclip — Other similar specialized statistical, mathematical, scientific, or domain-specific R packages will also be considered. Other similar specialized statistical or mathematical Python/Scilab packages are also welcome, such as statsmodels and PyMC. — Numerical computing and scientific modelling in Matlab/Scilab are also wanted. — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD required; PhD preferred, or MS with 10+ years of relevant experience) in statistics, applied mathematics, or a closely related quantitative field, with real hands-on experience using specialized computational packages — not just theoretical knowledge. — You have written code using one or more specialized statistical, mathematical, or scientific packages to solve actual research or professional problems, and you understand where these tools break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. Deep expertise with one or more specialized computational packages is more important than familiarity with the entire package list above. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in statistics, applied mathematics, a relevant STEM field, or equivalent research experience • Proven proficiency with at least one specialized statistical, mathematical, or scientific software package, demonstrated through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple computational domains or specialized software packages • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $90 / hourOpen / Referral verified
Medical / Playa Vista, CA or New York, NY Remote

Character and Facial Animation Consultant

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building a performance transfer model: one that takes an actor's original performance and carries it faithfully into a new output, holding on to the timing, the emotion, and the small physical choices that make it feel real. The team believes AI should honor the craft of performance, not flatten it. — To get the evaluation right, we are bringing in senior character animators to help define what "good" looks like, critique the current evaluation approach, and shape the pool of reviewers who score model outputs. — This is a part-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Establish evaluation standards: Define clear criteria for when a performance has been captured well, both in emotional resonance and physical fidelity, and point out where current evaluations fall short. • Curate benchmarks: Identify strong examples of performances that are easy to capture and ones that are genuinely hard, so the model has meaningful tests to measure against. • Shape the evaluation pool: Help define the right profile, expertise mix, and rubric for the human reviewers who assess outputs at scale. • Collaborate with other experts: Work alongside fellow subject-matter experts to keep evaluations consistent and accurate. — 3. Core Qualifications • 4+ years as a hands-on character animator, with human or character _performance_ as your core craft. • Verifiable credits on 2 or more feature films, animated features, or AAA game cinematics at recognized studios. • Direct experience with facial animation and/or performance capture — facial keyframe work, motion-capture solving, cleanup or motion editing, or character technical design of facial and performance rigs. • A meticulous eye for micro-expressions, emotional beats, and the subtleties that make a human performance feel authentic. • AI proficiency: familiarity with AI tools and workflows, including the ability to generate your own visual samples for testing and comparison. • The ability to translate complex visual details into clear, precise language that can be captured in data captions and evaluation templates. • Able to come on-site in Playa Vista (Los Angeles), CA or New York, NY for in-person working sessions once or twice a month. • (Optional) Reliable access to a 4K-resolution monitor for precise, pixel-level review. — 4. Who this role is not for — This role is specifically for animators whose primary craft is character performance. It is not a fit for directors of photography, lighting or camera specialists, colorists, compositors, editors, or generalist 3D artists. Layout, previsualization, environment and effects work will not qualify on their own. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / United States Remote

LLM Red Team Specialist — Failure Modes & Edge Cases

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs specialists who can find where those models break. Working in a red-teaming setup, you will design and probe complex, multi-step tasks to expose vulnerabilities, edge cases, and failure modes in frontier AI systems — the places where a model looks competent but is quietly wrong. — Each task represents one to two days of continuous, focused effort and spans multiple technical skills: coding, experimentation, and careful analysis. You will work in a tight feedback loop with the lab's researchers, turning the failure modes you find into stronger benchmark tasks. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Probe models: Explore how frontier AI models behave on coding, ML, and analysis tasks, and find the spots where they quietly get things wrong. • Design challenges: Turn the weaknesses you find into well-crafted tasks that are hard for models but fair to grade. • Document findings: Write up what you discover clearly, with evidence and steps others can reproduce. • Strengthen tasks: Team up with task authors to close loopholes, shortcuts, and grading gaps. • Work as a team: Share insights with researchers and fellow experts so the benchmark keeps getting better. — 3. Core Qualifications • MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding. • 1+ years of experience in a research, research-engineering, security, or AI-evaluation role. • Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems — through red teaming, adversarial testing, security research, or rigorous model evaluation. • Working proficiency in Python and Git, with the ability to script your own probes and analyses. • Strong familiarity with LLM capabilities, limitations, and evaluation techniques. • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in finding what others missed, strong written communication, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / United States Remote

QA/Test Engineer

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models, and complex multi-step tasks are only useful if they are airtight: unambiguous, correctly graded, and robust to shortcuts. We are seeking experienced QA and test engineers to be the quality backbone of this benchmark — designing the test cases and review processes that guarantee every task measures what it claims to measure. — Each task under review represents one to two days of expert effort and spans multiple technical skills, so quality review here means genuinely understanding the task: running it, probing its edge cases, and debugging its environment. You will work in a tight feedback loop with the lab's researchers and task authors. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design checks: Create test cases that confirm each task works as intended — including the tricky edge cases. • Review tasks: Give tasks and reference solutions a careful read before they're finalized, catching ambiguity and gaps early. • Debug: Roll up your sleeves in Python when a task or its checks don't behave the way they should. • Shape the process: Help build simple, repeatable quality checklists, and share feedback authors can act on right away. • Protect the results: Watch for shortcuts and grading gaps in AI agent runs so benchmark scores stay trustworthy. — 3. Core Qualifications • MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain. • 1+ years of experience in test engineering, quality assurance, or a research/software engineering role with strong quality ownership. • Demonstrated skill designing test cases and quality-review processes, and debugging complex systems end-to-end. • Working proficiency in Python and Git, and comfort navigating unfamiliar codebases and environments. • Exceptional attention to detail and clear written documentation habits. • Past experience in AI training, model evaluation, or quality review of AI-generated work is preferred. • A perfectionist mindset: creativity in finding what others missed, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Medical / United States Remote

STEM Researcher — Computational Fields

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and is recruiting researchers from computational STEM fields — as well as computationally heavy social sciences and humanities — to bring working-researcher rigor to benchmark design. You will translate the scientific method — experimental design, hypothesis testing, and rigorous evaluation — into complex, multi-step tasks that today's best models cannot yet complete reliably. — Each task represents one to two days of continuous, focused effort and spans multiple skills: study design, implementation in code, data analysis, and careful written conclusions. You will work in a tight feedback loop with the lab's researchers, surfacing the kinds of methodological mistakes a working researcher would catch immediately. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design tasks: Turn the research skills you use every day — designing studies, testing hypotheses, evaluating results — into engaging, multi-step tasks. • Author solutions: Work through your own tasks in Python and notebooks, at the level of rigor you'd expect from a careful colleague. • Define what good looks like: Help spell out what separates sound scientific reasoning from reasoning that merely sounds right. • Evaluate models: Review model attempts at your tasks and flag the mistakes a working researcher would spot right away. • Work as a team: Compare notes with researchers and fellow experts to keep evaluations consistent and accurate. — 3. Core Qualifications • MSc or PhD in a STEM field, or in a computational social-science or humanities discipline, or equivalent practical experience in a research-heavy domain requiring data analysis and coding. • 1+ years of experience in an active research role (academia, industry, or national labs). • Your own research involves significant computational work: Python-based analysis, simulation, modeling, or data pipelines. • Strong grounding in experimental design, hypothesis testing, and rigorous evaluation of results. • Working familiarity with Git, IDEs, and notebook environments (Jupyter or Colab). • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Math / United States Remote

Data Science & Quantitative Analysis Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced data scientists and quantitative analysts to act as ground-truth experts. You will design complex analysis tasks that simulate real research work — for example, comparing two anomaly-detection algorithms on a dataset, calculating correlations, performing manual spot checks, and summarizing the findings in a notebook clear enough to drive a researcher's decision. — Each task represents one to two days of continuous, focused effort and spans multiple skills: data cleaning, statistical analysis, interpretation, and clear reporting. You will work in a tight feedback loop with the lab's researchers, verifying exactly where and why frontier models fall short on rigorous analytical work. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design tasks: Create realistic data-analysis challenges — cleaning messy data, comparing methods, interpreting results — based on the kind of analysis you do every day. • Author notebooks: Work through your own tasks in Jupyter or Colab, producing clear, reproducible reference analyses. • Compare methods: Build tasks that ask for a fair comparison between analytical approaches, backed by spot checks and a clear recommendation. • Evaluate models: Review how models handle your tasks, and check whether their statistics and conclusions actually hold up. • Work as a team: Compare notes with researchers and fellow experts to keep evaluations consistent and accurate. — 3. Core Qualifications • MSc or PhD in statistics, data science, or another quantitative STEM field, or equivalent practical experience in a research-heavy analytical domain. • 1+ years of experience in a research, research-engineering, or heavy data-analysis role. • Deep hands-on data-analysis skills: data cleaning, statistical correlation, hypothesis testing, and careful interpretation of results. • Proficiency with Jupyter Notebooks or Google Colab for analysis and reporting. • Working proficiency in Python (pandas, NumPy, or similar) and Git. • Strong ability to communicate analytical findings in writing for decision-makers. • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in task design, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / United States Remote

Machine Learning Engineer — Model Evaluation & Experimentation

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation. You will author complex, multi-step ML tasks — for example, taking a vague research idea like "modify how an RL reward is computed," implementing the change, running the training experiment, and analyzing the results to determine success — and verify exactly where frontier models fall short. — Each task represents one to two days of continuous, focused effort and spans multiple technical skills: implementation, experiment setup and execution, and rigorous analysis. You will work in a tight feedback loop with the lab's researchers. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design tasks: Turn real ML research ideas — like tweaking how an RL reward is computed — into well-defined, multi-step tasks. • Run experiments: Implement changes, run training experiments, and analyze the results to show what a correct solution looks like. • Explore RL ideas: Build some of your tasks around reinforcement-learning basics such as reward functions and training behavior. • Evaluate models: See how frontier models handle your tasks, and note where and why they fall short. • Work as a team: Compare notes with researchers and fellow experts so tasks stay consistent, rigorous, and fair. — 3. Core Qualifications • MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy domain. • 1+ years of experience in a research or research-engineering role. • Hands-on experience training and evaluating ML models and running experiments end-to-end — experiment setup, execution, and analysis — not just using ML libraries superficially. • Strong familiarity with large language models: their capabilities, limitations, and evaluation techniques. • Working proficiency in Python and Git, with comfort in both scripting and notebook environments. • Basic understanding of reinforcement learning (reward functions, policy training) is preferred. • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / United States Remote

Software Engineering Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced software engineers to act as task curators. You will design, implement, and review complex, multi-step engineering tasks that simulate the real-world challenges research engineers face — realistic, genuinely hard problems that today's best AI coding agents cannot yet solve reliably. — Each task represents one to two days of continuous, focused effort and spans multiple technical skills: Python implementation, environment and tooling setup, debugging, and clear documentation. You will work in a tight feedback loop with the lab's researchers, verifying exactly where and why frontier models fail on your tasks. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design tasks: Create realistic, multi-step software engineering challenges — the kind of work you'd actually do on the job — that push the limits of today's best AI coding agents. • Build reference solutions: Solve your own tasks in Python, with the setup and checks needed so each task has a clear, verifiable answer. • Work with AI tools: Use AI coding assistants as part of your everyday workflow, and observe where they help and where they fall short. • Review and refine: Look over tasks built by fellow experts and share feedback on clarity, correctness, and difficulty. • Learn from failures: See how AI agents attempted your tasks and help the research team understand what tripped them up. — 3. Core Qualifications • MSc or PhD in computer science or another STEM field, or equivalent practical experience in a research-heavy domain requiring significant coding and data analysis. • 1+ years of experience in a research, research-engineering, or software engineering role. • Strong hands-on Python scripting and debugging skills, with clean-code habits and attention to readability. • Everyday fluency with version control (Git), IDEs, and standard software development workflows. • Experience with AI coding assistants, prompt engineering, or agent workflows is preferred. • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
STEM / US Remote

Architecture Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced architects to help evaluate and improve how AI systems understand and reason about architecture topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in architecture contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of architectural expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how architects actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Degree-qualified architect holding a professional architecture license in the US. • 2+ years of professional architecture experience, ideally across more than one project type (residential, commercial, institutional). • Familiarity with building codes and standards (e.g. IBC) and common tools (Revit, AutoCAD, BIM, or similar). • A portfolio of built or realized work — not purely academic or conceptual projects — is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Code / US Remote

Civil Engineering Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced civil engineers to help evaluate and improve how AI systems understand and reason about civil engineering topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in civil engineering contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of civil engineering expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how civil engineers actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Degree-qualified civil engineer holding a Professional Engineer (PE) license in the US. • 2+ years of professional civil engineering experience, ideally across more than one area (structural, geotechnical, transportation, water resources). • Familiarity with relevant codes and standards (e.g. ASCE, IBC, AASHTO) and common engineering tools (AutoCAD Civil 3D, STAAD, Revit, or similar). • Some experience writing technical specs, design reports, or training material is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Code / US Remote

Chemical Engineering Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced chemical engineers to help evaluate and improve how AI systems understand and reason about chemical engineering topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in chemical engineering contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of chemical engineering expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how chemical engineers actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Degree-qualified chemical engineer holding a Professional Engineer (PE) license in the US. • 2+ years of professional chemical engineering experience, ideally across more than one area (process design, process safety, plant operations). • Familiarity with process safety standards (e.g. OSHA PSM, API) and process simulation tools (Aspen, MATLAB, or similar). • Some experience writing process documentation, SOPs, or technical training material is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Code / US Remote

Mechanical Engineering Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced mechanical engineers to help evaluate and improve how AI systems understand and reason about mechanical engineering topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in mechanical engineering contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of mechanical engineering expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how mechanical engineers actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Degree-qualified mechanical engineer holding a Professional Engineer (PE) license in the US. • 2+ years of professional mechanical engineering experience, ideally across more than one area (thermodynamics/HVAC, structural/mechanical design, manufacturing). • Familiarity with relevant codes and standards (e.g. ASME) and common engineering tools (SolidWorks, ANSYS, AutoCAD, or similar). • Some experience writing technical documentation, specs, or training material is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Code / US Remote

Electrical Engineering Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced electrical engineers to help evaluate and improve how AI systems understand and reason about electrical engineering topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in electrical engineering contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of electrical engineering expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how electrical engineers actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Degree-qualified electrical engineer holding a Professional Engineer (PE) license in the US. • 2+ years of professional electrical engineering experience, ideally across more than one area (power systems, controls, electronics, signal processing). • Familiarity with relevant codes and standards (e.g. NEC, IEEE) and common engineering tools (MATLAB, AutoCAD Electrical, SPICE, or similar). • Some experience writing technical documentation, specs, or training material is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Finance / United States Remote

Finance Specialist

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Finance subject-matter experts (SMEs) with deep domain expertise and hands-on experience evaluating AI model outputs against rubrics to bring rigor and real-world financial judgment to our AI training data. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps in financial reasoning, analysis, and decision-making. • Design challenging, domain-relevant finance tasks and write accurate, well-reasoned solutions grounded in real financial practice. • Evaluate AI model outputs against structured rubrics and provide clear, written feedback on correctness, judgment, and reasoning quality. • Develop and refine evaluation guidelines and scoring rubrics specific to finance tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 8+ years of dedicated professional experience in finance (e.g., investment banking, asset management, corporate finance, financial advisory) at a recognized, top-tier organization (e.g., Goldman Sachs, JPMorgan, Morgan Stanley, BlackRock, Fidelity, Deloitte, PwC, EY, KPMG, or equivalent). • Prior hands-on experience evaluating LLM/AI model outputs against rubrics or structured scoring criteria — mandatory; please describe this experience in your application. • Demonstrable career progression (e.g., Analyst → Associate → VP/Director). • Ability to engage reliably for at least 35 hours/week during weekdays. • Verbal and written communication skills, problem-solving skills, and interpersonal skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Multimodal / Remote, United States (PST to EST hours)

Visual Quality Expert (Film, VFX & Animation)

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking senior image quality experts from film, VFX, animation, and professional photography — VFX and rendering supervisors, colorists, cinematographers, lighting artists, and high-end photographers — with a strong foundation in visual perception, light physics, and pixel-level scrutiny. You will judge, frame by frame, whether an enhanced or upscaled image actually holds up, catching the compression artifacts, noise, aliasing, grain-structure inconsistency, and over- or under-sharpening that typically go unnoticed by the average viewer. — This is a part-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Evaluate high-resolution video and stills at pixel level, identifying compression artifacts, noise, aliasing, banding, grain-structure inconsistency, and over- or under-sharpening. • Apply an uncompromising focus on the intricate details that typically go unnoticed, and flag where output diverges from professional image-quality standards. • Design and refine evaluation templates and rating guidelines so reviewers apply the same visual-quality bar consistently, including under edge cases. • Translate complex visual details into clear, precise descriptive language that can be captured in data captions and evaluation templates. • Generate your own visual samples for targeted testing and side-by-side comparison. • Collaborate with other subject matter experts and program leads to maintain visual consistency across sequences and across reviewers. — 3. Core Qualifications • 5+ years of professional experience in high-end visual evaluation across film, VFX, animation, or professional photography. • Direct professional experience in at least one of the following disciplines: • Visual Effects (VFX) Supervisor — the authority on final image quality, with an elite eye for pixel-level flaws, grain-structure inconsistencies, and visual artifacts. • Rendering Supervisor — accustomed to running dailies and taking ultimate responsibility for visual quality down to the absolute pixel. • Lighting Supervisor, Lead, or Artist — a specialist in illumination and the foundational building blocks of digital video and cinematic motion. • Colorist — deep expertise in color grading, artifact detection, and maintaining visual consistency across sequences. • Director of Photography (DP) or Cinematographer — skilled at framing, capturing visual performance, and critically evaluating over- and under-sharpness in high-resolution images. • Effects or Surfacing Artist — strong command of textures, materials, and the complex simulation of visual elements. • High-End Photographer — deep command of lighting dynamics, depth of field, and color artifacts in high-resolution captures. • Director — particularly those with deep experience directing crowd and background action. • Digital Projectionist — a rigorously trained eye for final-output quality, screen-level fidelity, and visual anomalies. • Training from a university or program with a highly respected film, visual effects, or digital media program. • Reliable access to a 4K-resolution monitor or display to perform precise, pixel-level evaluation. • Familiarity with AI tools and workflows, including the ability to generate your own visual samples for testing and comparison. • Ability to engage reliably for 20 hours per week, with working hours overlapping the PST to EST window. • Fluent in English, with the exceptional ability to translate complex visual details into clear, precise descriptive language. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$60 - $90 / hourOpen / Referral verified
Code / North America (US & Canada) Remote

CAD Engineer — ScreenSpot Plus (Screenshot Capture & UI Annotation)

About the Role — You've been selected for ScreenSpot Plus, where you'll capture screenshots of professional CAD software in realistic, expert use, annotate interactive UI elements, and write natural-language task instructions. — The project begins with a pilot phase, during which your initial submissions will be closely reviewed for quality before you ramp into full production. Detailed onboarding materials and capture guidelines will be shared once you accept. — What You'll Do • Capture high-quality screenshots of professional CAD software during realistic, expert workflows • Annotate interactive UI elements (buttons, menus, panels, toolbars, dialogs) accurately and consistently • Write clear, natural-language task instructions that reflect how an expert actually uses the software • Iterate on feedback during the pilot phase to meet quality standards before scaling to full production — Who We're Looking For • Hands-on, professional experience with CAD software (e.g., AutoCAD, SolidWorks, Fusion 360, CATIA, Revit, Siemens NX, or similar) • Strong working knowledge of the day-to-day workflows and interface of your CAD tools • Attention to detail and the ability to follow precise annotation and capture guidelines • Clear written English for task instructions • Reliable access to the relevant CAD software for capturing screenshots

$70 - $90 / hourOpen / Referral verified
Legal / Remote

Public Interest Law (Civil Law / Environmental Law)

Role Overview • Mercor is seeking senior public interest law professionals to build evaluation tasks for AI systems operating in civil rights, environmental, and public advocacy contexts. • The workflows are calibrated to the societal stakes, regulatory scope, and community impact of major public interest litigation and advocacy campaigns. • This role builds worlds on two tracks: a US track (federal civil rights statutes, NEPA, Clean Air Act and Clean Water Act) and an International track (international human rights law, EU environmental directives). Experts qualified in either or both tracks are encouraged to apply. • Contributors design public interest law scenarios, draft reference outputs, and write rubrics that capture how senior public interest attorneys think. — Key Responsibilities • Construct scenarios spanning civil rights litigation, environmental compliance and enforcement, and community advocacy or policy reform processes. • Build tasks across civil rights and constitutional law, environmental law and regulation, housing and consumer protection, immigration and asylum, and government accountability. • Develop scenarios involving tools such as legal research platforms, case management systems used by legal aid organizations and nonprofits, and regulatory filing systems (e.g., EPA dockets). • Apply public interest law methodologies (constitutional analysis, regulatory compliance review, impact litigation strategy) to the standards track a world targets (US: federal civil rights statutes, NEPA, environmental statutes; International: human rights treaties, EU directives), and produce reference legal briefs, regulatory comments, and advocacy memoranda. • Author rubrics that distinguish authentic public interest legal judgment from generic law school or bar exam-level recall. — Ideal Qualifications • 5+ years working as a public interest attorney at a legal aid organization, advocacy nonprofit, or government agency (ACLU, Earthjustice, NRDC, Legal Aid Society, or a state attorney general's office). • Direct ownership of impact litigation, regulatory advocacy, or policy reform initiatives. • Fluency in public interest legal tooling, plus understanding of how regulatory processes and legislative advocacy actually work. • A recognized professional credential is strongly preferred (JD with bar admission, or an international equivalent); prior rubric, training, or policy authorship is a plus.

$90 - $100 / hourOpen / Referral verified
Medical / Remote

Medical Revenue Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Charge Capture, Charge Integrity, and Revenue Integrity professionals to evaluate AI tools designed to identify revenue leakage, improve charge accuracy, and ensure coding and billing compliance. Your expertise in charge-capture workflows, CDM management, and revenue-integrity analysis will help us build AI systems that optimise revenue performance and reduce compliance risk. — Responsibilities • Oversee charge capture, charge integrity, and revenue integrity functions to ensure accurate and compliant charge submission. • Evaluate AI-generated charge review outputs, coding recommendations, and revenue integrity alerts for accuracy and compliance. • Conduct charge audits to identify missed charges, duplicate charges, and charge capture errors across clinical departments. • Manage and maintain the charge description master (CDM), ensuring accuracy of charge codes, revenue codes, and pricing. • Analyse charge patterns to identify revenue leakage opportunities and implement corrective action plans. • Collaborate with clinical, coding, and billing teams to resolve charge capture discrepancies. • Monitor KPIs including charge lag times, charge capture accuracy rates, and revenue integrity findings. • Ensure compliance with CMS billing guidelines, OIG work plan priorities, and payer-specific requirements. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in charge capture, charge integrity, revenue integrity, or healthcare compliance. • Deep knowledge of CDM management, revenue codes, charge capture workflows, and billing compliance. • Strong understanding of CMS billing guidelines, Medicare Part A/B billing rules, and OIG compliance requirements. • Proficiency with charge capture systems, EHR platforms, and revenue integrity tools. • Experience conducting charge audits and developing corrective action plans. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify billing errors and compliance risks in AI-generated outputs. — Preferred Qualifications • CPC, CCS, CHFP, or CRCS credential. • Experience with 340B drug program charge capture and compliance. • Background in hospital, health system, or physician group revenue integrity operations. • Familiarity with AI tools and comfort evaluating AI-generated charge and billing content. • Experience with revenue integrity software platforms and analytics tools. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in healthcare revenue integrity. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$88 / hourOpen / Referral verified
Legal / Remote

Corporate Law Expert

Role Overview • Mercor is seeking senior corporate law professionals to build evaluation tasks for AI systems operating in complex corporate transactions and governance contexts. • The workflows are calibrated to the deal complexity, stakeholder stakes, and regulatory scope of major M&A transactions, capital markets deals, and corporate governance matters. • This role builds worlds on two jurisdictional tracks: a US track (Delaware General Corporation Law, SEC regulations, Model Business Corporation Act) and an International track (UK Companies Act 2006, EU company law directives). Experts qualified in either or both tracks are encouraged to apply. • Contributors design corporate law scenarios, draft reference outputs, and write rubrics that capture how senior corporate lawyers think. — Key Responsibilities • Construct corporate law scenarios spanning large-scale M&A transactions, multi-party deal negotiation and regulatory review, and complex corporate governance or restructuring processes. • Build tasks across M&A and deal structuring, securities and capital markets, corporate governance, commercial contracts, and corporate restructuring. • Develop legal scenarios involving tools such as Westlaw, Lexis, virtual data rooms (Intralinks, Datasite), contract lifecycle management platforms, and document automation systems used on major transactions. • Apply corporate law methodologies (deal structuring, due diligence review, regulatory filing analysis) to the jurisdictional track a world targets (US: Delaware General Corporation Law, SEC rules, Model Business Corporation Act; International: UK Companies Act 2006, EU company law directives), and produce reference legal memoranda, transaction documents, and client/regulatory-facing narratives. • Author rubrics that distinguish authentic corporate law judgment from generic law school or bar exam-level recall. — Ideal Qualifications • 5+ years working as a corporate attorney or general counsel at a major law firm, investment bank, or corporation (Wachtell, Skadden, Cravath, Latham & Watkins, Kirkland & Ellis, Sullivan & Cromwell, or in-house legal at a large public company). • Direct ownership of M&A transactions, governance matters, or securities filings. • Fluency in corporate law tooling and methodologies, plus understanding of how regulatory approvals (SEC review, antitrust/HSR, or an international equivalent) and board/shareholder oversight actually work. • A recognized professional credential is strongly preferred (JD with bar admission, or an international equivalent such as Solicitor or Qualified Lawyer in England & Wales); prior rubric, legal training curriculum, or transaction documentation authorship is a plus.

$90 - $100 / hourOpen / Referral verified
Code / Remote

Litigation Expert

Role Overview • Mercor is seeking senior litigation professionals to build evaluation tasks for AI systems operating in civil litigation and dispute resolution contexts. • The workflows are calibrated to the case complexity, evidentiary stakes, and procedural scope of major commercial litigation and complex disputes. • This role builds worlds on two tracks: a US track (Federal Rules of Civil Procedure, Federal Rules of Evidence, state court analogs) and an International track (English Civil Procedure Rules, international arbitration rules such as ICC and LCIA). Experts qualified in either or both tracks are encouraged to apply. • Contributors design litigation scenarios, draft reference outputs, and write rubrics that capture how senior litigators think. — Key Responsibilities • Construct litigation scenarios spanning pre-trial discovery, motion practice, trial strategy, and settlement or arbitration processes. • Build tasks across commercial litigation, complex and class action litigation, discovery and evidence management, trial advocacy, and appellate practice. • Develop scenarios involving tools such as e-discovery platforms (Relativity, Everlaw), litigation management software, and deposition/trial preparation tools used on major matters. • Apply litigation methodologies (case strategy development, discovery and evidence analysis, motion drafting) to the standards track a world targets (US: FRCP, Federal Rules of Evidence, state court rules; International: CPR, international arbitration rules), and produce reference pleadings, discovery responses, motions, and trial memoranda. • Author rubrics that distinguish authentic litigation judgment from generic law school or bar exam-level recall. — Ideal Qualifications • 5+ years working as a litigator or trial attorney at a major litigation firm or corporate litigation department (Quinn Emanuel, Gibson Dunn, Boies Schiller Flexner, Kirkland & Ellis, or in-house litigation counsel). • Direct ownership of complex commercial litigation matters, with trial or arbitration experience. • Fluency in litigation tooling and methodologies, plus understanding of how court procedure, discovery, and evidentiary rules actually work. • A recognized professional credential is strongly preferred (JD with bar admission, or an international equivalent such as Barrister or Solicitor with advocacy rights); prior rubric or training authorship is a plus.

$90 - $100 / hourOpen / Referral verified
STEM / Remote

Biology Research Scientist (BA, MS, PhD's)

About the Role — Mercor is partnering with a leading AI research organization to verify protein target assignments in large-scale bioactivity databases (ChEMBL, BindingDB). You will read primary literature, apply scientific judgment, and determine whether UniProt IDs accurately reflect the proteins being studied — work that directly feeds AI drug discovery pipelines. — What You'll Do • Access primary sources (papers, patents) to verify protein target assignments against UniProt records • Flag and classify target assignment errors using a structured taxonomy • Propose correct UniProt accessions where assignments are wrong • Write concise, evidence-grounded notes explaining your reasoning — Requirements • BA/BS with 5+ years, MS with 2+ years, or PhD with industry or drug discovery research experience, in pharmacology, biochemistry, molecular biology, or chemical biology at a biotech, pharma, or CRO • Hands-on binding or functional assay experience (SPR, TR-FRET, radioligand binding, kinase assays, GPCR functional assays, IC50/Ki/KD) • Currently bench-active in a research, scientist, or associate scientist role • Working fluency with UniProt or adjacent workflows: SAR support, HTS, target validation, biochemical profiling, or IND-enabling studies — Nice to Have • Direct experience with ChEMBL, BindingDB, or PubChem • Selectivity profiling or counterscreening experience • Familiarity with agonist/antagonist vs. activator/inhibitor distinctions — Role Details • 10–20 hrs/week | Remote | Immediate start • 1–2 month minimum, extension likely • U.S. only

$50 - $70 / hourOpen / Referral verified
STEM / Remote

Generalist Expert

Overview — In this role, you will evaluate AI-generated responses and provide structured written feedback. This is a great opportunity for sharp, analytical thinkers to contribute to high-impact AI research projects. — Basic Qualifications • Bachelor's degree from a top-500 globally ranked university preferred • Strong analytical and written communication skills • Ability to work independently and follow detailed task guidelines — Required Skills • Strong critical reading skills with the ability to identify nuance, implicit meaning, and gaps in reasoning • Ability to write clear, precise, and well-evidenced written rationales that go beyond surface-level observations • Consistent and honest judgment, including the ability to give critical assessments when warranted • Strict attention to detail and accurate application of structured evaluation guidelines • Ability to work entirely without AI writing tools — Eligibility • Native English fluency required

$70 / hourOpen / Referral verified
Code / Remote

Computational Electrical Engineering & RF/Circuit Design Expert

Electrical Engineering & RF/Circuit Design Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience in: — Electrical Engineering & RF/Circuit Design — working with scikit-rf for RF and microwave network analysis, S-parameter characterization, and transmission-line modeling, or ngspice for circuit simulation, operating point analysis, and frequency response characterization. Candidates should be comfortable designing problems that involve recovering circuit parameters from measurement data. — _Experience with other specialized software in this domain will also be considered._ — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $85 / hourOpen / Referral verified
Code / Remote

Computational Structural & Mechanical Engineering Expert

Computational Structural & Mechanical Engineering Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard computational scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience with open-source, domain-specific computational tools such as FEniCSx/DOLFINx, scikit-fem, OpenFOAM, deal.II, MFEM, MOOSE, CalculiX, Elmer FEM, Code\_Aster, SfePy, FiPy, Devito, Cantera, CoolProp, Pyomo, or SimPy, for finite-element analysis, computational mechanics, structural analysis, elasticity, CFD, multiphysics simulation, heat and mass transfer, thermodynamics, combustion, fluid mechanics, HVAC/thermal systems, manufacturing simulation, optimization, or thermophysical-property calculations. — Relevant work may include beam, plate, and shell analysis; linear or nonlinear elasticity; finite-element and variational formulations; mesh refinement and convergence studies; continuum and solid mechanics; computational fluid dynamics; coupled multiphysics problems; thermal-fluid simulation; structural or system optimization; reliability analysis; and related numerical engineering workflows. — Experience with underlying theories and numerical methods — such as Euler–Bernoulli and Timoshenko beam theory, continuum mechanics, finite-element methods, Galerkin/variational methods, finite-volume methods, PDE discretization, constitutive modeling, thermodynamics, numerical linear algebra, and nonlinear solution methods — is valuable. — Experience with other open-source computational structural or mechanical engineering software will also be considered, including scientific codes and solver frameworks built with Python, C, C++, or Fortran. — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $85 / hourOpen / Referral verified
STEM / Remote

Computational Seismology & Geophysics Expert

Computational Seismology & Geophysics Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience in: — Computational Seismology & Geophysics — hands-on experience with open-source, domain-specific computational tools such as SPECFEM, ObsPy, Pyrocko, SimPEG, pyGIMLi, SeisBench, EQcorrscan, or Fatiando a Terra, for seismic wave propagation and numerical simulation, synthetic seismogram generation, full-waveform inversion (FWI), seismic imaging, travel-time tomography, moment tensor inversion, event detection/location, or related computational geophysics workflows. — Experience with other open-source computational seismology or geophysics software will also be considered, including tools and scientific codes built with Python, C, C++, or Fortran. — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $85 / hourOpen / Referral verified
STEM / Remote

Computational Pharmacokinetics & Systems Biology Expert

Computational Pharmacokinetics & Systems Biology Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience in: — Pharmacokinetics & Systems Biology — working with libRoadRunner, Tellurium, or SBML-based tools for compartmental PK/PD modeling, enzyme kinetics, or systems biology simulations. — _Experience with other specialized software in this domain will also be considered._ — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $85 / hourOpen / Referral verified
Finance / Remote

Payment-posting & Reconciliation Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Cash Posting and Payment Reconciliation Managers to evaluate AI tools designed to automate electronic remittance processing, payment posting, and cash reconciliation workflows. Your expertise in ERA processing, EOB interpretation, and payment reconciliation will directly inform AI systems that improve cash posting accuracy and accelerate revenue cycle close. — Responsibilities • Oversee cash posting and payment reconciliation operations, including electronic remittance advice (ERA/835) processing, manual EOB posting, and lockbox reconciliation. • Evaluate AI-generated payment posting outputs, ERA matching recommendations, and reconciliation reports for accuracy and completeness. • Manage electronic and manual payment posting workflows across multiple payers and payment types. • Reconcile posted payments against bank deposits, lockbox reports, and payer remittances to ensure accuracy. • Identify and resolve posting errors, misapplied payments, and unapplied cash. • Monitor cash posting KPIs, including posting accuracy rates, days to post, unapplied cash balances, and reconciliation variance. • Collaborate with billing, A/R, and finance teams to resolve payment discrepancies and ensure timely cash close. • Ensure compliance with internal controls, HIPAA, and audit requirements for cash handling. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in cash posting, payment reconciliation, or revenue cycle operations, with at least 2 years in a management role. • Deep knowledge of ERA/835 electronic remittance processing, EOB interpretation, and lockbox reconciliation. • Strong understanding of payer payment methodologies and remittance adjustment reason codes. • Experience managing high-volume payment posting operations across multiple payers. • Proficiency with billing systems and payment posting platforms (Epic, Athenahealth, or equivalent). • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify posting errors and discrepancies in AI-generated payment outputs. — Preferred Qualifications • CRCR, CPC, or CHFP certification. • Experience with automated ERA posting platforms and RPA-driven cash posting solutions. • Background in hospital or physician group cash posting operations with multi-payer complexity. • Familiarity with AI tools and comfort evaluating AI-generated remittance and payment content. • Experience developing cash posting SOPs and internal control frameworks. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in payment reconciliation and revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$85 / hourOpen / Referral verified
Finance / Remote

Underpayment & Managed-care Contract Specialist

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Underpayment and Managed Care Contract Specialists to evaluate AI tools designed to detect payment variances and automate contract compliance monitoring. Your expertise in payer contract interpretation, payment variance analysis, and underpayment recovery will help AI systems optimise net revenue and identify systematic underpayments. — Responsibilities • Identify, analyse, and recover underpayments and payment variances across commercial, Medicare Advantage, and Medicaid managed care payer contracts. • Evaluate AI-generated underpayment detection alerts, contract compliance outputs, and payment variance analyses for accuracy. • Interpret payer contract terms including fee schedules, carve-outs, outlier provisions, and payment methodologies to validate claim payments. • Conduct contract modelling and payment reconciliation to identify systematic underpayment patterns. • Develop and submit underpayment claims and recovery correspondence to payers. • Collaborate with managed care contracting teams to identify contract language gaps and renegotiation opportunities. • Monitor underpayment recovery KPIs including identified underpayment dollars, recovery rates, and payer response rates. • Ensure compliance with payer contract terms, HIPAA, and timely filing requirements for underpayment recovery. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in underpayment recovery, payment variance analysis, or managed care contracting, with demonstrated expertise in contract interpretation. • Deep knowledge of payer payment methodologies including DRG-based, per diem, per cent-of-billed charges, and fee schedule reimbursement. • Strong experience with contract modelling, payment reconciliation, and payer dispute resolution. • Familiarity with managed care contract management systems and payment variance detection tools. • Proficiency with Excel-based financial analysis and revenue cycle analytics platforms. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify payment discrepancies in complex contract structures and AI-generated outputs. — Preferred Qualifications • CHFP, CRCR, or managed care contracting certification. • Experience with contract management software (e.g., Experian Health, Recondo, or similar). • Background in hospital, health system, or large physician group underpayment recovery operations. • Familiarity with AI tools and comfort evaluating AI-generated payment variance content. • Experience negotiating payer contracts and presenting underpayment analysis to executive leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in managed care contracting and the revenue cycle. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$85 / hourOpen / Referral verified
Code / Remote

Applied Computer Science Benchmark Specialist

Role Overview — We are seeking expert computer scientists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core computer science domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of computer science expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Computer Science Domains Covered — Accelerator / GPU Kernel Engineering, Formal Methods & Automated Reasoning, Computer Architecture & Accelerators, Distributed Systems, DevOps & Site Reliability, Data Engineering & Databases, Cloud & Infrastructure, OS & Systems Kernel, Machine Learning Engineering, Web & API Development, Embedded Systems Engineering, Computer Graphics & Game Development, Mobile Engineering. — Key Responsibilities • Author original computer science questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Computer Science, Electrical Engineering, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level CS theory, algorithms, systems design, and/or machine learning • Research publications, industry experience at top tech companies, or competitive programming background is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$66 - $84 / hourOpen / Referral verified
Code / Remote

Atomic Layer Deposition (ALD) Experts

Mercor is seeking experts in Atomic Layer Deposition (ALD) and thin-film processes to support a frontier AI research lab building models for semiconductors and the physical sciences. This is hands-on, expert-level work: you'll apply deep, specialized knowledge to generate, structure, and evaluate the scientific data these models learn from — and your input will directly shape how advanced models reason about deposition, materials, and semiconductor processes. — Key Responsibilities: • Contribute domain expertise across ALD process development, precursor chemistry, and thin-film characterization to build high-quality training and evaluation data. • Review and evaluate AI-generated scientific reasoning, catching errors and improving technical accuracy. • Design and solve challenging, expert-level problems in ALD and semiconductor processing. • Rate and rank model outputs against defined scientific criteria, with clear written reasoning. • Structure technical knowledge — process parameters, recipes, characterization results — into well-organized, model-ready data. • Deliver reliable, high-quality work within defined timelines. — You're a strong fit if you have: • Hands-on experience developing, optimizing, or troubleshooting ALD processes. • Deep knowledge of thin-film deposition for semiconductor or advanced-packaging applications. • A strong grasp of precursor chemistry and surface reaction mechanisms. • Experience with materials characterization (XRD, SEM, TEM, XPS, ellipsometry, etc.). • An advanced degree (PhD/MS) or equivalent hands-on experience in materials science, chemistry, chemical engineering, or physics. • Clear written English and the ability to explain technical reasoning precisely. — Role Details: • Type: Long-term, ongoing engagement • Engagement: Up to 40 hours/week (minimum 10) • Work arrangement: Remote (US-based)

$84 / hourOpen / Referral verified
Policy & Safety / Remote

AI Safety Red Teamer

We are seeking experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous ("grey-area") topics. — Responsibilities • Design adversarial prompts to stress-test frontier AI models. • Identify jailbreaks, unsafe behaviours, hallucinations, and policy failures. • Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains. • Document vulnerabilities and contribute to safety benchmarking and red-teaming reports. • Collaborate with AI researchers to improve model alignment, robustness, and safety. — Required Qualifications • Bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related discipline. • 5+ years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a related field. • Strong analytical reasoning, prompt design, and written communication skills. • Experience designing adversarial prompts or evaluating frontier AI systems. — Preferred Qualifications • Experience with AI Red Teaming, RLHF, SFT, AI Alignment, or Trust & Safety. • Familiarity with jailbreak testing, prompt engineering, or adversarial evaluation methodologies. • Expertise in one or more grey-area domains, including cyber, biosecurity, political content, misinformation, or scientific safety. — Why Join? • Help secure and strengthen the next generation of frontier AI models. • Work on cutting-edge adversarial testing alongside leading AI researchers and safety teams. • Influence how AI systems respond to complex, real-world safety challenges.

$70 - $84 / hourOpen / Referral verified
Medical / Remote

Clinical Documentation Integrity (CDI) Leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Clinical Documentation Integrity (CDI) managers and leaders to evaluate AI tools designed to enhance clinical documentation accuracy and coding integrity workflows. Your expertise in DRG optimisation, HCC risk adjustment, physician query management, and clinical documentation best practices will directly inform AI systems that improve documentation quality, compliance, and revenue integrity. — Responsibilities • Lead clinical documentation integrity programs for inpatient and/or outpatient settings, overseeing concurrent and retrospective review workflows. • Evaluate AI-generated clinical documentation improvement suggestions, physician queries, and coding recommendations for clinical accuracy and compliance. • Conduct and review clinical documentation to ensure accurate capture of diagnoses, procedures, severity of illness, and risk of mortality. • Develop and manage physician query processes in alignment with AHIMA and ACDIS guidelines. • Monitor CDI program KPIs including query response rates, CC/MCC capture rates, case mix index, and DRG accuracy. • Collaborate with coding, compliance, and clinical teams to address documentation gaps and improve query processes. • Provide education to physicians and clinical staff on documentation requirements and best practices. • Ensure compliance with Official Coding Guidelines, CMS regulations, and payer-specific requirements. • Annotate AI outputs and provide structured clinical feedback to support AI training datasets. — Requirements • 5+ years of experience in clinical documentation integrity or improvement, with at least 2 years in a manager or leadership role. • Deep knowledge of MS-DRG methodology, CC/MCC hierarchies, and ICD-10-CM/PCS coding guidelines. • Expertise in physician query management per AHIMA and ACDIS compliant query standards. • Strong clinical background with the ability to interpret medical records and clinical documentation. • Proficiency with CDI software platforms (3M, Nuance, Optum360, or equivalent) and EHR systems. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to critically evaluate clinical documentation and AI-generated outputs. • Comfortable working independently in a fully remote environment. — Preferred Qualifications • CCDS (Certified Clinical Documentation Specialist) or CDIP (Clinical Documentation Improvement Practitioner) credential. • Experience with HCC risk adjustment and outpatient CDI programs. • Background in RN, RHIA, CCS, or similar clinical or coding credential. • Familiarity with AI-assisted CDI tools (e.g., Nuance DAX, 3M M\*Modal) and comfort evaluating AI-generated clinical content. • Experience presenting CDI performance metrics to clinical and executive leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in clinical documentation and revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$84 / hourOpen / Referral verified
Code / Remote

ML Engineer (Coding Agent Experience)

About the Role — \- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. — \- Contributors help evaluate and improve frontier AI coding models through structured technical assessments. — \- The work focuses on realistic machine learning engineering workflows and model evaluation. — \- Spots are limited and filling quickly on a first come, first serve basis. — What You'll Do — \- Use frontier AI coding agents to complete and evaluate complex machine learning and AI engineering tasks. — \- Review model-generated implementations involving model training, inference systems, MLOps, and LLM applications. — \- Identify bugs, edge cases, performance issues, and failure modes. — \- Compare outputs from multiple frontier models and assess their strengths and weaknesses. — \- Apply professional engineering judgment to realistic ML engineering scenarios. — Time Commitment — \- Sprint based project that runs in 12-24 hour stretches based on client requirement. — Compensation — \- $400 per accepted task. — \- Typical tasks take approximately 2–3 hours after ramp-up. — \- Compensation is tied to accepted work. — Who Should Apply — \- 2+ years of professional machine learning engineering experience. — \- Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products. — \- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. — \- Ability to evaluate model-generated machine learning implementations and technical tradeoffs. — \- Experience deploying ML systems to production is preferred.

$85 / hourOpen / Referral verified
Code / Remote

ML Research PhD Experts (ICML / NeurIPS / ICLR Publications)

Overview — We're looking to rapidly assemble a small group of world-class machine learning researchers for an initial pilot. This is a high-priority engagement with an accelerated timeline, so both exceptional candidate quality and fast turnaround are critical. We're specifically seeking researchers with demonstrated contributions to frontier ML research, particularly those driving algorithmic innovation rather than applied analytics. — Candidate Requirements — Required Qualifications • PhD in Machine Learning, Computer Science, AI, or a closely related field • Published at least one main conference paper at ICML, NeurIPS, or ICLR • Strong preference for candidates with 2+ publications at these venues • Demonstrated experience conducting original ML research — Preferred Research Areas • Reinforcement Learning (RL) • Meta-Learning • Recursive Self-Improvement • AI for Science (e.g. weather forecasting, protein modeling, scientific discovery) — Timeline — Priority: Urgent — Expected Schedule • Initial data/results: Monday / Tuesday 27th July • * * — Success Criteria — Candidates should have a proven record of advancing state-of-the-art machine learning through publications at top-tier conferences and possess deep expertise in frontier ML research. Speed of sourcing is important, but quality should not be compromised.

$60 - $100 / hourOpen / Referral verified
Code / Remote

DevOps / SRE / Cloud Engineer (Coding Agent Experience)

About the Role — \- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. — \- Contributors help evaluate and improve frontier AI coding models through structured technical assessments. — \- The work focuses on realistic infrastructure engineering workflows and model evaluation. — \- Spots are limited and filling quickly on a first come, first serve basis. — What You'll Do — \- Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. — \- Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation. — \- Identify bugs, edge cases, reliability issues, and failure modes. — \- Compare outputs from multiple frontier models and assess their strengths and weaknesses. — \- Apply professional engineering judgment to realistic infrastructure engineering scenarios. — Time Commitment — \- Sprint based project that runs in 12-24 hour stretches based on client requirement. — Compensation — \- $400 per accepted task. — \- Typical tasks take approximately 2–3 hours after ramp-up. — \- Compensation is tied to accepted work. — Who Should Apply — \- 2+ years of professional DevOps, SRE, or Cloud Engineering experience. — \- Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling. — \- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. — \- Ability to evaluate model-generated infrastructure and reliability engineering solutions. — \- Experience supporting production-scale systems is preferred.

$85 / hourOpen / Referral verified
Code / Remote

Cybersecurity Expert

Role Overview • Mercor is seeking senior cybersecurity professionals to build evaluation tasks for AI systems operating in security operations, incident response, and risk management contexts. • The workflows are calibrated to the threat sophistication, business risk stakes, and scope of major enterprise security programs. • This role builds worlds on two tracks: a US track (NIST Cybersecurity Framework, SOC 2) and an International track (ISO 27001, EU NIS2 Directive). Experts qualified in either or both tracks are encouraged to apply. • Contributors design cybersecurity scenarios, draft reference outputs, and write rubrics that capture how senior security leaders think. — Key Responsibilities • Construct cybersecurity scenarios spanning security operations center monitoring, incident response and forensics, vulnerability management, and security architecture design. • Build tasks across security operations and threat detection, incident response and digital forensics, vulnerability and penetration testing, security architecture and engineering, and governance/risk/compliance. • Develop scenarios involving tools such as SIEM platforms (Splunk, Microsoft Sentinel), EDR tools (CrowdStrike), vulnerability scanners (Tenable, Qualys), and GRC platforms used at major enterprises. • Apply cybersecurity methodologies (threat modeling, incident response playbooks, risk assessment) to the standards track a world targets (US: NIST CSF, SOC 2; International: ISO 27001, NIS2), and produce reference incident reports, security assessments, architecture designs, and compliance documentation. • Author rubrics that distinguish authentic security judgment from generic textbook or certification-exam-level recall. — Ideal Qualifications • 5+ years working as a security engineer or CISO at a major company or security firm (Mandiant, CrowdStrike, or an in-house CISO/security lead). • Direct ownership of incident response programs, security architecture, or compliance initiatives. • Fluency in security tooling, plus understanding of regulatory frameworks and the current threat landscape. • A recognized professional credential is strongly preferred (CISSP, CISM, or an international equivalent); prior rubric or training authorship is a plus.

$80 - $90 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Spanish - Mexico)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Mexico-based voice actors with native Mexican Spanish accents to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native Mexican Spanish speaker currently based in Mexico • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $100 / hourOpen / Referral verified
Medical / US Remote

Healthcare Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced healthcare professionals to help evaluate and improve how AI systems understand and reason about healthcare topics. You'll bring your real-world clinical expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in healthcare contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of clinical expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how healthcare professionals actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Licensed medical doctor (MD). • 2+ years of professional clinical experience; broad or generalist clinical exposure is a plus over a narrow subspecialty. • Comfortable engaging with medical literature and clinical guidelines (e.g. UpToDate, PubMed, published research). • Prior teaching or training experience (e.g. supervising residents, CME instruction, medical education) is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$90 - $100 / hourOpen / Referral verified
Business / Remote

Procurement Expert

About the Role — Mercor is building realistic, high-fidelity simulated environments to evaluate and train AI models on real-world procurement workflows for a leading spend-management technology company. We're looking for senior procurement and vendor-management professionals to author and validate tasks that mirror how procurement teams actually vet new vendors and review contract renewals. — Key Responsibilities • Review simulated new-vendor intake requests, benchmarking price and contract terms against realistic company reference data • Consolidate multi-lens reviews (finance, legal, security, IT) into a well-reasoned final approval recommendation • Author step-level rubrics and golden responses that capture how an experienced procurement lead would judge a request or renewal • Audit simulated company environments for realism and unintended contradictions before tasks open — Ideal Qualifications • 8+ years of professional experience in procurement, spend operations, or vendor management, with hands-on experience running real vendor requests and renewals • Sourced from mature procurement functions at companies with formal vendor-review processes • Strong written communication skills; comfortable producing structured, rubric-style feedback — Nice to Have • Experience with modern spend-management or e-procurement platforms • Prior task-writing, rubric-authoring, or AI-training data experience

$70 - $100 / hourOpen / Referral verified
Business / Remote

Sales & Marketing Experts

Role Overview • Mercor is seeking senior sales and marketing professionals to build evaluation tasks for AI systems operating in Fortune 500 go-to-market contexts. • The workflows are calibrated to the deal sizes, stakeholder complexity, and brand stakes of Fortune 500 and large public companies. • Contributors design enterprise GTM scenarios, draft reference outputs, and write rubrics that capture how senior F500 operators think. — Key Responsibilities • Construct enterprise sales scenarios spanning $1M+ ACV deals, multi-stakeholder buying committees, and complex procurement cycles at F500 accounts. • Build marketing tasks across F500 brand strategy, enterprise ABM, demand generation at scale, lifecycle, and category positioning. • Develop RevOps and GTM scenarios involving Salesforce Enterprise, Marketo, 6sense, Gong, and Outreach in F500 stacks. • Apply enterprise sales methodologies (MEDDIC, Challenger, Force Management) and produce reference deal strategies, account plans, and executive narratives. • Author rubrics that distinguish authentic enterprise GTM judgment from generic playbook recall. — Ideal Qualifications • 5+ years selling, marketing, or running RevOps at a Fortune 500 enterprise software vendor (Salesforce, Oracle, ServiceNow, SAP, Workday, Microsoft, AWS) or inside an F500 brand or marketing organization (P&G, JPMorgan, Unilever, Microsoft, PepsiCo). • Direct ownership of F500 accounts, F500 brand campaigns, or F500 demand programs. • Fluency in enterprise GTM tooling and methodologies, plus understanding of how F500 budgets, procurement, and legal review actually work. • Prior rubric, sales-enablement curriculum, or training-content authorship is a plus. — Compensation Note • Hourly Pay: $65 to $90 per hour, set by Mercor based on demonstrated expertise. • Minimum Commitment: 20 hours per week. • Onboarding via the Mercor Rubric Academy, a paid program that calibrates contributors to the quality bar before live work. • Advancement: strong contributors move into reviewer, lead, and domain SME roles with elevated rates.

$65 - $90 / hourOpen / Referral verified
Business / Remote

Product Management Expert

Role Overview • Mercor is seeking senior product management professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise product contexts. • The workflows are calibrated to the product scale, stakeholder complexity, and market stakes of Fortune 500 and large public company product organizations. • Contributors design enterprise product scenarios, draft reference outputs, and write rubrics that capture how senior F500 product leaders think. — Key Responsibilities • Construct enterprise product scenarios spanning multi-quarter roadmap planning, cross-functional stakeholder alignment, and complex go-to-market or platform decisions at F500 scale. • Build product tasks across F500 product strategy, discovery and user research, pricing and packaging, platform/API product management, and product-led growth. • Develop product operations scenarios involving tools such as Jira/Productboard, Amplitude/Mixpanel, Figma, and enterprise experimentation platforms in F500 stacks. • Apply enterprise product methodologies (JTBD, RICE/opportunity scoring, dual-track agile, OKRs) and produce reference PRDs, roadmap strategies, and executive-level product narratives. • Author rubrics that distinguish authentic enterprise product judgment from generic framework or bootcamp-level recall. — Ideal Qualifications • 5+ years as a product manager or product leader at a Fortune 500 technology or enterprise organization (Google, Amazon, Microsoft, Salesforce, Adobe) or inside an F500 product organization at a non-tech enterprise (JPMorgan, UPS, Unilever, PepsiCo). • Direct ownership of F500-scale products, platforms, or product lines with measurable business impact. • Fluency in enterprise product tooling and methodologies, plus understanding of how F500 budgets, cross-functional governance, and executive stakeholder alignment actually work. • Prior rubric, product-training curriculum, or PRD/documentation authorship is a plus.

$80 - $90 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Spanish - Latin America)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Latin America-based voice actors with native Latin American Spanish fluency and neutral, internationally intelligible delivery to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native Latin American Spanish speaker currently based in Latin America, with the ability to speak neutral, internationally intelligible Spanish — without a strong regional accent • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $100 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Spanish - Peninsular)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Spain-based voice actors with native Castilian/Peninsular Spanish accents to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native Castilian/Peninsular Spanish speaker currently based in Spain • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $100 / hourOpen / Referral verified
Business / United States Remote

Marketing Specialist

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Marketing subject-matter experts (SMEs) with deep domain expertise and hands-on experience evaluating AI model outputs against rubrics to bring rigor and real-world brand/growth judgment to our AI training data. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps in brand strategy, growth marketing, and campaign-reasoning tasks. • Design challenging, domain-relevant marketing tasks and write accurate, well-reasoned solutions grounded in real marketing practice. • Evaluate AI model outputs against structured rubrics and provide clear, written feedback on correctness, judgment, and reasoning quality. • Develop and refine evaluation guidelines and scoring rubrics specific to marketing tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 8+ years of dedicated professional experience in marketing (e.g., brand strategy, growth marketing, performance marketing) at a recognized, top-tier organization (e.g., P&G, Unilever, Nike, Ogilvy, WPP, Omnicom, or equivalent). • Prior hands-on experience evaluating LLM/AI model outputs against rubrics or structured scoring criteria — mandatory; please describe this experience in your application. • Demonstrable career progression (e.g., Marketing Manager → Senior Manager → Director/VP of Marketing). • Ability to engage reliably for at least 35 hours/week during weekdays. • Verbal and written communication skills, problem-solving skills, and interpersonal skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$60 - $80 / hourOpen / Referral verified
Medical / United States Remote

Insurance Specialist

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Insurance subject-matter experts (SMEs) with deep domain expertise and hands-on experience evaluating AI model outputs against rubrics to bring rigor and real-world underwriting judgment to our AI training data. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps in underwriting, claims, and risk-assessment reasoning. • Design challenging, domain-relevant insurance tasks and write accurate, well-reasoned solutions grounded in real underwriting/claims practice. • Evaluate AI model outputs against structured rubrics and provide clear, written feedback on correctness, judgment, and reasoning quality. • Develop and refine evaluation guidelines and scoring rubrics specific to insurance tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 8+ years of dedicated professional experience in insurance (e.g., underwriting, claims, actuarial, risk management) at a recognized, top-tier organization (e.g., AIG, Chubb, Allstate, Progressive, MetLife, Marsh McLennan, or equivalent). • Prior hands-on experience evaluating LLM/AI model outputs against rubrics or structured scoring criteria — mandatory; please describe this experience in your application. • Demonstrable career progression (e.g., Underwriter → Senior Underwriter → VP of Underwriting). • Ability to engage reliably for at least 35 hours/week during weekdays. • Verbal and written communication skills, problem-solving skills, and interpersonal skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$60 - $80 / hourOpen / Referral verified
Medical / United States Remote

Retail Specialist

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Retail subject-matter experts (SMEs) with deep domain expertise and hands-on experience evaluating AI model outputs against rubrics to bring rigor and real-world retail judgment to our AI training data. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps in retail merchandising, category management, and operations reasoning. • Design challenging, domain-relevant retail tasks and write accurate, well-reasoned solutions grounded in real retail practice. • Evaluate AI model outputs against structured rubrics and provide clear, written feedback on correctness, judgment, and reasoning quality. • Develop and refine evaluation guidelines and scoring rubrics specific to retail tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 8+ years of dedicated professional experience in retail (e.g., merchandising, category management, retail operations, buying/planning) at a recognized, top-tier organization (e.g., Amazon, Walmart, Target, Nike, Costco, Home Depot, or equivalent). • Prior hands-on experience evaluating LLM/AI model outputs against rubrics or structured scoring criteria — mandatory; please describe this experience in your application. • Demonstrable career progression (e.g., Category Manager → Senior Manager → Director of Merchandising). • Ability to engage reliably for at least 35 hours/week during weekdays. • Verbal and written communication skills, problem-solving skills, and interpersonal skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$60 - $80 / hourOpen / Referral verified
Medical / Remote

Medical Billing Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Billing and Claims Managers to support the evaluation of AI tools designed to automate and improve medical billing and claims submission workflows. Your expertise in claims processing, payer requirements, EDI transactions, and billing compliance will directly inform AI systems that improve first-pass claim acceptance rates and accelerate revenue cycle performance. — Responsibilities • Oversee end-to-end medical billing and claims submission operations across professional fee and/or facility billing environments. • Evaluate AI-generated billing outputs, claim edits, and coding validations for accuracy and payer compliance. • Manage claims submission workflows including electronic claim generation, clearinghouse edits, and payer-specific billing requirements. • Monitor clean claim rates, rejection rates, and first-pass acceptance rates and develop improvement strategies. • Coordinate with coding, CDI, and collections teams to resolve billing edits and claim rejections. • Ensure compliance with CMS billing guidelines, HIPAA 837 transaction standards, and payer-specific billing rules. • Manage billing staff workload, productivity, and quality performance metrics. • Develop and maintain billing SOPs and payer-specific billing reference guides. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in medical billing and claims management, with at least 2 years in a management role. • Deep knowledge of professional fee (CMS-1500/837P) and/or facility (UB-04/837I) billing requirements. • Expertise in HIPAA 837 transaction standards, clearinghouse operations, and payer-specific billing rules. • Strong understanding of Medicare, Medicaid, and commercial payer billing requirements. • Proficiency with billing platforms (Epic, Athenahealth, AdvancedMD, or equivalent) and clearinghouse tools. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify billing errors and compliance issues in AI-generated outputs. — Preferred Qualifications • CPC, CCS, CHFP, or CRCR certification. • Experience with automated billing platforms and RCM technology implementations. • Background in multi-speciality physician group, hospital, or health system billing operations. • Familiarity with AI tools and comfort evaluating AI-generated billing content. • Experience with payer contract interpretation and billing compliance program management. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in medical billing and revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$80 / hourOpen / Referral verified
Code / Remote

Coding Manager / HIM Coding Leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Coding Managers and HIM Coding leaders to evaluate AI-powered coding solutions and help train next-generation autonomous coding systems. Your expertise in professional fee (profee) and/or inpatient facility coding, ICD-10-CM/PCS, CPT/HCPCS, and HIM operations will directly inform AI tools designed to improve coding accuracy, productivity, and compliance across healthcare settings. — Responsibilities • Oversee professional fee and/or facility inpatient coding operations, ensuring accuracy, productivity, and compliance with coding guidelines. • Evaluate AI-generated coding assignments, including ICD-10-CM/PCS diagnoses, procedure codes, CPT/HCPCS codes, and DRG assignments, for accuracy and compliance. • Conduct coding quality audits and provide targeted feedback to coding staff and AI systems. • Monitor coding KPIs including coder productivity, accuracy rates, unbilled accounts, and claim denial rates attributable to coding errors. • Manage coding workflow queues, work distribution, and turnaround time compliance. • Ensure adherence to Official Coding Guidelines, CMS regulations, and payer-specific coding requirements. • Provide ongoing coding education and compliance training to coding staff. • Collaborate with CDI, billing, and compliance teams to address coding-related revenue integrity issues. • Annotate AI-generated coding outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in medical coding, with at least 2 years in a coding manager or HIM leadership role. • Expert knowledge of ICD-10-CM/PCS, CPT/HCPCS, and Official Coding Guidelines. • Proficiency in professional fee (profee) coding and/or facility inpatient coding with DRG assignment experience. • Experience conducting coding audits and developing coding quality improvement programs. • Proficiency with coding software (3M, Nuance, Optum360, TruCode) and EHR platforms. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify coding errors, compliance risks, and AI output inaccuracies. — Preferred Qualifications • CPC (Certified Professional Coder), CCS (Certified Coding Specialist), RHIA, or RHIT credential. • Experience with computer-assisted coding (CAC) tools and NLP-based coding platforms. • Background in inpatient facility coding with DRG optimisation experience. • Familiarity with AI coding tools and comfort evaluating AI-generated coding assignments. • Experience presenting coding performance data and quality metrics to leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in medical coding and health information management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$80 / hourOpen / Referral verified
Medical / Remote

Patient Access Leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Patient Access leaders to support the evaluation and improvement of AI tools designed for healthcare revenue cycle workflows. This is an opportunity to apply your expertise in patient registration, pre-registration, scheduling, and intake coordination to help shape AI systems that will transform front-end healthcare operations. — Responsibilities • Lead end-to-end patient access operations including pre-registration, registration, scheduling, and intake workflows across inpatient and outpatient settings. • Evaluate AI-generated outputs related to patient access processes, identifying errors and providing structured, actionable feedback. • Develop and enforce policies and procedures for registration accuracy, demographic capture, and point-of-service collections. • Monitor KPIs including registration error rates, pre-registration completion rates, insurance verification accuracy, and front-end denial rates. • Ensure clean claim submission from point of entry by identifying and resolving front-end deficiencies. • Collaborate with billing, clinical, and IT teams to streamline patient access workflows and reduce downstream claim errors. • Train and mentor staff on registration best practices, compliance requirements, and system usage. • Ensure compliance with HIPAA, CMS guidelines, and payer-specific requirements at the point of service. • Document observations, annotate AI outputs, and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in patient access, patient registration, or front-end revenue cycle management, with at least 2 years in a manager or director-level role. • Deep knowledge of pre-registration, insurance verification, scheduling intake, and point-of-service collections. • Strong understanding of HIPAA, CMS regulations, and commercial and government payer requirements. • Proficiency with Epic, Cerner, Meditech, or similar EHR and registration systems. • Exceptional written and verbal English communication skills. • High attention to detail and ability to identify inconsistencies in workflows, data, or AI-generated outputs. • Comfortable working independently in a fully remote environment. • Strong organisational skills with the ability to manage multiple priorities simultaneously. — Preferred Qualifications • NAHAM Certified Healthcare Access Manager (CHAM) or Certified Healthcare Access Associate (CHAA) credential. • Experience with denial prevention strategies at the point of registration. • Familiarity with AI tools such as ChatGPT, Claude, or similar systems. • Background in health system, hospital, or large physician group settings. • Experience developing SOPs and training materials for patient access teams. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in healthcare revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$80 / hourOpen / Referral verified
Finance / Remote

FP&A Expert

Role Overview • Mercor is seeking senior FP&A professionals to build evaluation tasks for AI systems operating in financial planning, forecasting, and business analysis contexts. • The workflows are calibrated to the forecasting complexity, decision stakes, and scope of major corporate financial planning functions. • This role builds worlds on two tracks: a US track (US GAAP-based reporting and corporate finance conventions) and an International track (IFRS-based reporting and international corporate finance conventions). Experts qualified in either or both tracks are encouraged to apply. • Contributors design FP&A scenarios, draft reference outputs, and write rubrics that capture how senior FP&A leaders think. — Key Responsibilities • Construct FP&A scenarios spanning budgeting and forecasting cycles, variance analysis and business partnering, and capital allocation and strategic planning decisions. • Build tasks across financial planning and forecasting, variance analysis and reporting, business partnering, capital budgeting and investment analysis, and financial modeling. • Develop scenarios involving tools such as FP&A platforms (Anaplan, Adaptive Insights, Planful), BI tools (Tableau, Power BI), and ERP systems used at large companies. • Apply FP&A methodologies (driver-based forecasting, variance analysis, scenario and sensitivity modeling) to the reporting basis a world targets (US GAAP-based; IFRS-based), and produce reference financial models, forecast decks, variance reports, and business case analyses. • Author rubrics that distinguish authentic FP&A judgment from generic textbook or spreadsheet-template-level work. — Ideal Qualifications • 5+ years working as an FP&A leader, Director of FP&A, or CFO at a major company. • Direct ownership of budgeting cycles, forecasting models, or capital allocation decisions. • Fluency in FP&A tooling, plus understanding of corporate finance and financial modeling best practices. • A recognized professional credential is a plus (CFA, MBA, or CPA); prior rubric or training authorship is a plus.

$80 - $90 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Swedish

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Swedish and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Swedish, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Swedish music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Swedish genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Swedish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$42 - $78 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Norwegian

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Norwegian and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Norwegian, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Norwegian music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Norwegian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Norwegian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$42 - $78 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Norwegian

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Norwegian and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Norwegian, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Norwegian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Norwegian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$42 - $78 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Spanish - Peru)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Peru-based female voice actors with native Peruvian Spanish accents to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native Peruvian female Spanish speaker currently based in Peru • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 - $100 / hourOpen / Referral verified
Code / Remote

Data Engineer (Coding Agent Experience)

About the Role — \- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. — \- Contributors help evaluate and improve frontier AI coding models through structured technical assessments. — \- The work focuses on realistic data engineering workflows and model evaluation. — \- Spots are limited and filling quickly on a first come, first serve basis. — What You'll Do — \- Use frontier AI coding agents to complete and evaluate complex data engineering tasks. — \- Review model-generated implementations involving ETL pipelines, data warehouses, analytics platforms, and distributed data systems. — \- Identify bugs, edge cases, scalability issues, and failure modes. — \- Compare outputs from multiple frontier models and assess their strengths and weaknesses. — \- Apply professional engineering judgment to realistic data engineering scenarios. — Time Commitment — \- Sprint based project that runs in 12-24 hour stretches based on client requirement. — Compensation — \- $400 per accepted task. — \- Typical tasks take approximately 2–3 hours after ramp-up. — \- Compensation is tied to accepted work. — Who Should Apply — \- 2+ years of professional data engineering experience. — \- Experience building ETL pipelines, data warehouses, analytics platforms, or distributed data systems. — \- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. — \- Ability to evaluate model-generated data infrastructure and pipeline implementations. — \- Experience operating large-scale data platforms is preferred.

$80 / hourOpen / Referral verified
Code / Remote

Applied Engineering Benchmark Specialist

Role Overview — We are seeking expert engineers to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core engineering domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of engineering expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Engineering Domains Covered — Semiconductor Design & Manufacturing (VLSI), Control Science and Engineering, Mechatronics, Reactor, Plant Design & Separations, Reservoir Engineering & Maintenance, Bioinstrumentation & Biotechnology. — Key Responsibilities • Author original engineering questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Engineering or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level engineering principles, applied mathematics, and domain-specific standards • Professional engineering licensure (PE) or industry experience is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$61 - $77 / hourOpen / Referral verified
Math / Remote

Applied Mathematics Benchmark Specialist

Role Overview — We are seeking expert mathematicians to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core mathematics domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of mathematical expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Mathematics Domains Covered — Signal Processing, Financial Mathematics & Actuarial Science, Mathematical Economics, Mathematical Modeling of Ecological & Biological Systems, Mathematical Programming & Combinatorial Optimization, Geomathematics & Climate Modeling. — Key Responsibilities • Author original math questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Mathematics, Applied Mathematics, Statistics, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level mathematical concepts and formal proof writing • Experience with rigorous academic problem design or mathematical competition writing is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$61 - $77 / hourOpen / Referral verified
STEM / Remote

Applied Chemistry Benchmark Specialist

Role Overview — We are seeking expert chemists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core chemistry domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of chemistry expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Chemistry Domains Covered — Materials, Polymer & Electronic Chemistry, Industrial & Process Chemistry, Energy Storage & Environmental Chemistry, Pharmaceutical & Agrochemical Chemistry, Consumer, Food & Specialty Chemicals. — Key Responsibilities • Author original chemistry questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Chemistry, Biochemistry, Chemical Engineering, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level chemistry concepts, reaction mechanisms, and quantitative analysis • Experience with rigorous academic problem design or chemistry olympiad writing is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$61 - $77 / hourOpen / Referral verified
STEM / Remote

Applied Physics Benchmark Specialist

Role Overview — We are seeking expert physicists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core physics domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of physics expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Physics Domains Covered — Semiconductor Physics, Nanoelectronics & Spintronics, Photonics, Quantum Optics & Ultrafast, Quantum Sensing & Metrology, Plasma Physics & Fusion Energy, Nonlinear Dynamics & Turbulence, Geophysics & Reservoir Simulation. — Key Responsibilities • Author original physics questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Physics, Applied Physics, Astrophysics, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level physics concepts and mathematical formalism • Experience with rigorous academic problem design or physics olympiad writing is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$61 - $77 / hourOpen / Referral verified
STEM / Remote

Chess Expert (1800+ ELO) — Game Reconstruction & Transcript Correction

We're hiring strong chess players to help build a high-quality chess dataset for an AI research partner: — 1. Transcript Correction — Watch a video of a strong player narrating a game and fix speech-to-text errors in the commentary transcript (piece names, squares, moves). The form syncs the transcript to the video (click a word to jump there). 2. Game Reconstruction (the main task) — Working ONLY from the corrected transcript (no video, no searching for the source game), rebuild the full game on an analysis board (lichess/chess.com) and paste the PGN. Then determine whether the transcript uniquely specifies the complete game; if not, itemize the minimal missing information needed. — You're a fit if you: • Are rated 1800+ ELO (FIDE or online equivalent) • Read algebraic notation fluently and are comfortable using an analysis board • Have a strong ear for spoken chess commentary and careful attention to detail

$90 / hourOpen / Referral verified
Code / Remote

Civil Engineering Expert

Role Overview • Mercor is seeking senior civil engineering professionals to build evaluation tasks for AI systems operating in large-scale infrastructure and construction contexts. • The workflows are calibrated to the design complexity, public safety stakes, and regulatory scope of major infrastructure projects and large public/private construction programs. • This role builds worlds on two standards tracks: an International track (Eurocodes and ISO) and a US track (ASCE, ACI, AISC, AASHTO). Experts qualified in either or both tracks are encouraged to apply. • Contributors design civil engineering scenarios, draft reference outputs, and write rubrics that capture how senior civil engineering leaders think. — Key Responsibilities • Construct civil engineering scenarios spanning large-scale infrastructure design cycles, multi-stakeholder permitting and regulatory review, and complex construction management or public works processes. • Build tasks across structural design and analysis, transportation and highway engineering, water resources and environmental engineering, geotechnical engineering, and construction project management. • Develop engineering scenarios involving tools such as AutoCAD Civil 3D, Revit, STAAD.Pro/ETABS, HEC-RAS, and enterprise project management platforms used on major infrastructure programs. • Apply civil engineering methodologies (structural load analysis, geotechnical site assessment, hydrology/hydraulic modeling) to the standards track a world targets (US: ASCE 7, ACI 318, AISC 360, AASHTO LRFD; International: Eurocodes EN 1990-1998 and ISO), and produce reference design plans, engineering calculations, and stakeholder/regulatory-facing technical narratives. • Author rubrics that distinguish authentic civil engineering judgment from generic textbook or coursework-level recall. — Ideal Qualifications • 5+ years working as a civil engineer or engineering lead • Direct ownership of large-scale infrastructure designs, permitting processes, or construction project delivery. • Fluency in civil engineering tooling and methodologies, plus understanding of how regulatory approvals, environmental review (NEPA or an international equivalent), and public agency oversight actually work. • A recognized professional engineering credential is strongly preferred (US PE, or an international equivalent such as CEng, EUR ING, or P.Eng); prior rubric, engineering-training curriculum, or design documentation authorship is a plus.

$70 - $80 / hourOpen / Referral verified
STEM / Remote

Applied Biology Benchmark Specialist

Role Overview — We are seeking expert biologists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core biology domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of biology expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Biology Domains Covered — Pharmaceutical Manufacturing, Industrial and Synthetic Biology, Medical Research & Drug Discovery, Agricultural, Environmental & Food Biology. — Key Responsibilities • Author original biology questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Biology, Molecular Biology, Biochemistry, Neuroscience, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level biological concepts, experimental design, and data interpretation • Research publications or laboratory experience in biological sciences is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$60 - $75 / hourOpen / Referral verified
Finance / Remote

A/R Follow-up Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced A/R Follow-Up Managers to evaluate AI tools designed to automate accounts receivable follow-up and payer collections workflows. Your expertise in claim status follow-up, payer correspondence, and A/R management will directly inform AI systems that reduce days in accounts receivable and improve revenue recovery. — Responsibilities • Lead A/R follow-up operations across commercial, Medicare, Medicaid, and managed care payers, ensuring timely resolution of outstanding claims. • Evaluate AI-generated A/R follow-up recommendations, claim status inquiry outputs, and payer correspondence drafts for accuracy and effectiveness. • Manage claim status follow-up workflows including electronic claim status inquiries (276/277 EDI), payer portal follow-up, and phone-based resolution. • Prioritise A/R queues by ageing bucket, payer, and dollar value to maximise revenue recovery. • Identify and resolve claim payment discrepancies, payer processing errors, and underpayments. • Monitor A/R KPIs including days in A/R, ageing bucket distribution, collection rates, and write-off rates. • Develop and implement payer-specific follow-up strategies to accelerate claim resolution. • Ensure compliance with FDCPA, HIPAA, and payer-specific follow-up and timely filing requirements. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in A/R follow-up, payer collections, or revenue cycle operations, with at least 2 years in a management role. • Deep knowledge of claim status follow-up workflows, EDI 276/277 transactions, and payer-specific collections processes. • Strong understanding of Medicare, Medicaid, and commercial payer claims processing and payment timelines. • Experience prioritising and managing high-volume A/R queues across multiple payers. • Proficiency with billing systems and A/R management platforms. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify payment errors and discrepancies in AI-generated A/R content. — Preferred Qualifications • CRCR, CPC, or CHFP certification. • Experience with RCM technology platforms featuring automated A/R follow-up capabilities. • Background in multi-payer follow-up operations in hospital or physician group settings. • Familiarity with AI tools and comfort evaluating AI-generated A/R follow-up content. • Experience developing A/R reduction action plans and presenting performance to leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in accounts receivable and revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$75 / hourOpen / Referral verified
Medical / Remote

Pharmacy Prior Authorization & Specialty-Medication Access Specialist

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Medication and Pharmacy Prior Authorisation specialists with expertise in speciality medication access to evaluate AI tools designed to automate pharmacy benefit management and speciality drug authorisation workflows. Your deep knowledge of drug formularies, step therapy, and payer PA criteria will directly shape AI systems that improve patient access to critical therapies. — Responsibilities • Process and manage prior authorisation requests for speciality medications, biologics, and high-cost pharmaceuticals across commercial, Medicare Part D, and Medicaid payers. • Review clinical documentation and pharmacy benefit criteria to determine medical necessity for speciality drug authorisations. • Evaluate AI-generated pharmacy prior authorisation recommendations for accuracy, completeness, and payer compliance. • Coordinate with prescribers, speciality pharmacies, and payers to obtain timely medication authorisations. • Manage appeals for denied pharmacy authorisations, including peer-to-peer requests and exception processes. • Navigate payer formularies, step therapy requirements, and speciality drug coverage policies. • Track authorisation status, approval rates, and denial trends to identify process improvement opportunities. • Ensure compliance with payer-specific pharmacy benefit requirements and CMS Part D guidelines. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in pharmacy prior authorisation, speciality medication access, or pharmacy benefit management. • Deep knowledge of speciality drug formularies, step therapy protocols, and payer PA criteria across commercial and government payers. • Experience with speciality pharmacy platforms and pharmacy benefit management (PBM) systems. • Familiarity with Medicare Part D coverage determination and exception processes. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to evaluate clinical documentation and AI-generated outputs. — Preferred Qualifications • Certified Pharmacy Technician (CPhT) or Certified Prior Authorisation Professional (CPAP) credential. • Experience with hub services and patient assistance programs for speciality medications. • Background in speciality pharmacy, infusion services, or oncology/rheumatology drug access. • Familiarity with AI tools and comfort evaluating AI-generated pharmacy content. • Experience with manufacturer copay assistance and patient access programs. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in healthcare revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$75 / hourOpen / Referral verified
Business / Remote

HR Expert

Role Overview • Mercor is seeking senior HR professionals to build evaluation tasks for AI systems operating in talent management, employee relations, and organizational design contexts. • The workflows are calibrated to the organizational scale, compliance stakes, and scope of major workforce programs at large companies. • This role builds worlds on two tracks: a US track (Title VII, FLSA, ADA, and state employment law) and an International track (UK Employment Rights Act, EU labor directives). Experts qualified in either or both tracks are encouraged to apply. • Contributors design HR scenarios, draft reference outputs, and write rubrics that capture how senior HR leaders think. — Key Responsibilities • Construct HR scenarios spanning talent acquisition and workforce planning, employee relations and performance management, and compensation and benefits design. • Build tasks across recruiting and talent acquisition, employee and labor relations, compensation and benefits, HR compliance and employment law, and organizational development. • Develop scenarios involving tools such as HRIS platforms (Workday, SAP SuccessFactors), applicant tracking systems, and compensation benchmarking tools (Radford, Mercer) used at large organizations. • Apply HR methodologies (workforce planning, compensation benchmarking, employee relations investigation) to the standards track a world targets (US: Title VII, FLSA, ADA, state law; International: Employment Rights Act, EU labor directives), and produce reference HR policies, investigation reports, compensation analyses, and organizational design proposals. • Author rubrics that distinguish authentic HR judgment from generic textbook or certification-exam-level recall. — Ideal Qualifications • 5+ years working as an HR leader or Chief People Officer at a major company or HR consulting firm (Mercer, Korn Ferry, or in-house VP/CHRO at a large company). • Direct ownership of talent programs, employee relations matters, or compensation design. • Fluency in HR tooling and methodologies, plus understanding of how employment law compliance actually works. • A recognized professional credential is strongly preferred (SHRM-SCP, SPHR, or an international equivalent); prior rubric or training authorship is a plus.

$70 - $80 / hourOpen / Referral verified
Language / Remote

Design Expert

Role Overview • Mercor is seeking senior product design and AI-assisted development professionals to build evaluation tasks for AI systems operating in rapid prototyping, product design, and human-AI collaborative coding contexts. • The workflows are calibrated to the design complexity, user experience stakes, and scope of major product development and design systems work. • This role builds worlds across two subdomains: a Vibecoding track (AI-assisted, natural-language-driven software development and rapid prototyping) and a Design track (product, UX/UI, and visual design systems). Experts qualified in either or both subdomains are encouraged to apply. • Contributors design scenarios, draft reference outputs, and write rubrics that capture how senior product builders and designers think. — Key Responsibilities • Construct scenarios spanning rapid prototyping and iteration cycles, design system development, and cross-functional product-design-engineering collaboration. • Build tasks across AI-assisted application development, natural-language-to-code workflows, and rapid prototyping (Vibecoding); and UX/UI design, design systems and component libraries, user research, and usability testing (Design). • Develop scenarios involving tools such as AI coding assistants (Cursor, Claude Code, v0, Replit) for the Vibecoding subdomain, and Figma, prototyping tools (Framer), and design systems tooling for the Design subdomain. • Apply subdomain-specific methodologies (iterative prompt-driven development for Vibecoding; human-centered design principles and heuristic evaluation for Design), and produce reference prototypes, code artifacts, design specs, and usability evaluation reports. • Author rubrics that distinguish authentic senior-level product judgment from generic tutorial-level or template-driven work. — Ideal Qualifications • 5+ years working as a product designer, design lead, or AI-native engineer/vibecoder at a major product- or design-forward company (Figma, Airbnb, Notion, Linear, or a notable independent/startup product builder). • Direct ownership of product design systems or AI-assisted development workflows shipped to production. • Fluency in vibecoding or design tooling, plus understanding of user-centered design principles and rapid prototyping practices. • A portfolio of shipped work is strongly preferred in place of a formal credential; prior rubric, training, or design critique authorship is a plus.

$70 - $80 / hourOpen / Referral verified
Finance / Remote

Accounting Expert

Role Overview • Mercor is seeking senior accounting professionals to build evaluation tasks for AI systems operating in financial reporting, audit, and technical accounting contexts. • The workflows are calibrated to the reporting complexity, materiality stakes, and scope of major public company and enterprise accounting functions. • This role builds worlds on two tracks: a US track (US GAAP, PCAOB standards) and an International track (IFRS, International Standards on Auditing). Experts qualified in either or both tracks are encouraged to apply. • Contributors design accounting scenarios, draft reference outputs, and write rubrics that capture how senior accountants and auditors think. — Key Responsibilities • Construct accounting scenarios spanning financial statement preparation, technical accounting research, external and internal audit processes, and complex transaction accounting (M&A, revenue recognition, leases). • Build tasks across financial reporting, technical accounting and research, audit and assurance, complex transaction accounting, and internal controls/SOX compliance. • Develop scenarios involving tools such as ERP systems (SAP, Oracle), consolidation software, audit management platforms, and research tools (Bloomberg Tax, RIA Checkpoint) used at major firms. • Apply accounting methodologies (revenue recognition, lease accounting, consolidation) to the standards track a world targets (US: US GAAP/ASC codification, PCAOB standards; International: IFRS, ISA), and produce reference financial statements, technical accounting memos, audit workpapers, and internal control assessments. • Author rubrics that distinguish authentic accounting judgment from generic textbook or CPA exam-level recall. — Ideal Qualifications • 5+ years working as an accountant, controller, or audit partner at a major accounting firm or corporation (Big Four, or a corporate controller/CFO). • Direct ownership of financial reporting, audit engagements, or technical accounting matters. • Fluency in accounting tooling, plus understanding of regulatory reporting (SEC filings) and audit standards. • A recognized professional credential is strongly preferred (CPA, or an international equivalent such as ACCA or CA); prior rubric or training authorship is a plus.

$70 - $80 / hourOpen / Referral verified
Legal / US Remote

Legal Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced legal professionals to help evaluate and improve how AI systems understand and reason about legal topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in legal contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of legal expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how legal professionals actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Admitted to a United States state bar. A background as a legal librarian is a plus, though not required. • 2+ years of professional legal experience, ideally spanning more than one practice area (e.g. contracts, litigation, regulatory/compliance, corporate, IP). • Comfortable researching and writing using standard legal tools and references (case law databases, statutes, treatises). • Some experience writing or publishing legal content — memos, briefs, CLE materials, or similar — is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Business / Remote

Wealth Management & Asset Management Expert

Role Overview • Mercor is seeking senior wealth management and asset management professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise investment and advisory contexts. • The workflows are calibrated to the portfolio complexity, regulatory stakes, and client sophistication of Fortune 500 asset managers and large private wealth institutions. • Contributors design enterprise wealth and asset management scenarios, draft reference outputs, and write rubrics that capture how senior F500 investment professionals think. — Key Responsibilities • Construct enterprise wealth management scenarios spanning high-net-worth portfolio construction, multi-stakeholder trust and estate planning, and complex fiduciary or regulatory review cycles. • Build asset management tasks across F500 portfolio strategy, alternative investments, institutional client servicing, risk and performance analytics, and manager due diligence. • Develop investment operations scenarios involving tools such as Bloomberg Terminal, Aladdin, Morningstar Direct, and enterprise portfolio/order management systems (OMS) in F500 stacks. • Apply enterprise investment methodologies (modern portfolio theory, factor investing, fiduciary duty standards, ESG integration) and produce reference investment policy statements, portfolio strategies, and client-facing executive narratives. • Author rubrics that distinguish authentic enterprise wealth and asset management judgment from generic textbook or CFA-exam-level recall. — Ideal Qualifications • 5+ years working as a portfolio manager, wealth advisor, or investment professional at a Fortune 500 asset manager, private bank, or wirehouse (BlackRock, Vanguard, Morgan Stanley, Goldman Sachs, JPMorgan). • Direct ownership of F500-scale client portfolios, institutional mandates, or investment strategy decisions. • Fluency in enterprise investment tooling and methodologies, plus understanding of how F500 regulatory compliance (SEC, FINRA), fiduciary standards, and client reporting actually work. • Prior rubric, investment-training curriculum, or portfolio documentation authorship is a plus.

$70 - $80 / hourOpen / Referral verified
Finance / US Remote

Finance Specialist — CFA/ACA/ACCA/CPA Required

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking Finance subject-matter experts (SMEs) with solid domain expertise to bring rigor and real-world financial judgment to our AI training data. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps in financial reasoning, analysis, and decision-making. • Design challenging, domain-relevant finance tasks and write accurate, well-reasoned solutions grounded in real financial practice. • Evaluate AI model outputs against structured rubrics and provide clear, written feedback on correctness, judgment, and reasoning quality. • Develop and refine evaluation guidelines and scoring rubrics specific to finance tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 2+ years of dedicated professional experience in finance (e.g., investment banking, asset management, corporate finance, accounting/audit, corporate treasury, financial planning & analysis) — not a generalist role that only touches finance peripherally. • Professional finance credential required: CFA, ACA, ACCA, or CPA. • Ability to engage reliably for at least 10 hours/week during weekdays. • Verbal and written communication skills, problem-solving skills, and interpersonal skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Business / Remote

Denials Management & Appeals Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Denials Management and Appeals Managers to evaluate AI tools designed to automate denial prevention, appeal writing, and root cause analysis workflows. Your expertise in payer denial patterns, clinical and technical appeals, and denial management analytics will directly inform AI systems that reduce denial rates and maximise revenue recovery. — Responsibilities • Lead denials management and appeals operations, overseeing the identification, categorisation, and resolution of claim denials. • Evaluate AI-generated appeal letters, denial root cause analyses, and denial prevention recommendations for accuracy and effectiveness. • Analyse denial trends by payer, denial code (CARC/RARC), and denial category to identify systemic root causes. • Develop and manage clinical and technical appeal strategies across multiple payer types. • Coordinate with clinical, coding, billing, and compliance teams to implement denial prevention initiatives. • Monitor denial management KPIs including denial rates, appeal overturn rates, revenue recovery, and days in A/R. • Manage the appeals calendar to ensure timely submission within payer and regulatory deadlines. • Ensure compliance with payer appeal requirements, CMS regulations, and timely filing deadlines. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in denials management, appeals, or revenue cycle operations, with at least 2 years in a management role. • Deep knowledge of CARC/RARC denial codes, payer denial patterns, and appeal strategies across commercial, Medicare, and Medicaid payers. • Strong understanding of clinical and technical appeal processes including peer-to-peer reviews and external reviews. • Experience with denial analytics platforms and revenue cycle reporting tools. • Proficiency with EHR systems and billing platforms. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to evaluate appeal quality and identify errors in AI-generated denial management content. — Preferred Qualifications • CPC, CCS, CRCR, or CHFP certification. • Experience with AI-assisted denial management platforms (e.g., Waystar, Experian Health, Nthrive). • Background in complex clinical appeals including medical necessity, experimental/investigational, and level of care denials. • Familiarity with AI tools and comfort evaluating AI-generated appeal and denial content. • Experience presenting denial management performance to revenue cycle leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in denials management and revenue cycle. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$70 - $93 / hourOpen / Referral verified
Language / Remote

Mechanical Engineering Writer - Engineering FRQ

We're looking for a senior mechanical engineering expert to help build a benchmark of the hardest reasoning questions in the field — questions specifically designed to be difficult enough that today's most capable AI models still get them wrong. You'll author original, free-response engineering problems grounded in real industry scenarios, write a complete expert-level solution for each, and validate difficulty by running it against three frontier language models until at least one fails. — Requirements: PhD in Mechanical Engineering, plus 8+ years of hands-on professional/industry experience in design, analysis, or R&D. Candidates from well-known/blue-chip employers preferred.

$85 / hourOpen / Referral verified
Language / Remote

Generalist Expert (UK/Europe)

We’re looking for UK or Europe-based generalist experts to help train and evaluate frontier AI models. In this role, you’ll review and assess a wide range of everyday professional content — documents, slides, spreadsheets, and other written materials — judging them for quality, accuracy, clarity, and completeness. Your feedback directly shapes how AI systems reason about and produce real-world work product. — This is a great fit if you’re a sharp, detail-oriented generalist who’s comfortable moving across different formats and subject areas. You don’t need deep specialization in any one field — strong judgment, careful reading, and clear written reasoning matter most. — What you’ll do • Review and evaluate documents, slides, spreadsheets, and similar materials for quality and correctness • Provide clear, well-reasoned written feedback and ratings • Compare and rank AI-generated outputs against defined criteria • Flag errors, inconsistencies, and gaps in reasoning or formatting — What we’re looking for • Bachelor’s degree (minimum requirement) • Based in the UK or Europe • Excellent reading comprehension and written communication in English • Strong attention to detail and sound judgment across varied subject matter • Comfort working independently across common productivity tools (docs, slides, sheets) — No prior AI or machine-learning experience is required — we’ll provide the guidelines and context you need to succeed.

$50 - $70 / hourOpen / Referral verified
Code / Remote

Network Engineer - Data for Autonomous Systems annotation

Are you a Level 3 / Tier 3 network support engineer interested in data science and autonomous infrastructure? Our client is building vertically integrated networking systems and using the data they generate to power the next generation of AI-driven infrastructure. They're looking for engineers experienced in final escalations, packet analysis, troubleshooting, and RCA workflows to help label, annotate, and structure networking data from real production systems. — This is a hands-on role that blends your L3 troubleshooting and incident-response experience with a growing understanding of how data pipelines are built and used in AI systems. — In this role, you'll: • Review real-world data from deployed networks: logs, configs, telemetry, event streams • Label and classify network behaviors, issues, anomalies, and incident patterns • Help define schemas and structure for large-scale data pipelines that downstream ML models will train on — You're a strong fit if you: • Work today as a Level 3 / Tier 3 / Principal Support Engineer keeping existing enterprise infrastructure online and stable — on-call rotation, RCAs, final escalations, troubleshooting outages — Must Have • Have hands-on experience with end-customer enterprise networks (switches, APs, firewalls in retail, healthcare, financial, manufacturing, university, hospitality, etc.) — Must Have • Bring hands-on Wi-Fi/wireless proficiency — enterprise WLAN controllers (Cisco WLC, Aruba, or Meraki), 802.1X/RADIUS, and wireless troubleshooting — Must Have • Do packet-level troubleshooting yourself — Wireshark, tcpdump, SPAN captures • Are curious about how raw infra data becomes machine learning input — This is a maintainer role — likely not the right fit if your current work is mainly network design/architecture, cloud/SRE, security/SOC, or IT helpdesk. — Your work will directly feed into the pipelines that power client's AI models, and help shape how intelligent systems reason about networks in the real world. — Here are more details about the role: • You will interface directly with the client team. • You are expected to work 30-40 hours/week, with your hours overlapping the Pacific (PT) business day. • This is an individual 1099 contract paid to a personal account — no corp-to-corp or agency billing. • You must be authorized to work in the US or Canada without sponsorship.

$50 - $70 / hourOpen / Referral verified
Business / United States Remote

Project Coordinator - AI & Data Projects

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — Are you ready to help shape the future of artificial intelligence? Join a leading AI lab's cutting-edge GenAI team, where you'll be at the forefront of building groundbreaking AI models. We're seeking talented Project Coordinators to support and accelerate world-class AI research and data operations — acting as the operational bridge between program leadership and a growing team of domain experts. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. What You'll Do • Act as a day-to-day point of contact for AI data projects, helping keep workflows and operations running smoothly. • Work closely with program leads and domain experts — take inputs from leadership, turn them into clear guidelines, and share them with expert teams. • Help onboard and support new experts as the program grows. • Join weekly business reviews, keep notes and action items organized, and lead calls when needed. • Help spot quality issues, understand root causes, and share findings across teams. • Contribute to creating and refining project guidelines with research and product partners. • Flag potential roadblocks early and help get them resolved. • Track project deliverables using Google Sheets and Excel. • Stay flexible and adapt as project needs evolve. — 3. Qualifications • Location: Must be based in the USA. • Education: STEM background strongly preferred. • Experience: 3+ years in project coordination or project management, with the ability to coordinate large, cross-functional teams. • AI fluency: hands-on experience on AI training-data or human-data projects — as a project coordinator, team lead, EPM, or expert contributor. • Ideal: prior project coordination on AI/data programs at organizations like Meta, TikTok, or Amazon, or EPM/project-lead experience on Mercor projects or other AI tranining projects. • Skills: strong data management experience (Excel, SQL) and working knowledge of coding. • Leadership: prior people management experience or demonstrated ability to lead large groups effectively. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$45 - $70 / hourOpen / Referral verified
Policy & Safety / Remote

AI Safety Practitioner

We are seeking experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback. — Responsibilities • Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality. • Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains. • Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking. • Identify unsafe outputs, hallucinations, reasoning failures, and policy violations. • Provide structured feedback to improve model alignment and safety performance. • Collaborate with AI researchers and safety teams on ongoing evaluation initiatives. — Required Qualifications • Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline. • 5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field. • Excellent written English, critical thinking, and analytical reasoning skills. • Ability to consistently evaluate nuanced and policy-sensitive scenarios. — Preferred Qualifications • Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation. • Familiarity with safety policies, content moderation, or evaluation rubric development. • Experience reviewing complex, high-risk, or ambiguous content. — Why Join? • Shape the safety and behaviour of frontier AI models used by millions worldwide. • Work on challenging, real-world safety evaluations across nuanced and high-impact domains. • Collaborate with leading AI researchers, engineers, and safety teams.

$60 - $70 / hourOpen / Referral verified
Business / Remote

IT Services & Consulting Expert

Role Overview • Mercor is seeking senior IT services and consulting professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise consulting and technology delivery contexts. • The workflows are calibrated to the engagement scale, stakeholder complexity, and delivery stakes of Fortune 500 and large public company consulting engagements. • Contributors design enterprise consulting scenarios, draft reference outputs, and write rubrics that capture how senior F500 consulting and delivery leaders think. — Key Responsibilities • Construct enterprise consulting scenarios spanning $1M+ engagement values, multi-stakeholder steering committees, and complex statement-of-work (SOW) and procurement cycles at F500 accounts. • Build consulting tasks across F500 digital transformation strategy, systems integration, managed services delivery, IT advisory, and technology change management. • Develop delivery and engagement scenarios involving tools such as Jira/Confluence, ServiceNow, SAP, Salesforce, and enterprise PMO platforms in F500 client environments. • Apply enterprise consulting methodologies (Agile/SAFe delivery, business case development, RACI/governance models) and produce reference engagement plans, account strategies, and executive-level deliverables. • Author rubrics that distinguish authentic enterprise consulting judgment from generic framework recall. — Ideal Qualifications • 5+ years in IT consulting, systems integration, or managed services delivery at a Fortune 500 technology or consulting firm (Accenture, Deloitte, IBM, Capgemini, Cognizant, TCS) or inside an F500 enterprise IT/transformation organization (JPMorgan, UPS, Unilever, Microsoft, PepsiCo). • Direct ownership of F500 client engagements, F500 delivery programs, or F500 transformation initiatives. • Fluency in enterprise consulting methodologies and delivery tooling, plus understanding of how F500 budgets, procurement, and legal/SOW review actually work. • Prior rubric, consulting-enablement curriculum, or training-content authorship is a plus.

$70 - $80 / hourOpen / Referral verified
Language / Remote

Generalist (Macbook User)

Role Overview: — Mercor is looking for detail-oriented individuals with 2-3 years of experience in STEM, non-STEM fields, or currently enrolled postgrad college students to support a research project with a leading AI lab. You will help benchmark and improve cutting-edge AI models. — Qualifications: • Required: All experts on this project must have a Macbook device with an ‘M’ series chip to perform tasks on. • Currently pursuing or recently completed Masters studies, or have 2-3 years of relevant experience + Bachelors from a prestigious institution • Strong online research and communication skills • Ability to gather and clearly summarize information from diverse sources • Excellent written communication skills • Interdisciplinary degree is a nice to have • Must have a Macbook with one of the following requirements: • Apple Silicon (ARM) Macs • M‑series MacBook Pro, or • MacBook Air models running macOS 15 or higher — Job Details: • Part-time commitment of approximately 10-20 hours per week • Responsibilities include creating high-quality research questions, answers, and evaluation materials to train advanced language models • Clearly structured responsibilities without the high pressure typically found in internships or full-time consulting roles • Immediate start preferred — Application and Onboarding Process: • Submit your resume, followed by a brief (15-minute) conversation with our AI interviewer to assess research and reasoning skills • Complete a brief paid assessment to further evaluate your fit for the role • You will receive follow-up communication within a few days regarding your application status and next steps

$50 - $60 / hourOpen / Referral verified
Code / Remote — US-based

Atomistic & Surface Modeling Experts (Computational Materials & Catalysis)

Mercor is seeking computational scientists specializing in atomistic and surface modeling to support a frontier AI research lab building models for materials science and the physical sciences. This is hands-on, expert-level work: you'll apply deep, specialized knowledge to generate, structure, and evaluate the scientific data these models learn from — and your input will directly shape how advanced models reason about materials, surfaces, and chemical processes. — Key Responsibilities: • Contribute domain expertise across first-principles and molecular simulation — electronic structure, surface and interface modeling, adsorption, and reaction energetics — to build high-quality training and evaluation data. • Review and evaluate AI-generated scientific reasoning, catching errors and improving technical accuracy. • Design and solve challenging, expert-level problems in atomistic and surface modeling. • Rate and rank model outputs against defined scientific criteria, with clear written reasoning. • Structure technical knowledge — simulation setups, methods, and results — into well-organized, model-ready data. • Deliver reliable, high-quality work within defined timelines. — You're a strong fit if you have: • Hands-on experience with atomistic modeling using first-principles or molecular methods (DFT, ab initio molecular dynamics, classical MD, or Monte Carlo). • Experience modeling surfaces, interfaces, and adsorption or reaction phenomena (slab models, surface reconstructions, transition states, NEB, microkinetics). • Experience modeling semiconductor-relevant materials, or a background in computational (heterogeneous) catalysis. • Proficiency with standard tooling (e.g., VASP, Quantum ESPRESSO, CP2K, GPAW, LAMMPS, ASE, pymatgen). • A PhD in materials science, chemistry, physics, chemical engineering, or a related field, ideally with several years of research experience beyond the PhD. • Clear written English and the ability to explain technical reasoning concisely. — Role Details: • Type: Long-term, ongoing engagement • Engagement: Up to 40 hours/week (minimum 10) • Work arrangement: Remote (US-based)

$84 / hourOpen / Referral verified
Code / United States Remote

Software Engineer, Full Stack (Python, Java, Rust, C#, C++)

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Full-Stack Software Engineers with hands-on production experience across modern application stacks — Python, Java, Rust, C#, C++, and TypeScript — to build, integrate, and stress-test real software on top of frontier models before they ship. Engineers work directly with the AI lab's engineering manager on fast-moving projects and are expected to onboard quickly and contribute working code early. This is a full-time commitment of 40 hours per week. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Build and ship full-stack applications, services, and internal tools — in Python, Java, Rust, C#, C++, or TypeScript — that exercise frontier model capabilities end to end. • Integrate pre-release model APIs into working software, including the surrounding scaffolding: tool interfaces, evaluation harnesses, and telemetry. • Diagnose and document model and integration failure modes surfaced while building, translating engineering observations into clear written feedback for the research team. • Prototype quickly against shifting requirements, delivering working increments with limited up-front specification. • Collaborate with the engineering manager and other engineers to maintain consistency in code quality, architecture, and technical documentation. — 3. Core Qualifications • 3+ years of dedicated professional experience building and shipping production software at a recognized, top-tier organization. • Production experience in at least one of Python, Java, Rust, C#, or C++, plus working proficiency in a second language — this cohort is intentionally staffed across multiple language ecosystems. • Hands-on experience across the stack: backend services and APIs, a modern front-end framework (React or equivalent), relational or document databases, and cloud deployment. • Demonstrated ability to onboard onto an unfamiliar codebase and deliver working code quickly with minimal ramp-up. • Demonstrable career progression. • Ability to engage reliably for at least 40 hours/week during weekdays. • Strong written communication skills and the ability to explain complex technical decisions clearly. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$50 - $65 / hourOpen / Referral verified
Language / United States Remote

AI Rater Guidelines Writer (Linguist / Instructional Designer)

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Linguists and Instructional Designers with deep experience translating ambiguous program requirements into clear, unambiguous rater guidelines to bring precision and consistency to our AI training data — working across domains from finance and retail to insurance, legal, and sports. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and program teams to close the gap between ambiguous program requirements and rater-ready instructions across a range of subject-matter domains. • Design clear, non-contradictory rating guidelines and rubrics that raters can apply consistently, including under edge cases. • Evaluate draft guidelines for ambiguity, internal contradiction, and coverage gaps, and revise until raters can apply them without escalation. • Translate program specifications from domains such as finance, retail, insurance, legal, and sports into precise, discipline-specific rater instructions. • Collaborate with other subject matter experts and program leads to ensure consistency and accuracy across guideline sets. — 3. Core Qualifications • 3+ years of professional experience in linguistics, instructional design, technical writing, or a closely related field, with direct experience writing or refining guidelines/rubrics for human raters in a GenAI/RLHF context. • Demonstrated ability to work across multiple subject-matter domains (e.g., finance, retail, insurance, legal, sports) and translate domain-specific nuance into clear, unambiguous instructions. • Strong track record of resolving ambiguity and contradiction in written specifications — able to point to concrete before/after examples. • Demonstrable career progression. • Ability to engage reliably for at least 35 hours/week during weekdays. • Strong written communication skills and the ability to explain complex or nuanced guidance clearly and precisely. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$45 - $65 / hourOpen / Referral verified
Code / Remote

Inorganic Materials, Semiconductor & Superconductor Experts

Mercor is seeking experimental scientists and engineers across inorganic synthesis, characterization, superconductors, and semiconductors (including advanced packaging) to support a frontier AI research lab building models for materials science and the physical sciences. This is hands-on, expert-level work: you'll apply deep, specialized knowledge to generate, structure, and evaluate the scientific data these models learn from — and your input will directly shape how advanced models reason about materials, devices, and processes. — Key Responsibilities: • Contribute domain expertise across synthesis, characterization, fabrication, and device physics to build high-quality training and evaluation data. • Review and evaluate AI-generated scientific reasoning, catching errors and improving technical accuracy. • Design and solve challenging, expert-level problems in your area of specialization. • Rate and rank model outputs against defined scientific criteria, with clear written reasoning. • Structure technical knowledge — experimental procedures, characterization results, process data — into well-organized, model-ready data. • Deliver reliable, high-quality work within defined timelines. — You're a strong fit if you have: • Hands-on experimental experience in one or more of: inorganic synthesis (solid-state, solution, solvothermal, sol-gel), superconducting materials, or semiconductors and advanced packaging. • Strong materials or device characterization skills (XRD, SEM, TEM, spectroscopy, electrical/transport measurements). • Experience with thin-film growth or device fabrication (MBE/epitaxy, MOCVD, CVD, sputtering, IBAD, etch, clean-room microfabrication) — a plus. • An advanced degree (PhD/MS) or equivalent hands-on experience in materials science, chemistry, physics, or a related engineering field. • Clear written English and the ability to explain technical reasoning concisely. — Role Details: • Type: Long-term, ongoing engagement • Engagement: Up to 40 hours/week (minimum 10) • Work arrangement: Remote (US-based)

$84 / hourOpen / Referral verified
Medical / Remote

Applied Psychology Benchmark Specialist

Role Overview — We are seeking expert psychologists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core psychology domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of psychology expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Psychology Domains Covered — Neuromorphic Engineering, Human Factors and Engineering Psychology, Consumer and Market Psychology, Psychometrics, Digital Health, Psychedelic-assisted Therapy. — Key Responsibilities • Author original psychology questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise reasoning in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD, PsyD, or doctoral candidate in Psychology or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level psychological theory, research methodology, and empirical literature • Clinical licensure or research publications in psychology is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$50 - $63 / hourOpen / Referral verified
STEM / Remote

Applied Philosophy Benchmark Specialist

Role Overview — We are seeking expert philosophers to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core philosophy domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of philosophy expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Philosophy Domains Covered — Formal Ontology & Knowledge Representation, AI Ethics, Applied Epistemology, Philosophy of Technology & Robotics, Philosophy of Science. — Key Responsibilities • Author original philosophy questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise reasoning in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Philosophy or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of philosophical argumentation, formal logic, and canonical texts across traditions • Research publications or teaching experience in philosophy is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$50 - $63 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Korean

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Korean and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Korean, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Korean music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Korean genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Korean • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$35 - $62 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Japanese

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Japanese and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Japanese, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Japanese music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Japanese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Japanese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$35 - $62 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Korean

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Korean and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Korean, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Korean genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Korean • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$35 - $62 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Norwegian

Location: Remote — Fluent Language Skills Required: English & Norwegian. Native fluency in English and Norwegian is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$48 - $62 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Danish

Location: Remote — Fluent Language Skills Required: English & Danish. Native fluency in English and Danish is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$48 - $62 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Finnish

Location: Remote — Fluent Language Skills Required: English & Finnish. Native fluency in English and Finnish is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$48 - $62 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Swedish

Location: Remote — Fluent Language Skills Required: English & Swedish. Native fluency in English and Swedish is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$48 - $62 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Dutch

Location: Remote — Fluent Language Skills Required: English & Dutch. Native fluency in English and Dutch is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$48 - $62 / hourOpen / Referral verified
Finance / Remote

Fraud Detection Experts

Role Overview • Mercor is seeking senior fraud detection professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise fraud and financial crimes contexts. • The workflows are calibrated to the transaction scale, adversarial sophistication, and financial-loss stakes of Fortune 500 banks, payment companies, and large enterprises. • Contributors design enterprise fraud detection scenarios, draft reference outputs, and write rubrics that capture how senior F500 fraud and financial crimes operators think. — Key Responsibilities • Construct enterprise fraud scenarios spanning large-scale transaction monitoring, multi-stakeholder investigation workflows, and complex regulatory reporting cycles (SARs, CTRs) at F500 institutions. • Build tasks across F500 payment fraud, account takeover and identity theft, anti-money laundering (AML), synthetic identity fraud, and merchant/card network fraud. • Develop fraud operations scenarios involving tools such as SAS Fraud Management, Actimize, Feedzai, and enterprise case management/transaction monitoring platforms in F500 stacks. • Apply enterprise fraud detection methodologies (rules-based and ML-driven detection models, network/link analysis, KYC/AML frameworks, BSA compliance) and produce reference investigation reports, risk models, and executive-level fraud narratives. • Author rubrics that distinguish authentic enterprise fraud investigation judgment from generic textbook or certification-exam-level recall. — Ideal Qualifications • 5+ years working in fraud detection, financial crimes investigation, or AML compliance at a Fortune 500 bank, payments company, or financial institution (JPMorgan Chase, Visa, PayPal, American Express, Wells Fargo). • Direct ownership of F500-scale fraud detection programs, investigation caseloads, or transaction monitoring systems. • Fluency in enterprise fraud detection tooling and methodologies, plus understanding of how F500 regulatory reporting (FinCEN, BSA/AML), law enforcement coordination, and case escalation actually work. • Prior rubric, fraud-investigation training curriculum, or case documentation authorship is a plus.

$60 - $70 / hourOpen / Referral verified
Finance / Remote

Retail Banking Expert

Role Overview • Mercor is seeking senior retail banking professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise consumer banking contexts. • The workflows are calibrated to the customer volume, regulatory complexity, and operational stakes of Fortune 500 and large national retail banks. • Contributors design enterprise retail banking scenarios, draft reference outputs, and write rubrics that capture how senior F500 retail banking operators think. — Key Responsibilities • Construct enterprise retail banking scenarios spanning large-scale deposit and lending operations, multi-stakeholder branch network decisions, and complex regulatory examination or compliance review cycles. • Build tasks across F500 consumer lending (mortgage, auto, personal), branch and digital banking operations, credit risk and underwriting, fraud prevention, and customer experience/retention strategy. • Develop retail banking operations scenarios involving tools such as core banking platforms (FIS, Fiserv, Temenos), loan origination systems (LOS), CRM platforms, and fraud detection/AML monitoring tools in F500 stacks. • Apply enterprise retail banking frameworks (credit risk scoring models, CFPB/regulatory compliance standards, omnichannel service design, KYC/AML protocols) and produce reference lending policies, risk assessments, and executive-level operational narratives. • Author rubrics that distinguish authentic enterprise retail banking judgment from generic textbook or licensing-exam-level recall. — Ideal Qualifications • 5+ years working in consumer lending, branch operations, credit risk, or compliance at a Fortune 500 retail bank (JPMorgan Chase, Bank of America, Wells Fargo, Citi, U.S. Bank). • Direct ownership of F500-scale lending portfolios, branch/digital banking operations, or regulatory compliance programs. • Fluency in enterprise retail banking tooling and frameworks, plus understanding of how F500 regulatory examinations (OCC, CFPB, FDIC), KYC/AML compliance, and credit risk governance actually work. • Prior rubric, banking-training curriculum, or policy/risk documentation authorship is a plus.

$60 - $70 / hourOpen / Referral verified
STEM / Remote

Education Expert - Sourcing Funnel (Private)

Role Overview — Mercor is collaborating with a leading AI lab to engage experienced educators across all areas of practice — including K-12 and higher-education teaching, curriculum development and instructional design, assessment and psychometrics, special education, tutoring and academic support, and educational technology. Contributors help build AI systems that reason about real educational work by translating everyday teaching, assessment, and instructional-design workflows, judgments, and decision-making into structured, high-quality training data. — Key Responsibilities — \- Design realistic educational scenarios and tasks drawn from your day-to-day work (e.g., lesson planning, assessment and item writing, grading against a rubric, differentiation and intervention, curriculum and unit design, IEPs and accommodations, student feedback) — \- Review and compare AI-generated educational outputs for accuracy, standards alignment (e.g., Common Core, NGSS, state or discipline standards), grade-level appropriateness, and sound pedagogical judgment — \- Create structured examples that reflect how educators actually reason through problems — \- Provide clear written feedback that improves how AI performs teaching and instructional tasks — \- Collaborate asynchronously with the research team — Ideal Qualifications — \- 3+ years of professional experience in education (K-12 or higher-ed teaching, curriculum/instructional design, assessment, special education, tutoring, or a related field) — \- A state teaching license/certification, National Board Certification, or an advanced degree (MEd/EdD/PhD or a subject master's) preferred, but not required — \- Bachelor's degree in Education or a subject-matter field — \- Comfortable with common classroom and instructional tools (e.g., Google Classroom, Canvas, an SIS such as PowerSchool, assessment and curriculum platforms) — \- Strong written communication and attention to detail — More About the Opportunity — \- Open to all education specialties and grade bands — contribute where your expertise is strongest — \- Work spans task design, evaluation, and structured feedback on AI educational outputs — \- Strong contributors advance into reviewer, lead, and domain-expert roles — Application Process — \- Submit a resume or a short summary of your teaching or education experience — \- Complete a short form on your field, specialties, and credentials — \- Selected applicants may complete a brief sample task — \- Follow-up typically provided within a few days

$45 - $60 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Dutch

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Dutch and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Dutch, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Dutch music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Dutch genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Dutch • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$28 - $60 / hourOpen / Referral verified
Finance / United States Remote

Finance Program Coordinator — AI Training Data Operations

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking a talented Program Coordinator with a finance background to bring hands-on coordination and organizational rigor to our AI training data program, keeping finance-specific tasks and rater teams on track and consistent. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide day-to-day coordination between finance-domain raters and the broader training-data program to close execution gaps and keep work on schedule. • Design and maintain task-tracking, escalation, and quality-check workflows specific to finance evaluation tasks. • Evaluate rater throughput and quality signals on finance tasks and provide clear, written status updates and feedback. • Triage and resolve finance-specific rater questions, escalating ambiguous cases to subject matter experts. • Collaborate with subject matter experts and program leads to ensure consistency and accuracy across finance training data. — 3. Core Qualifications • 5–10 years of professional experience in finance or finance operations, with a track record of coordinating cross-functional teams or projects. • Direct experience managing or coordinating a team of contributors/reviewers (raters, analysts, or similar) against deadlines and quality bars. • Demonstrable career progression (e.g., Analyst → Senior Analyst → Program/Project Coordinator). • Ability to engage reliably for at least 35 hours/week during weekdays. • Strong verbal and written communication skills, organizational skills, and problem-solving skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$40 - $60 / hourOpen / Referral verified
Finance / Remote

Expert Project Manager

We're hiring an Expert Project Manager to help run projects that push the frontier of LLM browsing capabilities. Day to day: tracking project and annotator performance in Google Sheets, reviewing annotator quality, handling contributor communications, and more. We're looking for someone with high agency, strong organization, an analytical mindset, and real investment in seeing LLM training projects succeed. Ops or project management background preferred, but the bigger thing is being able to pick up an unfamiliar problem and own it start to finish.

$50 - $60 / hourOpen / Referral verified
Finance / Remote

Supply Chain Expert

Role Overview • Mercor is seeking senior supply chain professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise supply chain contexts. • The workflows are calibrated to the network complexity, volume scale, and operational stakes of Fortune 500 and large public company supply chain operations. • Contributors design enterprise supply chain scenarios, draft reference outputs, and write rubrics that capture how senior F500 supply chain operators think. — Key Responsibilities • Construct enterprise supply chain scenarios spanning large-scale demand planning, multi-tier supplier networks, and complex logistics or procurement cycles at F500 accounts. • Build supply chain tasks across F500 sourcing and procurement strategy, inventory optimization, distribution and logistics network design, supplier risk management, and S&OP (sales and operations planning). • Develop supply chain and operations scenarios involving tools such as SAP, Oracle SCM, Blue Yonder, Kinaxis, and enterprise TMS/WMS platforms in F500 stacks. • Apply enterprise supply chain methodologies (Lean/Six Sigma, network optimization modeling, risk-adjusted sourcing strategies) and produce reference sourcing plans, network strategies, and executive-level supply chain briefings. • Author rubrics that distinguish authentic enterprise supply chain judgment from generic textbook or framework recall. — Ideal Qualifications • 5+ years working in supply chain, procurement, or logistics at a Fortune 500 manufacturer, retailer, or logistics provider (Amazon, Walmart, UPS, Procter & Gamble, Unilever) or inside an F500 supply chain/operations organization. • Direct ownership of F500 supplier relationships, F500 distribution networks, or F500 procurement programs. • Fluency in enterprise supply chain tooling and methodologies, plus understanding of how F500 sourcing, procurement, and vendor contract review actually work. • Prior rubric, supply chain training curriculum, or operations documentation authorship is a plus.

$60 - $70 / hourOpen / Referral verified
Math / Remote

Data Science and Analytics Experts

Role Overview • Mercor is seeking senior data science and analytics professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise data and analytics contexts. • The workflows are calibrated to the data scale, model complexity, and business-critical stakes of Fortune 500 and large public company data operations. • Contributors design enterprise data science scenarios, draft reference outputs, and write rubrics that capture how senior F500 data leaders think. — Key Responsibilities • Construct enterprise data science scenarios spanning large-scale predictive modeling, multi-stakeholder analytics governance, and complex data infrastructure decisions at F500 accounts. • Build analytics tasks across F500 machine learning model development, enterprise data pipelines, business intelligence at scale, experimentation/causal inference, and data strategy. • Develop data and MLOps scenarios involving tools such as Snowflake, Databricks, Python/R, SQL, Tableau/Power BI, and enterprise ML platforms (SageMaker, Vertex AI, MLflow) in F500 stacks. • Apply enterprise data science methodologies (statistical rigor, A/B testing frameworks, model validation, MLOps best practices) and produce reference analyses, model documentation, and executive-level insights. • Author rubrics that distinguish authentic enterprise data science judgment from generic textbook or tutorial-level recall. — Ideal Qualifications • 5+ years working as a data scientist, analytics leader, or ML engineer at a Fortune 500 technology or enterprise organization (Google, Meta, Amazon, Microsoft, Netflix) or inside an F500 data/analytics organization (JPMorgan, UPS, Unilever, PepsiCo, Walmart). • Direct ownership of F500 data products, F500 analytics initiatives, or F500 machine learning systems in production. • Fluency in enterprise data science tooling and methodologies, plus understanding of how F500 data governance, privacy compliance, and cross-functional stakeholder alignment actually work. • Prior rubric, technical curriculum, or model documentation authorship is a plus.

$60 - $70 / hourOpen / Referral verified
Business / Remote

Higher Education Expert

Role Overview • Mercor is seeking senior higher education professionals to build evaluation tasks for AI systems operating in large university and higher education institution contexts. • The workflows are calibrated to the academic complexity, stakeholder diversity, and institutional stakes of large research universities and higher education systems. • Contributors design higher education scenarios, draft reference outputs, and write rubrics that capture how senior faculty and university administrators think. — Key Responsibilities • Construct higher education scenarios spanning curriculum and program design, multi-stakeholder faculty senate or accreditation review, and complex institutional budget or enrollment planning cycles. • Build education tasks across academic program development, student affairs and retention, research administration and grant compliance, faculty governance, and institutional advancement/fundraising. • Develop higher-ed operations scenarios involving tools such as Banner/Workday Student, Canvas/Blackboard, CRM platforms for admissions and advancement, and research administration systems (Cayuse, InfoEd). • Apply higher education frameworks (accreditation standards, shared governance models, learning outcomes assessment, enrollment management strategy) and produce reference academic plans, accreditation documentation, and administrator/board-level communications. • Author rubrics that distinguish authentic higher education judgment from generic academic or policy-manual recall. — Ideal Qualifications • 5+ years working as faculty, an academic administrator, or in a senior operational role at a college, university, or higher education system. • Direct ownership of academic programs, enrollment/retention initiatives, research administration, or institutional strategic priorities. • Fluency in higher education tooling and frameworks, plus understanding of how university budgets, accreditation compliance, and shared governance actually work. • Prior rubric, curriculum-design, or faculty-training content authorship is a plus.

$60 - $70 / hourOpen / Referral verified
Code / Remote

Moonlight MCQA - Norwegian Generalist

Write original general-knowledge multiple-choice questions in Norwegian for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: master's or PhD, 2 to 6 years of professional experience, educated or based in Norway, native-level Norwegian.

$48.51 - $59.29 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Norwegian Law

Write original multiple-choice questions on Norwegian law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Norwegian equivalent, ideally 2+ years practising, educated or based in Norway, native-level Norwegian.

$48.51 - $59.29 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - German

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in German and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in German, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the German music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary German genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in German • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$30 - $58 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Swedish

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Swedish and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Swedish, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Swedish genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Swedish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$42 - $78 / hourOpen / Referral verified
Business / Remote

Sales and Marketing Expert

Role Overview • Mercor is seeking senior sales and marketing professionals to build evaluation tasks for AI systems operating in Fortune 500 go-to-market contexts. • The workflows are calibrated to the deal sizes, stakeholder complexity, and brand stakes of Fortune 500 and large public companies. • Contributors design enterprise GTM scenarios, draft reference outputs, and write rubrics that capture how senior F500 operators think. — Key Responsibilities • Construct enterprise sales scenarios spanning $1M+ ACV deals, multi-stakeholder buying committees, and complex procurement cycles at F500 accounts. • Build marketing tasks across F500 brand strategy, enterprise ABM, demand generation at scale, lifecycle, and category positioning. • Develop RevOps and GTM scenarios involving Salesforce Enterprise, Marketo, 6sense, Gong, and Outreach in F500 stacks. • Apply enterprise sales methodologies (MEDDIC, Challenger, Force Management) and produce reference deal strategies, account plans, and executive narratives. • Author rubrics that distinguish authentic enterprise GTM judgment from generic playbook recall. — Ideal Qualifications • 5+ years selling, marketing, or running RevOps at a Fortune 500 enterprise software vendor (Salesforce, Oracle, ServiceNow, SAP, Workday, Microsoft, AWS) or inside an F500 brand or marketing organization (P&G, JPMorgan, Unilever, Microsoft, PepsiCo). • Direct ownership of F500 accounts, F500 brand campaigns, or F500 demand programs. • Fluency in enterprise GTM tooling and methodologies, plus understanding of how F500 budgets, procurement, and legal review actually work. • Prior rubric, sales-enablement curriculum, or training-content authorship is a plus.

$60 - $70 / hourOpen / Referral verified
STEM / Remote

Applied History & Political Science Benchmark Specialist

Role Overview — We are seeking experts in history and political science to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core history and political science domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — History & Political Science Domains Covered — National Security, Public Policy, Business History, Environmental History, Latin American History. — Key Responsibilities • Author original history and political science questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise reasoning in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in History, Political Science, International Relations, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of historiographical methods, political theory, and comparative analysis • Research publications or policy experience is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$44 - $56 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Spanish (MX)

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Spanish (MX) and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Spanish (MX), and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Spanish (MX) music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Spanish (MX) genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Spanish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$13 - $54 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Spanish (ES)

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Spanish (ES) and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Spanish (ES), and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Spanish (ES) music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Spanish (ES) genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Spanish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$39 - $54 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Mandarin Chinese

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Mandarin Chinese and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Mandarin Chinese, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Mandarin Chinese music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Mandarin Chinese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Mandarin Chinese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$18 - $54 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - French

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in French and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in French, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the French music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary French genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in French • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $54 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - English (US/UK/Australia)

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Experience as a songwriter, lyricist, performer, composer, or music journalist • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Domain Expert Interview • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$23 - $54 / hourOpen / Referral verified
Language / Remote

Music Production Expert - French

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in French and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in French, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary French genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in French • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $54 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Spanish (MX)

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Spanish (MX) and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Spanish (MX), and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Spanish (MX) genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Spanish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$13 - $54 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Canadian French)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Canada-based voice actors with native Canadian French fluency and international French (neutral, accent-free) delivery to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native Canadian French speaker currently based in Canada, with the ability to speak international French — neutral, without strong regional accent • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 / hourOpen / Referral verified
Code / Remote

Audio Engineer (Speech / TTS Audio Specialist) – German

Overview We're looking for audio engineers with strong experience editing and quality-checking voice recordings for speech and text-to-speech (TTS) applications, specifically for German-language audio. This role focuses on preparing high-quality audio datasets for machine learning systems, ensuring clarity, consistency, and technical compliance across large volumes of German speech data. Native or near-native German fluency is required to accurately assess speech quality, pronunciation, and naturalness. • * * — What You'll Do • Edit and process German voice recordings for use in TTS and speech AI systems • Perform audio cleanup including silence trimming, breath reduction/removal, de-clicking, de-essing, and noise reduction • Conduct detailed quality assurance checks for: • Clipping and distortion • Background noise and artifacts • Loudness consistency and level balancing • Ensure audio meets strict technical specifications for ML training pipelines • Evaluate recordings for German pronunciation accuracy, naturalness, and fluency • Work with large batches of recordings, maintaining consistency and throughput • Collaborate with teams working on German speech datasets, voice talent recordings, and model training workflows • * * — Basic Qualifications • Native or near-native German speaker with strong listening intuition for German speech nuances • Experience editing speech or voice recordings (not just music production) • Strong understanding of speech audio quality standards and common issues in recorded dialogue • Hands-on experience with tools like iZotope RX, Pro Tools, or similar audio restoration software • Ability to identify and fix artifacts, inconsistencies, and technical defects in audio • Experience working with high-volume audio datasets or structured workflows • * * — Preferred Qualifications • Experience working with TTS vendors or platforms • Prior involvement in speech ML, voice AI, or dataset creation/annotation workflows • Familiarity with audio QA pipelines, labeling, or evaluation processes • Experience directing or editing German voice talent recordings • * * — Ideal Candidate Profile • Detail-oriented with strong critical listening skills in German • Comfortable working with repetitive, high-precision tasks at scale • Able to balance speed and quality in production environments • Familiar with the nuances of human German speech vs synthetic voice requirements

$50 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Indonesian Bahasa - Female)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Indonesia-based female voice actors who are native Bahasa Indonesia speakers to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. Fluency in English is a plus. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native Bahasa Indonesia speaker currently based in Indonesia • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability over the project duration • \[IMP\]: Your voice may be cloned for the clients CX AI Agent so please only apply if you are okay with voice cloning — Preferred Qualifications • Fluency in English in addition to Bahasa Indonesia • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 / hourOpen / Referral verified
Code / Remote

Audio Engineer (Speech / TTS Audio Specialist) – French

Overview We're looking for audio engineers with strong experience editing and quality-checking voice recordings for speech and text-to-speech (TTS) applications, specifically for French-language audio. This role focuses on preparing high-quality audio datasets for machine learning systems, ensuring clarity, consistency, and technical compliance across large volumes of French speech data. Native or near-native French fluency is required to accurately assess speech quality, pronunciation, and naturalness. • * * — What You'll Do • Edit and process French voice recordings for use in TTS and speech AI systems • Perform audio cleanup including silence trimming, breath reduction/removal, de-clicking, de-essing, and noise reduction • Conduct detailed quality assurance checks for: • Clipping and distortion • Background noise and artifacts • Loudness consistency and level balancing • Ensure audio meets strict technical specifications for ML training pipelines • Evaluate recordings for French pronunciation accuracy, naturalness, and fluency • Work with large batches of recordings, maintaining consistency and throughput • Collaborate with teams working on French speech datasets, voice talent recordings, and model training workflows • * * — Basic Qualifications • Native or near-native French speaker with strong listening intuition for French speech nuances • Experience editing speech or voice recordings (not just music production) • Strong understanding of speech audio quality standards and common issues in recorded dialogue • Hands-on experience with tools like iZotope RX, Pro Tools, or similar audio restoration software • Ability to identify and fix artifacts, inconsistencies, and technical defects in audio • Experience working with high-volume audio datasets or structured workflows • * * — Preferred Qualifications • Experience working with TTS vendors or platforms • Prior involvement in speech ML, voice AI, or dataset creation/annotation workflows • Familiarity with audio QA pipelines, labeling, or evaluation processes • Experience directing or editing French voice talent recordings • * * — Ideal Candidate Profile • Detail-oriented with strong critical listening skills in French • Comfortable working with repetitive, high-precision tasks at scale • Able to balance speed and quality in production environments • Familiar with the nuances of human French speech vs synthetic voice requirements

$50 / hourOpen / Referral verified
Multimodal / Remote

Russian Audio Generalist Evaluator Expert (San Francisco Bay Area)

Mercor is seeking a Russian Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project with a leading research lab. In this role, you will work on transcription, annotation, and evaluation tasks that help train and benchmark advanced language models. This is a short-term, structured engagement ideal for candidates with strong academic or analytical backgrounds who are fluent in Russian and English and are based in the San Francisco Bay Area, enabling occasional in-person collaboration if required. — Job Responsibilities — Transcribe and Optimise Audio & Video • Listen to, analyse, and transcribe audio and video content in Russian, following detailed constraints and instructions. • Produce high-quality written outputs in Russian, with supporting work in English when required. • Ensure clarity, accuracy, and strict adherence to formatting and stylistic guidelines. • Capture nuances such as tone, intent, formal vs. informal register, regional expressions, dialectal variations, and contemporary Russian usage where relevant. — Define and Document Evaluation Standards • Establish clear expectations for correct and high-quality responses in general consumer audio contexts. • Develop detailed evaluation rubrics and grading guidelines in Russian and English. • Document standards to ensure consistency across reviewers and model evaluations. • Identify linguistic nuances, grammatical complexities, colloquialisms, and edge cases specific to Russian. — Conduct Model Testing and Grading • Run prompts through language models and assess generated outputs. • Evaluate responses against predefined criteria for accuracy, completeness, fluency, and instructional clarity. • Provide structured feedback to improve model performance in Russian audio tasks. — Support Benchmarking and Quality Assurance • Participate in QA and review cycles to ensure tasks, rubrics, and outputs meet Mercor’s quality bar. • Maintain consistency and reliability before datasets are integrated into official benchmarks. • Collaborate with project leads to resolve ambiguities and improve task design. — Minimum Qualifications • Strong writing, editing, and critical thinking skills. • Ability to work independently, manage time effectively, and meet deadlines. • Native or near-native fluency in Russian (spoken and written) and professional fluency in English. • Strong familiarity with spoken Russian, regional vocabulary, dialects, and contemporary language usage. • Ability to accurately transcribe and analyse Russian audio content across general consumer contexts. • Must be based in the San Francisco Bay Area. • Available to commit 10–20 hours per week. — Preferred Qualifications • College students or recent graduates. • Background in linguistics, humanities, social sciences, journalism, translation/localization, or technical disciplines. • Prior experience with transcription, annotation, localisation, evaluation, or research workflows in Russian. • Familiarity with regional variations of Russian and contemporary digital language usage. • Interest in AI, language models, or applied research environments. — Application & Onboarding Process • Complete a short AI-led interview (approximately 15 minutes). • If selected, you will be onboarded and invited to begin project work.

$50 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Russian

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Russian and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Russian, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Russian music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Russian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Russian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$35 - $49 / hourOpen / Referral verified
Code / Remote

Moonlight MCQA - Chinese (Simplified) Generalist

Write original general-knowledge multiple-choice questions in Chinese (Simplified) for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: master's or PhD, 2 to 6 years of professional experience, educated or based in China, native-level Chinese (Simplified).

$39.69 - $48.51 / hourOpen / Referral verified
Finance / Remote

Moonlight MCQA - Chinese (Traditional) Finance

Write original multiple-choice questions on Taiwanese or Hong Kong accounting, auditing and tax, in Chinese (Traditional), for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: master's or higher in finance, accounting or business, ideally 2+ years' experience, educated or based in Taiwan or Hong Kong, native-level Chinese (Traditional).

$39.69 - $48.51 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Chinese (Traditional) Law

Write original multiple-choice questions on Taiwanese or Hong Kong law, in Chinese (Traditional), for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or national equivalent, ideally 2+ years practising, educated or based in Taiwan or Hong Kong, native-level Chinese (Traditional).

$39.69 - $48.51 / hourOpen / Referral verified
Code / Remote

Moonlight MCQA - Chinese (Traditional) Generalist

Write original general-knowledge multiple-choice questions in Chinese (Traditional) for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: master's or PhD, 2 to 6 years of professional experience, educated or based in Taiwan or Hong Kong, native-level Chinese (Traditional).

$39.69 - $48.51 / hourOpen / Referral verified
Finance / Remote

Moonlight MCQA - Chinese (Simplified) Finance

Write original multiple-choice questions on Chinese accounting, auditing and tax for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: master's or higher in finance, accounting or business, ideally 2+ years' experience, educated or based in China, native-level Chinese (Simplified).

$39.69 - $48.51 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Chinese (Simplified) Law

Write original multiple-choice questions on Chinese law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Chinese equivalent, ideally 2+ years practising, educated or based in China, native-level Chinese (Simplified).

$39.69 - $48.51 / hourOpen / Referral verified
Business / Remote

Insurance Experts

Role Overview • Mercor is seeking senior insurance professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise insurance and risk contexts. • The workflows are calibrated to the underwriting complexity, regulatory stakes, and claims scale of Fortune 500 and large public insurance carriers. • Contributors design enterprise insurance scenarios, draft reference outputs, and write rubrics that capture how senior F500 insurance operators think. — Key Responsibilities • Construct enterprise insurance scenarios spanning complex commercial underwriting, multi-stakeholder claims adjudication, and regulatory filing or compliance review cycles at F500 carriers. • Build insurance tasks across F500 underwriting and risk assessment, claims management, actuarial pricing and reserving, reinsurance strategy, and regulatory/compliance affairs. • Develop insurance operations scenarios involving tools such as Guidewire, Duck Creek, actuarial modeling platforms (SAS, R), and enterprise claims/policy administration systems in F500 stacks. • Apply enterprise insurance methodologies (actuarial risk modeling, loss reserving, ERM frameworks, NAIC compliance standards) and produce reference underwriting guidelines, claims strategies, and executive-level risk narratives. • Author rubrics that distinguish authentic enterprise insurance judgment from generic textbook or licensing-exam recall. — Ideal Qualifications • 5+ years working in underwriting, claims, actuarial, or risk management at a Fortune 500 insurance carrier or reinsurer (State Farm, Allstate, Chubb, AIG, Berkshire Hathaway) or inside an F500 enterprise risk/insurance organization. • Direct ownership of F500-scale underwriting portfolios, claims operations, or actuarial pricing models. • Fluency in enterprise insurance tooling and methodologies, plus understanding of how F500 regulatory compliance, reinsurance treaties, and reserving practices actually work. • Prior rubric, actuarial/underwriting training curriculum, or claims documentation authorship is a plus.

$50 - $60 / hourOpen / Referral verified
Legal / Remote

Senior Civil Legal Paralegal / Accredited Representative

Mercor is seeking experienced Senior Civil Legal Paralegals and Accredited Representatives to review AI-generated legal guidance involving civil legal services workflows. Experts will assess procedural accuracy, documentation quality, and practical usefulness across common civil legal matters. • * * — Responsibilities • Review AI-generated legal guidance. • Evaluate procedural correctness and documentation. • Assess housing, family, consumer debt, and bankruptcy scenarios. • Identify filing errors and missing procedural steps. • Provide structured written feedback. • * * — Required Qualifications • Significant experience as a Senior Paralegal or Accredited Representative. • Strong knowledge of civil legal procedures. • Experience supporting housing, family, consumer, or debt matters. • Excellent written communication skills. • * * — Preferred Qualifications • Legal aid or nonprofit legal services experience. • Familiarity with jurisdiction-specific filing procedures. • Experience supporting underserved populations. • Experience mentoring legal support staff. • * * — Why Join Mercor? • Bring practical legal operations expertise into AI development. • Help improve AI-generated legal guidance for real-world users. • Competitive consulting rates. • Opportunity to contribute to next-generation legal technology.

$70 / hourOpen / Referral verified
Finance / Remote

Moonlight MCQA - Danish Finance

Write original multiple-choice questions on Danish accounting, auditing and tax for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: master's or higher in finance, accounting or business, ideally 2+ years' experience, educated or based in Denmark, native-level Danish.

$38.59 - $47.17 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Swedish Law

Write original multiple-choice questions on Swedish law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Swedish equivalent, ideally 2+ years practising, educated or based in Sweden, native-level Swedish.

$38.59 - $47.17 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Danish Law

Write original multiple-choice questions on Danish law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Danish equivalent, ideally 2+ years practising, educated or based in Denmark, native-level Danish.

$38.59 - $47.17 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Finnish Law

Write original multiple-choice questions on Finnish law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Finnish equivalent, ideally 2+ years practising, educated or based in Finland, native-level Finnish.

$38.59 - $47.17 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Turkish

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Turkish and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Turkish, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Turkish music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Turkish genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Turkish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $45 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Portuguese (global)

Location: Remote — Fluent Language Skills Required: English & Portuguese (global, excluding Brazilian Portuguese). Native fluency in English and Portuguese (global, excluding Brazilian Portuguese) is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$29 - $45 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Dutch

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Dutch and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Dutch, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Dutch genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Dutch • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$28 - $60 / hourOpen / Referral verified
Medical / Remote

Adult Inpatient Nurses (RN)

We're hiring experienced Adult Inpatient Nurses (RNs) to help train and evaluate AI systems used in clinical and healthcare settings. This role is ideal for nurses who want to apply their frontline experience to improve the accuracy, safety, and reliability of medical AI tools. — You'll work on projects that require deep clinical judgment, attention to detail, and the ability to translate real-world bedside documentation practices into structured feedback for AI systems. — Key Responsibilities • Review and evaluate AI-generated clinical outputs based on nursing flowsheet documentation • Validate accuracy, completeness, and adherence to current documentation standards • Annotate and structure inpatient nursing assessment data for AI training datasets • Provide expert feedback on nursing assessments and documentation practices • Identify gaps, inconsistencies, or risks in AI-generated responses • Ask clarifying questions when annotation guidance is ambiguous, and contribute to guideline refinement • Collaborate with technical teams to improve model performance • Contribute to the development of high-quality clinical benchmarks — Basic Qualifications • Active RN license (U.S., outside California) • Recent adult acute care inpatient bedside experience, ideally in a non-procedural area of specialty • Experience within the last 5–10 years, reflecting familiarity with current workflows and documentation practices (current documentation experience is prioritized over total years of nursing experience) • Comfortable performing and documenting comprehensive nursing assessments • Comfortable following detailed annotation guidelines consistently • Excellent written communication skills and responsiveness to feedback • Ability to work independently and meet deadlines — Preferred Qualifications • Epic EHR experience (most common platform, closely aligns with many workflows) • Comfortable learning new annotation tools and web-based platforms • Able to navigate transcripts efficiently and use AI-assisted tools appropriately while still verifying outputs • Experience reviewing charts for quality improvement, utilization review, CDI, or clinical informatics • Previous annotation, chart abstraction, or healthcare AI experience

$55 - $65 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Urdu

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Urdu and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Urdu, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Urdu music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Urdu genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Urdu • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$14 - $42 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Portuguese

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Portuguese and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Portuguese, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Portuguese music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Portuguese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Portuguese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$18 - $42 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Hindi

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Hindi and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Hindi, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Hindi music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Hindi genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Hindi • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$14 - $42 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Hebrew

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Hebrew and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Hebrew, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Hebrew music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Hebrew genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Hebrew • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $42 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Greek

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Greek and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Greek, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Greek music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Greek genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Greek • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $42 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Hebrew

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Hebrew and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Hebrew, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Hebrew genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Hebrew • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $42 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Urdu

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Urdu and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Urdu, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Urdu genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Urdu • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$14 - $42 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Hindi

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Hindi and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Hindi, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Hindi genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Hindi • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$14 - $42 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Portuguese

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Portuguese and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Portuguese, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Portuguese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Portuguese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$18 - $42 / hourOpen / Referral verified
Code / Remote

Moonlight MCQA - Czech Generalist

Write original general-knowledge multiple-choice questions in Czech for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: master's or PhD, 2 to 6 years of professional experience, educated or based in Czechia, native-level Czech.

$33.96 - $41.5 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Czech Law

Write original multiple-choice questions on Czech law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Czech equivalent, ideally 2+ years practising, educated or based in Czechia, native-level Czech.

$33.96 - $41.5 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Japanese

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Japanese and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Japanese, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Japanese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Japanese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$35 - $62 / hourOpen / Referral verified
Medical / Remote

Social Work Expert

Role Overview • Mercor is seeking senior social work professionals to build evaluation tasks for AI systems operating in large human services and clinical social work contexts. • The workflows are calibrated to the case complexity, stakeholder sensitivity, and client-outcome stakes of large public agencies, healthcare systems, and community-based organizations. • Contributors design social work scenarios, draft reference outputs, and write rubrics that capture how senior clinicians and case management leaders think. — Key Responsibilities • Construct social work scenarios spanning complex case assessment, multi-stakeholder care coordination, and crisis intervention or mandated reporting cycles. • Build tasks across clinical social work practice, child and family welfare, healthcare/medical social work, substance use and mental health services, and community-based case management. • Develop practice scenarios involving tools such as electronic health/case record systems (EHR, CCWIS), risk and needs assessment instruments, and interagency referral platforms. • Apply social work frameworks (biopsychosocial assessment, trauma-informed care, strengths-based practice, NASW Code of Ethics) and produce reference case plans, clinical documentation, and interdisciplinary team communications. • Author rubrics that distinguish authentic clinical and ethical judgment from generic textbook or policy-manual recall. — Ideal Qualifications • 5+ years working as a licensed clinical social worker (LCSW/LMSW) or senior case management professional at a hospital system, child welfare agency, behavioral health organization, or large nonprofit. • Direct ownership of clinical caseloads, family/child welfare cases, or program-level service delivery. • Fluency in social work practice standards and case management tooling, plus understanding of how mandated reporting, interagency coordination, and ethical/legal compliance actually work. • Prior rubric, clinical training curriculum, or case documentation authorship is a plus.

$40 - $50 / hourOpen / Referral verified
Business / Remote

K-12 Education Expert

Role Overview • Mercor is seeking senior K-12 education professionals to build evaluation tasks for AI systems operating in large school district and public education contexts. • The workflows are calibrated to the instructional complexity, stakeholder diversity, and student-outcome stakes of large public school districts and state education systems. • Contributors design K-12 education scenarios, draft reference outputs, and write rubrics that capture how senior educators and district leaders think. — Key Responsibilities • Construct K-12 scenarios spanning curriculum design and adoption, multi-stakeholder IEP/504 planning, and complex district budget or policy review cycles. • Build education tasks across instructional design, classroom differentiation and special education, assessment and standards alignment, student support services, and family/community engagement. • Develop education operations scenarios involving tools such as PowerSchool, Canvas/Google Classroom, state assessment platforms, and district-level MTSS/RTI systems. • Apply K-12 pedagogical frameworks (Universal Design for Learning, standards-based grading, data-driven instruction, restorative practices) and produce reference lesson plans, IEP documentation, and administrator-level communications. • Author rubrics that distinguish authentic K-12 educator judgment from generic pedagogical or textbook recall. — Ideal Qualifications • 5+ years teaching, administering, or leading instruction at a public or private K-12 school, district office, or state education agency. • Direct ownership of classroom instruction, IEP/special education caseloads, curriculum programs, or school/district-level initiatives. • Fluency in K-12 education tooling and frameworks, plus understanding of how school budgets, state standards compliance, and family/community engagement actually work. • Prior rubric, curriculum-writing, or teacher-training content authorship is a plus.

$40 - $50 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Hungarian Law

Write original multiple-choice questions on Hungarian law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Hungarian equivalent, ideally 2+ years practising, educated or based in Hungary, native-level Hungarian.

$30.87 - $37.73 / hourOpen / Referral verified
Language / Remote

Music Production Expert - German

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in German and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in German, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary German genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in German • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$30 - $58 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Slovak

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Slovak and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Slovak, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Slovak music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Slovak genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Slovak • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$20 - $36 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Italian

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Italian and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Italian, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Italian music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Italian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Italian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$20 - $36 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Hungarian

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Hungarian and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Hungarian, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Hungarian music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Hungarian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Hungarian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$20 - $36 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Hungarian

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Hungarian and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Hungarian, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Hungarian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Hungarian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$20 - $36 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Slovak

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Slovak and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Slovak, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Slovak genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Slovak • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$20 - $36 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Italian

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Italian and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Italian, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Italian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Italian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$20 - $36 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Thai

Location: Remote — Fluent Language Skills Required: English & Thai. Native fluency in English and Thai is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$24 - $35 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Spanish (ES)

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Spanish (ES) and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Spanish (ES), and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Spanish (ES) genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Spanish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$39 - $54 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Arabic

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Arabic and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Arabic, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Arabic music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Arabic genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Arabic • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $34 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Ukrainian Law

Write original multiple-choice questions on Ukrainian law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Ukrainian equivalent, ideally 2+ years practising, educated or based in Ukraine, native-level Ukrainian.

$27.78 - $33.96 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Mandarin Chinese

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Mandarin Chinese and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Mandarin Chinese, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Mandarin Chinese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Mandarin Chinese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$18 - $54 / hourOpen / Referral verified
Business / United States Remote

Project Coordinator

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — Are you ready to help shape the future of artificial intelligence? Join a leading AI lab's cutting-edge GenAI team, where you'll be at the forefront of building groundbreaking AI models. We're seeking talented Project Coordinators to support and accelerate world-class AI research and data operations — acting as the operational bridge between program leadership and a growing team of domain experts. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. What You'll Do • Act as a day-to-day point of contact for AI data projects, helping keep workflows and operations running smoothly. • Work closely with program leads and domain experts — take inputs from leadership, turn them into clear guidelines, and share them with expert teams. • Help onboard and support new experts as the program grows. • Join weekly business reviews, keep notes and action items organized, and lead calls when needed. • Help spot quality issues, understand root causes, and share findings across teams. • Contribute to creating and refining project guidelines with research and product partners. • Flag potential roadblocks early and help get them resolved. • Track project deliverables using Google Sheets and Excel. • Stay flexible and adapt as project needs evolve. — 3. Qualifications • Location: Must be based in the USA. • Education: STEM background strongly preferred. • Experience: 3+ years in project coordination or project management, with the ability to coordinate large, cross-functional teams. • AI fluency: hands-on experience on AI training-data or human-data projects — as a project coordinator, team lead, EPM, or expert contributor. • Ideal: prior project coordination on AI/data programs at organizations like Meta, TikTok, or Amazon, or EPM/project-lead experience on Mercor projects or other AI tranining projects. • Skills: strong data management experience (Excel, SQL) and working knowledge of coding. • Leadership: prior people management experience or demonstrated ability to lead large groups effectively. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$45 - $55 / hourOpen / Referral verified
Language / Remote

Music Production Expert - English (US/UK/Australia)

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Domain Expert Interview • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$23 - $54 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Indonesian

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Indonesian and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Indonesian, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Indonesian music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Indonesian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Indonesian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$11 - $30 / hourOpen / Referral verified
Code / India Remote

Software Engineer, Full Stack — India

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Full-Stack Software Engineers with hands-on production experience across modern application stacks — Python, Java, Rust, C#, C++, and TypeScript — to build, integrate, and stress-test real software on top of frontier models before they ship. Engineers work directly with the AI lab's engineering manager on fast-moving projects and are expected to onboard quickly and contribute working code early. — This is a full-time engagement of 40 hours per week. — 2. Key Responsibilities • Build and ship full-stack applications, services, and internal tools — in Python, Java, Rust, C#, C++, or TypeScript — that exercise frontier model capabilities end to end. • Integrate pre-release model APIs into working software, including the surrounding scaffolding: tool interfaces, evaluation harnesses, and telemetry. • Diagnose and document model and integration failure modes surfaced while building, translating engineering observations into clear written feedback for the research team. • Prototype quickly against shifting requirements, delivering working increments with limited up-front specification. • Collaborate with the engineering manager and other engineers to maintain consistency in code quality, architecture, and technical documentation. — 3. Core Qualifications • 3+ years of dedicated professional experience building and shipping production software at a recognized, top-tier organization. • Production experience in at least one of Python, Java, Rust, C#, or C++, plus working proficiency in a second language — this cohort is intentionally staffed across multiple language ecosystems. • Hands-on experience across the stack: backend services and APIs, a modern front-end framework (React or equivalent), relational or document databases, and cloud deployment. • Demonstrated ability to onboard onto an unfamiliar codebase and deliver working code quickly with minimal ramp-up. • Demonstrable career progression. • Ability to engage reliably for at least 40 hours/week during weekdays. • Strong written communication skills and the ability to explain complex technical decisions clearly.

$25 - $30 / hourOpen / Referral verified
Finance / Remote

Moonlight MCQA - Arabic Finance

Write original multiple-choice questions on accounting, auditing and tax in your jurisdiction, in Arabic, for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: master's or higher in finance, accounting or business, ideally 2+ years' experience, educated or based in an Arabic-speaking country, native-level Arabic.

$24.26 - $29.65 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Arabic Law

Write original multiple-choice questions on the law of your jurisdiction, in Arabic, for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or national equivalent, ideally 2+ years practising, educated or based in an Arabic-speaking country, native-level Arabic.

$24.26 - $29.65 / hourOpen / Referral verified
Code / Remote

Moonlight MCQA - Arabic Generalist

Write original general-knowledge multiple-choice questions in Arabic for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: master's or PhD, 2 to 6 years of professional experience, educated or based in an Arabic-speaking country, native-level Arabic.

$24.26 - $29.65 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Russian

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Russian and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Russian, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Russian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Russian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$35 - $49 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Malay

Location: Remote — Fluent Language Skills Required: English & Malay. Native fluency in English and Malay is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$17 - $25 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Indonesian

Location: Remote — Fluent Language Skills Required: English & Indonesian. Native fluency in English and Indonesian is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$17 - $25 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Vietnamese

Location: Remote — Fluent Language Skills Required: English & Vietnamese. Native fluency in English and Vietnamese is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$17 - $25 / hourOpen / Referral verified
Finance / Hyderabad Onsite

Fraud Analyst – Content & Reviews Abuse

About the Role — You'll join a trust & safety function protecting a high-traffic, consumer-facing platform from fraud and abuse. Partnering with product, engineering, and operations stakeholders, you'll investigate ambiguous, evolving abuse patterns — applying investigative rigor and creative analysis to get to root cause. Over time, you'll build specialized subject-matter expertise across the many forms this abuse can take. — We're looking for candidates with a track record in fraud detection, investigation, and risk assessment — reviewing transactions, spotting anomalies in data, and evaluating exposure. This background could come from insurance fraud, forensic accounting, banking/financial crimes, or similar. Certifications such as CFE or ACFE are a nice-to-have, not required. — What You'll Do • Investigate ambiguous cases end-to-end, turning findings into data-backed conclusions that keep user trust as the top priority • Pull together varied data sources and signals to trace deceptive or manipulative activity, and help design new detection experiments • Document findings clearly — a well-reasoned narrative backed by concrete evidence — so others can act on your conclusions — Work Setup • Full-time, onsite in India • Must be available to work India Standard Time hours • Requires a private, secure workspace (no shared or public settings), given the sensitivity of the content involved • Prior experience working on a Mercor project is mandatory — Minimum Qualifications • Bachelor's degree in a technical discipline, or equivalent hands-on experience • 4+ years analyzing data — spotting trends, building summary stats, turning raw data into decisions • 4+ years working end-to-end on ambiguous, open-ended analytical projects • Comfortable with SQL and Python for analysis; strong problem-solving instincts and efficient execution • Strong written/verbal communication, including translating dense policy concepts into plain language • Track record presenting analytical findings to senior stakeholders • Able to produce thorough case write-ups and detailed observation notes — Preferred Qualifications • Familiarity with how fraud/abuse evolves on online platforms over time • Background in program or policy management, customer/merchant experience, or business strategy • Exposure to content moderation, policy enforcement, customer support, or policy design • Sharp problem-solving and critical-thinking instincts with strong attention to detail • Some familiarity with AI/ML or generative AI concepts • CFE/ACFE certification a plus

$15 - $25 / hourOpen / Referral verified
Code / Remote

AI Data Generalist — Egocentric Video Review

About the Role — We are looking for 5 detail-oriented generalists to support a series of short-term projects involving egocentric (first-person) video data. The engagement will run for approximately 2 weeks, covering 3 projects, with each individual project expected to last between 3 days and 1 week. The first project is expected to begin Monday morning 17th August (PST). — What You’ll Do — You will review egocentric video data and complete structured quality-review tasks according to project-specific guidelines. The work requires strong attention to detail, consistency, and the ability to make precise judgments about actions occurring in first-person video. — Project 1: Action Segment Boundary Review — For the first project, you will review the boundaries of action segments within egocentric videos. — Dataset • 50 egocentric videos • Approximately 1–2 minutes per video • Approximately 38 action segments per video on average • Estimated review time: ~30 minutes per video — Responsibilities • Review action segments identified within each video • Determine whether segment start and end boundaries accurately correspond to the action being performed • Adjust or flag incorrect boundaries according to provided guidelines • Maintain consistent judgment across a high volume of short action segments • Complete assigned videos within the project timeline — Additional egocentric-data projects will follow during the 2-week engagement. Detailed instructions will be provided at the start of each project. — Who We’re Looking For — You may be a strong fit if you: • Have excellent attention to detail and visual comprehension • Can carefully distinguish between closely related actions and moments in video • Are comfortable performing structured, repetitive review work while maintaining accuracy • Can quickly learn and consistently apply detailed annotation guidelines • Have strong written English communication skills — Prior experience with video annotation, data labeling, computer vision datasets, or egocentric video is helpful but not required. — Engagement Details • Openings: 5 • Duration: Approximately 2 weeks • Projects: 3 short-term projects • Individual project duration: Approximately 3 days to 1 week • First project start: Monday morning PST

$20 - $25 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Turkish

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Turkish and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Turkish, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Turkish genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Turkish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $45 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Greek

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Greek and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Greek, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Greek genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Greek • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $42 / hourOpen / Referral verified
Multimodal / Remote

Household Activity Video Contributor (US Based)

About this opportunity — Mercor is partnering with a leading physical AI company to collect first-person video of everyday household activities for training the next generation of robotics and physical AI systems. In this role, you will record short, first-person clips of ordinary tasks around your home, helping build and refine the datasets these models learn from. Your recordings will directly support the development of systems that can understand how people move, reach, and handle objects in real-world environments, the everyday actions that come naturally to people but remain difficult for machines to learn. — What you will do • Record short first-person clips of normal everyday activities at home, using a simple head mount phone holder so the camera sees what you see • Example tasks: loading the dishwasher, folding laundry, wiping down counters, putting away groceries, tidying a room, making a bed • Upload your recordings through a provided app • Contributors typically record & upload 10 to 20 hours a week — Eligibility • You must be located in the United States and at least 18 years old • This opportunity is not currently available to residents of California, Illinois, Texas, or Washington • You have an iPhone (12 or newer, not including 16e or 17e), Google Pixel (6 or newer), or Samsung Galaxy (S21 or newer) • You are able to wear a simple head mounted phone holder. If you are accepted and need a head mount, we can provide one at no cost to you • You have a home environment where you can record everyday activities — How it works — 1. Apply to the opportunity listing & complete a brief intake form (~5-10 min) 2. Complete a short, AI-led interview (~15 min) 3. Most applicants will hear back within 1 to 14 days of applying

$10 - $15 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Thai Law

Write original multiple-choice questions on Thai law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Thai equivalent, ideally 2+ years practising, educated or based in Thailand, native-level Thai.

$15.79 - $19.29 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Telugu

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Telugu and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Telugu, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Telugu music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Telugu genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Telugu • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$11 - $19 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Tamil

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Tamil and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Tamil, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Tamil music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Tamil genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Tamil • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$15 - $19 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Bengali

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Bengali and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Bengali, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Bengali music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Bengali genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Bengali • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$11 - $19 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Bengali

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Bengali and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Bengali, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Bengali genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Bengali • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$11 - $19 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Vietnamese

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Vietnamese and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Vietnamese, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Vietnamese music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Vietnamese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Vietnamese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$18 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Thai

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Thai and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Thai, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Thai music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Thai genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Thai • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$18 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - English (India)

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Experience as a songwriter, lyricist, performer, composer, or music journalist • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Domain Expert Interview • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$13 - $17 / hourOpen / Referral verified
Language / Remote

Music Production Expert - English (India)

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Domain Expert Interview • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$13 - $17 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Vietnamese Law

Write original multiple-choice questions on Vietnamese law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Vietnamese equivalent, ideally 2+ years practising, educated or based in Vietnam, native-level Vietnamese.

$12.92 - $15.8 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Punjabi

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Punjabi and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Punjabi, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Punjabi music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Punjabi genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Punjabi • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$15 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Malayalam

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Malayalam and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Malayalam, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Malayalam music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Malayalam genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Malayalam • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$15 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Arabic

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Arabic and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Arabic, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Arabic genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Arabic • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $34 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Indonesian

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Indonesian and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Indonesian, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Indonesian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Indonesian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$11 - $30 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Punjabi

Location: Remote — Fluent Language Skills Required: English & Punjabi. Native fluency in English and Punjabi is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$20 - $22 / hourOpen / Referral verified
Business / United States (Remote)

Dishwasher — ONET Occupation Study

About this study — Mercor is building a new, more accurate version of O\*NET — the U.S. government's framework for describing occupations and the tasks they involve. We are gathering input directly from experienced practitioners to improve how the work of dishwashers is described. — About the role (SOC 35-9021.00 — Dishwashers) — Clean dishes, kitchen, food preparation equipment, and utensils. — What you'll do • Complete a one-time online survey/interview about your day-to-day work as a dishwasher. It takes up to 30 minutes. • There may be an optional follow-up survey, task, or short interview afterward, which would be paid separately. — Who we're looking for (eligible applicants have) • 2+ years of experience working specifically as a dishwasher. • Current work, or work within the last 12 months, in the role. • Direct knowledge of daily kitchen tasks, equipment, and workflow. • U.S.-based. — Not eligible (common confounders) • Cooks and food preparation workers whose primary role is preparing food (SOC 35-2014.00, 35-2021.00). • Servers, hosts, or general front-of-house staff. • Kitchen or restaurant managers. — Reference for the occupation: [https://www.onetonline.org/link/summary/35-9021.00](https://www.onetonline.org/link/summary/35-9021.00)

$16 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Odia

Location: Remote — Fluent Language Skills Required: English & Odia. Native fluency in English and Odia is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$20 - $22 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Bengali

Location: Remote — Fluent Language Skills Required: English & Bengali. Native fluency in English and Bengali is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$20 - $22 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Assamese

Location: Remote — Fluent Language Skills Required: English & Assamese. Native fluency in English and Assamese is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$20 - $22 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Telugu

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Telugu and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Telugu, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Telugu genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Telugu • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$11 - $19 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Tamil

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Tamil and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Tamil, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Tamil genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Tamil • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$15 - $19 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Thai

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Thai and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Thai, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Thai genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Thai • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$18 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Vietnamese

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Vietnamese and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Vietnamese, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Vietnamese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Vietnamese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$18 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Malayalam

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Malayalam and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Malayalam, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Malayalam genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Malayalam • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$15 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Punjabi

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Punjabi and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Punjabi, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Punjabi genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Punjabi • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$15 / hourOpen / Referral verified
Code / Sri Lanka Remote

Sinhala Voice & QA Experts

Mercor is hiring on behalf of a leading AI company for Sinhala Voice & QA Experts. You will place and evaluate Sinhala-language voice conversations to help build and quality-check AI-powered voice agents. — This is a remote, hourly engagement open to candidates based in Sri Lanka, with an immediate start. — Responsibilities • Make outbound calls in Sinhala following provided scenarios • QA those calls for accuracy, tone, clarity, and naturalness • QA other Sinhala conversations and flag linguistic or quality issues • Provide clear written feedback • Participate in check-ins as needed — Requirements • Native or fluent Sinhala speaker • Based in Sri Lanka • Strong attention to detail • Clear written communication — Nice to Have • Customer service experience • Data annotation / QA experience — Engagement Details • Remote, hourly • Immediate start — onboarding as soon as possible • Short onboarding • Ongoing scenario testing and QA

$8 - $12 / hourOpen / Referral verified