Role directory

$50-$99/hr jobs

128 active, referral-verified opportunities.

Finance / Remote

Applied Economics & Finance Benchmark Specialist

Role Overview — We are seeking expert economists and finance professionals to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core economics and finance domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Economics & Finance Domains Covered — Algorithmic Trading & Market Microstructure, Macroprudential Policy, Behavioral Finance & Experimental Economics, Urban Economics, Tokenomics & Decentralized Finance. — Key Responsibilities • Author original economics and finance questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Economics, Finance, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level economic theory, quantitative methods, and financial modeling • Experience with academic research, CFA/CPA certification, or financial industry expertise is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$77 - $98 / hourOpen / Referral verified
Business / Remote

Applied Business & Commerce Benchmark Specialist

Role Overview — We are seeking business and commerce experts to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core business domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of business expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Business & Commerce Domains Covered — Logistics / Supply Chain Management, Operations Research Analysis, Sustainability, Management Analysis, Product Management and Strategy, Digital Marketing. — Key Responsibilities • Author original business questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise reasoning in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD, DBA, or doctoral candidate in Business Administration, Management, Marketing, or a closely related field • MBA or Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level business strategy, organizational theory, and quantitative methods • Industry leadership experience or research publications in business fields is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$77 - $98 / hourOpen / Referral verified
Finance / Remote

Patient Financial Services Leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Patient Collections and Patient Financial Services leaders to evaluate AI tools designed to improve self-pay revenue recovery and patient financial engagement. Your expertise in self-pay collections, payment plan administration, and patient financial advocacy will directly inform AI systems that optimise patient payment experiences while maximising revenue recovery. — Responsibilities • Lead patient collections and self-pay operations including early-out collections, bad debt management, and patient payment plan administration. • Evaluate AI-generated patient financial communication drafts, payment plan recommendations, and self-pay resolution strategies for accuracy and compliance. • Develop and implement self-pay collection strategies across the revenue cycle, including pre-service, point-of-service, and post-service collections. • Manage patient payment plan enrollment, monitoring, and compliance processes. • Coordinate with financial counselling, billing, and bad debt recovery teams to optimise self-pay revenue capture. • Monitor self-pay KPIs including self-pay collection rates, payment plan conversion rates, bad debt write-off rates, and patient satisfaction scores. • Ensure compliance with FDCPA, HIPAA, state collections laws, and internal patient financial assistance policies. • Oversee relationships with collection agencies and early-out vendors as applicable. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in patient collections, self-pay revenue cycle, or patient financial services, with at least 2 years in a leadership role. • Deep knowledge of self-pay collection workflows, FDCPA compliance, and patient financial engagement best practices. • Experience managing early-out and bad debt collection programs, including vendor oversight. • Familiarity with propensity-to-pay tools and patient payment technology platforms. • Proficiency with EHR systems and patient collections/billing platforms. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to evaluate collections strategies and identify issues in AI-generated patient communication content. — Preferred Qualifications • CRCR, CHAM, or similar revenue cycle certification. • Experience with digital patient payment platforms and text/email-based collections outreach. • Background in hospital, health system, or physician group patient financial services. • Familiarity with AI tools and comfort evaluating AI-generated patient financial communication content. • Experience implementing propensity-to-pay analytics to prioritise collection efforts. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in patient financial services and the revenue cycle. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$92 / hourOpen / Referral verified
Code / Remote

Frontend Engineer — Web Replication Preference Rater

About the work — We're building a high-quality dataset of human preference judgments on AI-generated frontend code. You'll be shown a reference web page alongside two candidate replications produced by AI models, and you'll decide which replication is better — then explain why in writing that a model can learn from. — This is evaluation work, not authoring. You won't be building sites from scratch. You'll be reading someone else's HTML and CSS, running it locally, comparing it pixel-by-pixel against a target, and articulating exactly where and why it falls short. — What you'll do • Render a reference page and two candidate replications side by side at desktop width and judge which is the closer reproduction. • Diff layout fidelity in detail: box model and spacing, typography (family, size, weight, line-height, letter-spacing), color and border treatment, image and asset handling, z-order and overflow. • Inspect the underlying markup with browser devtools to distinguish a replication that is genuinely correct from one that merely looks correct at one viewport — hardcoded pixel offsets, absolute positioning standing in for real layout, and inline styles that will not survive a resize. • Evaluate responsive behavior and semantic quality: whether flexbox and grid are used where they belong, whether legacy float or table layouts in the reference were reproduced faithfully, whether headings and landmarks carry real semantic meaning. • Write a structured rationale for every judgment — the specific defects you found, ranked by how much they matter, in language precise enough to be actionable. • Flag ties, ambiguous cases, and broken task items rather than forcing a preference. — You're a fit if you have • 3+ years of professional web development experience, primarily in frontend or full-stack work. • Fluency in hand-written HTML and CSS: semantic markup, flexbox, grid, media queries, and older float- and table-based layouts you can still read and reason about. • Working command of browser devtools — element inspection, computed styles, the box model, and the network panel. • Enough JavaScript to read a page's scripts and understand what they do to the DOM, even if you don't write JS daily. • Comfort in a terminal: cloning a folder and serving it over a local static server without help. • Strong written English and the discipline to justify a judgment rather than assert it. — Equipment • A desktop or laptop with a browser window that opens to at least 1920px wide. • Administrator rights on your own machine, so you can install and run a local server. — Nice to have • Prior RLHF, preference labeling, or model evaluation work. • A code review or technical assessment background. • Pixel-perfect design-to-code experience — translating Figma or PSD comps into production markup. • Web accessibility expertise (WCAG, ARIA, screen reader testing). • Familiarity with how LLMs typically fail at codegen. • Web scraping or DOM parsing experience. — Note: this seat is for practicing web developers. Backend-only, mobile-native-only, data science, DevOps, and design-without-code backgrounds are out of scope for this project.

$90 / hourOpen / Referral verified
Business / US Remote

Enterprise Sales Domain Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is teaching its AI to do real enterprise sales work — the kind of multi-step, judgment-heavy workflows that only someone who has actually carried a quota knows how to get right. To do that well, they need seasoned sellers to act as the ground truth: people who can show the model what excellent looks like, catch where it goes wrong, and set the bar its work is measured against. — That is where you come in. As a Sales Domain Expert, you will bring years of real selling experience to auditing sales workflows, building the "golden" reference examples the model learns from, and shaping the rubrics used to evaluate it. This is a role for accomplished enterprise sellers — not general data annotators. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States. — 2. What You'll Do • Set the standard for "good": Review multi-step enterprise sales workflows and judge whether the AI handled them the way a strong seller would, flagging what is off and why. • Build golden trajectories: Work real sales tasks end-to-end — prospecting, pipeline management, CRM hygiene, reporting, and customer-facing documents — to create the expert reference examples the model is trained on. • Shape the evaluation: Pressure-test and refine the rubrics used to score AI outputs so they capture what actually separates great selling from average. • Work in realistic sandboxes: Operate inside test environments that mirror a real sales tech stack, completing and simulating workflows across CRM, communication, and billing tools. • Produce senior-level artifacts: Create polished, customer-facing materials — pitches, slide decks, and summaries — at the quality a senior sales leader would put their name on. — 3. Who We're Looking For • Around 10 years of hands-on experience selling in enterprise environments, with a real feel for the full sales cycle. • A consistent track record of meeting or exceeding your sales quota, which you can speak to in specifics. • Sellers at every level are welcome — quota-carrying individual contributors, managers who lead selling teams, and VPs running regional or global sales organizations. • Day-to-day fluency with the standard enterprise stack (Salesforce, Jira, Confluence, Google Workspace, and Microsoft Office), and the comfort to pick up tools like Slack, HubSpot, Notion, Aircall, Stripe, Intercom, Zoho, and Microsoft Teams. • Genuine depth in CRM administration, sales process management, pipeline tracking, and data-driven reporting. • Strong writing and a good eye for polish — you can turn a complex situation into a clear, professional, customer-facing artifact. — Nice to Have • Experience working within or alongside a two-sided, AI-powered B2B SaaS creator marketplace — for example, running outcome-based campaigns against contracted performance targets (tracking views, engagement, installs, and sign-ups in real time) and troubleshooting across a connected stack such as Zoho CRM, Jira, Confluence, Freshdesk, Microsoft Teams, and Google Workspace. If this sounds like you, tell us — we are especially keen to talk. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / Remote

Computational Statistics and Applied Mathematics Expert (R, Python, and Matlab/Scilab)

Computational Statistics and Applied Mathematics Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized statistical, mathematical, or scientific software packages. Some will ask the AI to compute reproducible numerical answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We welcome statisticians and applied mathematicians working across a wide range of specializations. You do not need experience with every package listed below; strong expertise with one or more specialized computational packages is sufficient. — We're especially interested in experts with deep, hands-on experience using one or more specialized R or Python packages, including examples such as: • Bayesian statistics: rstan, cmdstanr, rjags, runjags, brms, rstanarm, nimble, bayesplot, posterior, loo • Item response theory and psychometrics: TAM, sirt, mirt, mirtCAT, eRm, ltm, lordif, psych • Structural equation and latent variable modelling: lavaan, semTools, OpenMx • Topological data analysis: TDAstats, TDApplied • Differential equations and dynamical systems: deSolve, pomp, FME • State-space and time-series modelling: KFAS, MARSS, forecast, vars, urca, rugarch, rmgarch, tseries, timeSeries • Survival and event-history analysis: survival, flexsurv, timereg, mets • Mixed, additive, and advanced regression models: lme4, nlme, mgcv, glmmTMB, TMB, quantreg, scam • Spatial statistics and geostatistics: spatstat, spatstat.geom, spatstat.linnet, spdep, gstat, geoR, spBayes, sf, stars, terra, lwgeom • Statistical learning and specialized modelling: mclust, kernlab, earth, pROC, multcomp, sandwich, effectsize, irr • Optimization and mathematical programming: lpSolve, linprog, nloptr, DEoptimR, SQUAREM • Numerical linear algebra and high-precision computation: RSpectra, Rmpfr, gmp, pracma • Computational geometry: geometry, deldir, polyclip — Other similar specialized statistical, mathematical, scientific, or domain-specific R packages will also be considered. Other similar specialized statistical or mathematical Python/Scilab packages are also welcome, such as statsmodels and PyMC. — Numerical computing and scientific modelling in Matlab/Scilab are also wanted. — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD required; PhD preferred, or MS with 10+ years of relevant experience) in statistics, applied mathematics, or a closely related quantitative field, with real hands-on experience using specialized computational packages — not just theoretical knowledge. — You have written code using one or more specialized statistical, mathematical, or scientific packages to solve actual research or professional problems, and you understand where these tools break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. Deep expertise with one or more specialized computational packages is more important than familiarity with the entire package list above. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in statistics, applied mathematics, a relevant STEM field, or equivalent research experience • Proven proficiency with at least one specialized statistical, mathematical, or scientific software package, demonstrated through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple computational domains or specialized software packages • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $90 / hourOpen / Referral verified
Medical / Playa Vista, CA or New York, NY Remote

Character and Facial Animation Consultant

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building a performance transfer model: one that takes an actor's original performance and carries it faithfully into a new output, holding on to the timing, the emotion, and the small physical choices that make it feel real. The team believes AI should honor the craft of performance, not flatten it. — To get the evaluation right, we are bringing in senior character animators to help define what "good" looks like, critique the current evaluation approach, and shape the pool of reviewers who score model outputs. — This is a part-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Establish evaluation standards: Define clear criteria for when a performance has been captured well, both in emotional resonance and physical fidelity, and point out where current evaluations fall short. • Curate benchmarks: Identify strong examples of performances that are easy to capture and ones that are genuinely hard, so the model has meaningful tests to measure against. • Shape the evaluation pool: Help define the right profile, expertise mix, and rubric for the human reviewers who assess outputs at scale. • Collaborate with other experts: Work alongside fellow subject-matter experts to keep evaluations consistent and accurate. — 3. Core Qualifications • 4+ years as a hands-on character animator, with human or character _performance_ as your core craft. • Verifiable credits on 2 or more feature films, animated features, or AAA game cinematics at recognized studios. • Direct experience with facial animation and/or performance capture — facial keyframe work, motion-capture solving, cleanup or motion editing, or character technical design of facial and performance rigs. • A meticulous eye for micro-expressions, emotional beats, and the subtleties that make a human performance feel authentic. • AI proficiency: familiarity with AI tools and workflows, including the ability to generate your own visual samples for testing and comparison. • The ability to translate complex visual details into clear, precise language that can be captured in data captions and evaluation templates. • Able to come on-site in Playa Vista (Los Angeles), CA or New York, NY for in-person working sessions once or twice a month. • (Optional) Reliable access to a 4K-resolution monitor for precise, pixel-level review. — 4. Who this role is not for — This role is specifically for animators whose primary craft is character performance. It is not a fit for directors of photography, lighting or camera specialists, colorists, compositors, editors, or generalist 3D artists. Layout, previsualization, environment and effects work will not qualify on their own. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / United States Remote

LLM Red Team Specialist — Failure Modes & Edge Cases

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs specialists who can find where those models break. Working in a red-teaming setup, you will design and probe complex, multi-step tasks to expose vulnerabilities, edge cases, and failure modes in frontier AI systems — the places where a model looks competent but is quietly wrong. — Each task represents one to two days of continuous, focused effort and spans multiple technical skills: coding, experimentation, and careful analysis. You will work in a tight feedback loop with the lab's researchers, turning the failure modes you find into stronger benchmark tasks. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Probe models: Explore how frontier AI models behave on coding, ML, and analysis tasks, and find the spots where they quietly get things wrong. • Design challenges: Turn the weaknesses you find into well-crafted tasks that are hard for models but fair to grade. • Document findings: Write up what you discover clearly, with evidence and steps others can reproduce. • Strengthen tasks: Team up with task authors to close loopholes, shortcuts, and grading gaps. • Work as a team: Share insights with researchers and fellow experts so the benchmark keeps getting better. — 3. Core Qualifications • MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding. • 1+ years of experience in a research, research-engineering, security, or AI-evaluation role. • Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems — through red teaming, adversarial testing, security research, or rigorous model evaluation. • Working proficiency in Python and Git, with the ability to script your own probes and analyses. • Strong familiarity with LLM capabilities, limitations, and evaluation techniques. • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in finding what others missed, strong written communication, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / United States Remote

QA/Test Engineer

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models, and complex multi-step tasks are only useful if they are airtight: unambiguous, correctly graded, and robust to shortcuts. We are seeking experienced QA and test engineers to be the quality backbone of this benchmark — designing the test cases and review processes that guarantee every task measures what it claims to measure. — Each task under review represents one to two days of expert effort and spans multiple technical skills, so quality review here means genuinely understanding the task: running it, probing its edge cases, and debugging its environment. You will work in a tight feedback loop with the lab's researchers and task authors. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design checks: Create test cases that confirm each task works as intended — including the tricky edge cases. • Review tasks: Give tasks and reference solutions a careful read before they're finalized, catching ambiguity and gaps early. • Debug: Roll up your sleeves in Python when a task or its checks don't behave the way they should. • Shape the process: Help build simple, repeatable quality checklists, and share feedback authors can act on right away. • Protect the results: Watch for shortcuts and grading gaps in AI agent runs so benchmark scores stay trustworthy. — 3. Core Qualifications • MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain. • 1+ years of experience in test engineering, quality assurance, or a research/software engineering role with strong quality ownership. • Demonstrated skill designing test cases and quality-review processes, and debugging complex systems end-to-end. • Working proficiency in Python and Git, and comfort navigating unfamiliar codebases and environments. • Exceptional attention to detail and clear written documentation habits. • Past experience in AI training, model evaluation, or quality review of AI-generated work is preferred. • A perfectionist mindset: creativity in finding what others missed, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Medical / United States Remote

STEM Researcher — Computational Fields

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and is recruiting researchers from computational STEM fields — as well as computationally heavy social sciences and humanities — to bring working-researcher rigor to benchmark design. You will translate the scientific method — experimental design, hypothesis testing, and rigorous evaluation — into complex, multi-step tasks that today's best models cannot yet complete reliably. — Each task represents one to two days of continuous, focused effort and spans multiple skills: study design, implementation in code, data analysis, and careful written conclusions. You will work in a tight feedback loop with the lab's researchers, surfacing the kinds of methodological mistakes a working researcher would catch immediately. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design tasks: Turn the research skills you use every day — designing studies, testing hypotheses, evaluating results — into engaging, multi-step tasks. • Author solutions: Work through your own tasks in Python and notebooks, at the level of rigor you'd expect from a careful colleague. • Define what good looks like: Help spell out what separates sound scientific reasoning from reasoning that merely sounds right. • Evaluate models: Review model attempts at your tasks and flag the mistakes a working researcher would spot right away. • Work as a team: Compare notes with researchers and fellow experts to keep evaluations consistent and accurate. — 3. Core Qualifications • MSc or PhD in a STEM field, or in a computational social-science or humanities discipline, or equivalent practical experience in a research-heavy domain requiring data analysis and coding. • 1+ years of experience in an active research role (academia, industry, or national labs). • Your own research involves significant computational work: Python-based analysis, simulation, modeling, or data pipelines. • Strong grounding in experimental design, hypothesis testing, and rigorous evaluation of results. • Working familiarity with Git, IDEs, and notebook environments (Jupyter or Colab). • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Math / United States Remote

Data Science & Quantitative Analysis Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced data scientists and quantitative analysts to act as ground-truth experts. You will design complex analysis tasks that simulate real research work — for example, comparing two anomaly-detection algorithms on a dataset, calculating correlations, performing manual spot checks, and summarizing the findings in a notebook clear enough to drive a researcher's decision. — Each task represents one to two days of continuous, focused effort and spans multiple skills: data cleaning, statistical analysis, interpretation, and clear reporting. You will work in a tight feedback loop with the lab's researchers, verifying exactly where and why frontier models fall short on rigorous analytical work. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design tasks: Create realistic data-analysis challenges — cleaning messy data, comparing methods, interpreting results — based on the kind of analysis you do every day. • Author notebooks: Work through your own tasks in Jupyter or Colab, producing clear, reproducible reference analyses. • Compare methods: Build tasks that ask for a fair comparison between analytical approaches, backed by spot checks and a clear recommendation. • Evaluate models: Review how models handle your tasks, and check whether their statistics and conclusions actually hold up. • Work as a team: Compare notes with researchers and fellow experts to keep evaluations consistent and accurate. — 3. Core Qualifications • MSc or PhD in statistics, data science, or another quantitative STEM field, or equivalent practical experience in a research-heavy analytical domain. • 1+ years of experience in a research, research-engineering, or heavy data-analysis role. • Deep hands-on data-analysis skills: data cleaning, statistical correlation, hypothesis testing, and careful interpretation of results. • Proficiency with Jupyter Notebooks or Google Colab for analysis and reporting. • Working proficiency in Python (pandas, NumPy, or similar) and Git. • Strong ability to communicate analytical findings in writing for decision-makers. • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in task design, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / United States Remote

Machine Learning Engineer — Model Evaluation & Experimentation

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation. You will author complex, multi-step ML tasks — for example, taking a vague research idea like "modify how an RL reward is computed," implementing the change, running the training experiment, and analyzing the results to determine success — and verify exactly where frontier models fall short. — Each task represents one to two days of continuous, focused effort and spans multiple technical skills: implementation, experiment setup and execution, and rigorous analysis. You will work in a tight feedback loop with the lab's researchers. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design tasks: Turn real ML research ideas — like tweaking how an RL reward is computed — into well-defined, multi-step tasks. • Run experiments: Implement changes, run training experiments, and analyze the results to show what a correct solution looks like. • Explore RL ideas: Build some of your tasks around reinforcement-learning basics such as reward functions and training behavior. • Evaluate models: See how frontier models handle your tasks, and note where and why they fall short. • Work as a team: Compare notes with researchers and fellow experts so tasks stay consistent, rigorous, and fair. — 3. Core Qualifications • MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy domain. • 1+ years of experience in a research or research-engineering role. • Hands-on experience training and evaluating ML models and running experiments end-to-end — experiment setup, execution, and analysis — not just using ML libraries superficially. • Strong familiarity with large language models: their capabilities, limitations, and evaluation techniques. • Working proficiency in Python and Git, with comfort in both scripting and notebook environments. • Basic understanding of reinforcement learning (reward functions, policy training) is preferred. • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / United States Remote

Software Engineering Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced software engineers to act as task curators. You will design, implement, and review complex, multi-step engineering tasks that simulate the real-world challenges research engineers face — realistic, genuinely hard problems that today's best AI coding agents cannot yet solve reliably. — Each task represents one to two days of continuous, focused effort and spans multiple technical skills: Python implementation, environment and tooling setup, debugging, and clear documentation. You will work in a tight feedback loop with the lab's researchers, verifying exactly where and why frontier models fail on your tasks. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design tasks: Create realistic, multi-step software engineering challenges — the kind of work you'd actually do on the job — that push the limits of today's best AI coding agents. • Build reference solutions: Solve your own tasks in Python, with the setup and checks needed so each task has a clear, verifiable answer. • Work with AI tools: Use AI coding assistants as part of your everyday workflow, and observe where they help and where they fall short. • Review and refine: Look over tasks built by fellow experts and share feedback on clarity, correctness, and difficulty. • Learn from failures: See how AI agents attempted your tasks and help the research team understand what tripped them up. — 3. Core Qualifications • MSc or PhD in computer science or another STEM field, or equivalent practical experience in a research-heavy domain requiring significant coding and data analysis. • 1+ years of experience in a research, research-engineering, or software engineering role. • Strong hands-on Python scripting and debugging skills, with clean-code habits and attention to readability. • Everyday fluency with version control (Git), IDEs, and standard software development workflows. • Experience with AI coding assistants, prompt engineering, or agent workflows is preferred. • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
STEM / US Remote

Architecture Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced architects to help evaluate and improve how AI systems understand and reason about architecture topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in architecture contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of architectural expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how architects actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Degree-qualified architect holding a professional architecture license in the US. • 2+ years of professional architecture experience, ideally across more than one project type (residential, commercial, institutional). • Familiarity with building codes and standards (e.g. IBC) and common tools (Revit, AutoCAD, BIM, or similar). • A portfolio of built or realized work — not purely academic or conceptual projects — is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Code / US Remote

Civil Engineering Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced civil engineers to help evaluate and improve how AI systems understand and reason about civil engineering topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in civil engineering contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of civil engineering expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how civil engineers actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Degree-qualified civil engineer holding a Professional Engineer (PE) license in the US. • 2+ years of professional civil engineering experience, ideally across more than one area (structural, geotechnical, transportation, water resources). • Familiarity with relevant codes and standards (e.g. ASCE, IBC, AASHTO) and common engineering tools (AutoCAD Civil 3D, STAAD, Revit, or similar). • Some experience writing technical specs, design reports, or training material is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Code / US Remote

Chemical Engineering Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced chemical engineers to help evaluate and improve how AI systems understand and reason about chemical engineering topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in chemical engineering contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of chemical engineering expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how chemical engineers actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Degree-qualified chemical engineer holding a Professional Engineer (PE) license in the US. • 2+ years of professional chemical engineering experience, ideally across more than one area (process design, process safety, plant operations). • Familiarity with process safety standards (e.g. OSHA PSM, API) and process simulation tools (Aspen, MATLAB, or similar). • Some experience writing process documentation, SOPs, or technical training material is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Code / US Remote

Mechanical Engineering Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced mechanical engineers to help evaluate and improve how AI systems understand and reason about mechanical engineering topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in mechanical engineering contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of mechanical engineering expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how mechanical engineers actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Degree-qualified mechanical engineer holding a Professional Engineer (PE) license in the US. • 2+ years of professional mechanical engineering experience, ideally across more than one area (thermodynamics/HVAC, structural/mechanical design, manufacturing). • Familiarity with relevant codes and standards (e.g. ASME) and common engineering tools (SolidWorks, ANSYS, AutoCAD, or similar). • Some experience writing technical documentation, specs, or training material is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Code / US Remote

Electrical Engineering Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced electrical engineers to help evaluate and improve how AI systems understand and reason about electrical engineering topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in electrical engineering contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of electrical engineering expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how electrical engineers actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Degree-qualified electrical engineer holding a Professional Engineer (PE) license in the US. • 2+ years of professional electrical engineering experience, ideally across more than one area (power systems, controls, electronics, signal processing). • Familiarity with relevant codes and standards (e.g. NEC, IEEE) and common engineering tools (MATLAB, AutoCAD Electrical, SPICE, or similar). • Some experience writing technical documentation, specs, or training material is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Finance / United States Remote

Finance Specialist

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Finance subject-matter experts (SMEs) with deep domain expertise and hands-on experience evaluating AI model outputs against rubrics to bring rigor and real-world financial judgment to our AI training data. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps in financial reasoning, analysis, and decision-making. • Design challenging, domain-relevant finance tasks and write accurate, well-reasoned solutions grounded in real financial practice. • Evaluate AI model outputs against structured rubrics and provide clear, written feedback on correctness, judgment, and reasoning quality. • Develop and refine evaluation guidelines and scoring rubrics specific to finance tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 8+ years of dedicated professional experience in finance (e.g., investment banking, asset management, corporate finance, financial advisory) at a recognized, top-tier organization (e.g., Goldman Sachs, JPMorgan, Morgan Stanley, BlackRock, Fidelity, Deloitte, PwC, EY, KPMG, or equivalent). • Prior hands-on experience evaluating LLM/AI model outputs against rubrics or structured scoring criteria — mandatory; please describe this experience in your application. • Demonstrable career progression (e.g., Analyst → Associate → VP/Director). • Ability to engage reliably for at least 35 hours/week during weekdays. • Verbal and written communication skills, problem-solving skills, and interpersonal skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Multimodal / Remote, United States (PST to EST hours)

Visual Quality Expert (Film, VFX & Animation)

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking senior image quality experts from film, VFX, animation, and professional photography — VFX and rendering supervisors, colorists, cinematographers, lighting artists, and high-end photographers — with a strong foundation in visual perception, light physics, and pixel-level scrutiny. You will judge, frame by frame, whether an enhanced or upscaled image actually holds up, catching the compression artifacts, noise, aliasing, grain-structure inconsistency, and over- or under-sharpening that typically go unnoticed by the average viewer. — This is a part-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Evaluate high-resolution video and stills at pixel level, identifying compression artifacts, noise, aliasing, banding, grain-structure inconsistency, and over- or under-sharpening. • Apply an uncompromising focus on the intricate details that typically go unnoticed, and flag where output diverges from professional image-quality standards. • Design and refine evaluation templates and rating guidelines so reviewers apply the same visual-quality bar consistently, including under edge cases. • Translate complex visual details into clear, precise descriptive language that can be captured in data captions and evaluation templates. • Generate your own visual samples for targeted testing and side-by-side comparison. • Collaborate with other subject matter experts and program leads to maintain visual consistency across sequences and across reviewers. — 3. Core Qualifications • 5+ years of professional experience in high-end visual evaluation across film, VFX, animation, or professional photography. • Direct professional experience in at least one of the following disciplines: • Visual Effects (VFX) Supervisor — the authority on final image quality, with an elite eye for pixel-level flaws, grain-structure inconsistencies, and visual artifacts. • Rendering Supervisor — accustomed to running dailies and taking ultimate responsibility for visual quality down to the absolute pixel. • Lighting Supervisor, Lead, or Artist — a specialist in illumination and the foundational building blocks of digital video and cinematic motion. • Colorist — deep expertise in color grading, artifact detection, and maintaining visual consistency across sequences. • Director of Photography (DP) or Cinematographer — skilled at framing, capturing visual performance, and critically evaluating over- and under-sharpness in high-resolution images. • Effects or Surfacing Artist — strong command of textures, materials, and the complex simulation of visual elements. • High-End Photographer — deep command of lighting dynamics, depth of field, and color artifacts in high-resolution captures. • Director — particularly those with deep experience directing crowd and background action. • Digital Projectionist — a rigorously trained eye for final-output quality, screen-level fidelity, and visual anomalies. • Training from a university or program with a highly respected film, visual effects, or digital media program. • Reliable access to a 4K-resolution monitor or display to perform precise, pixel-level evaluation. • Familiarity with AI tools and workflows, including the ability to generate your own visual samples for testing and comparison. • Ability to engage reliably for 20 hours per week, with working hours overlapping the PST to EST window. • Fluent in English, with the exceptional ability to translate complex visual details into clear, precise descriptive language. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$60 - $90 / hourOpen / Referral verified
Code / North America (US & Canada) Remote

CAD Engineer — ScreenSpot Plus (Screenshot Capture & UI Annotation)

About the Role — You've been selected for ScreenSpot Plus, where you'll capture screenshots of professional CAD software in realistic, expert use, annotate interactive UI elements, and write natural-language task instructions. — The project begins with a pilot phase, during which your initial submissions will be closely reviewed for quality before you ramp into full production. Detailed onboarding materials and capture guidelines will be shared once you accept. — What You'll Do • Capture high-quality screenshots of professional CAD software during realistic, expert workflows • Annotate interactive UI elements (buttons, menus, panels, toolbars, dialogs) accurately and consistently • Write clear, natural-language task instructions that reflect how an expert actually uses the software • Iterate on feedback during the pilot phase to meet quality standards before scaling to full production — Who We're Looking For • Hands-on, professional experience with CAD software (e.g., AutoCAD, SolidWorks, Fusion 360, CATIA, Revit, Siemens NX, or similar) • Strong working knowledge of the day-to-day workflows and interface of your CAD tools • Attention to detail and the ability to follow precise annotation and capture guidelines • Clear written English for task instructions • Reliable access to the relevant CAD software for capturing screenshots

$70 - $90 / hourOpen / Referral verified
Medical / Remote

Medical Revenue Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Charge Capture, Charge Integrity, and Revenue Integrity professionals to evaluate AI tools designed to identify revenue leakage, improve charge accuracy, and ensure coding and billing compliance. Your expertise in charge-capture workflows, CDM management, and revenue-integrity analysis will help us build AI systems that optimise revenue performance and reduce compliance risk. — Responsibilities • Oversee charge capture, charge integrity, and revenue integrity functions to ensure accurate and compliant charge submission. • Evaluate AI-generated charge review outputs, coding recommendations, and revenue integrity alerts for accuracy and compliance. • Conduct charge audits to identify missed charges, duplicate charges, and charge capture errors across clinical departments. • Manage and maintain the charge description master (CDM), ensuring accuracy of charge codes, revenue codes, and pricing. • Analyse charge patterns to identify revenue leakage opportunities and implement corrective action plans. • Collaborate with clinical, coding, and billing teams to resolve charge capture discrepancies. • Monitor KPIs including charge lag times, charge capture accuracy rates, and revenue integrity findings. • Ensure compliance with CMS billing guidelines, OIG work plan priorities, and payer-specific requirements. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in charge capture, charge integrity, revenue integrity, or healthcare compliance. • Deep knowledge of CDM management, revenue codes, charge capture workflows, and billing compliance. • Strong understanding of CMS billing guidelines, Medicare Part A/B billing rules, and OIG compliance requirements. • Proficiency with charge capture systems, EHR platforms, and revenue integrity tools. • Experience conducting charge audits and developing corrective action plans. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify billing errors and compliance risks in AI-generated outputs. — Preferred Qualifications • CPC, CCS, CHFP, or CRCS credential. • Experience with 340B drug program charge capture and compliance. • Background in hospital, health system, or physician group revenue integrity operations. • Familiarity with AI tools and comfort evaluating AI-generated charge and billing content. • Experience with revenue integrity software platforms and analytics tools. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in healthcare revenue integrity. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$88 / hourOpen / Referral verified
STEM / Remote

Biology Research Scientist (BA, MS, PhD's)

About the Role — Mercor is partnering with a leading AI research organization to verify protein target assignments in large-scale bioactivity databases (ChEMBL, BindingDB). You will read primary literature, apply scientific judgment, and determine whether UniProt IDs accurately reflect the proteins being studied — work that directly feeds AI drug discovery pipelines. — What You'll Do • Access primary sources (papers, patents) to verify protein target assignments against UniProt records • Flag and classify target assignment errors using a structured taxonomy • Propose correct UniProt accessions where assignments are wrong • Write concise, evidence-grounded notes explaining your reasoning — Requirements • BA/BS with 5+ years, MS with 2+ years, or PhD with industry or drug discovery research experience, in pharmacology, biochemistry, molecular biology, or chemical biology at a biotech, pharma, or CRO • Hands-on binding or functional assay experience (SPR, TR-FRET, radioligand binding, kinase assays, GPCR functional assays, IC50/Ki/KD) • Currently bench-active in a research, scientist, or associate scientist role • Working fluency with UniProt or adjacent workflows: SAR support, HTS, target validation, biochemical profiling, or IND-enabling studies — Nice to Have • Direct experience with ChEMBL, BindingDB, or PubChem • Selectivity profiling or counterscreening experience • Familiarity with agonist/antagonist vs. activator/inhibitor distinctions — Role Details • 10–20 hrs/week | Remote | Immediate start • 1–2 month minimum, extension likely • U.S. only

$50 - $70 / hourOpen / Referral verified
STEM / Remote

Generalist Expert

Overview — In this role, you will evaluate AI-generated responses and provide structured written feedback. This is a great opportunity for sharp, analytical thinkers to contribute to high-impact AI research projects. — Basic Qualifications • Bachelor's degree from a top-500 globally ranked university preferred • Strong analytical and written communication skills • Ability to work independently and follow detailed task guidelines — Required Skills • Strong critical reading skills with the ability to identify nuance, implicit meaning, and gaps in reasoning • Ability to write clear, precise, and well-evidenced written rationales that go beyond surface-level observations • Consistent and honest judgment, including the ability to give critical assessments when warranted • Strict attention to detail and accurate application of structured evaluation guidelines • Ability to work entirely without AI writing tools — Eligibility • Native English fluency required

$70 / hourOpen / Referral verified
Code / Remote

Computational Electrical Engineering & RF/Circuit Design Expert

Electrical Engineering & RF/Circuit Design Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience in: — Electrical Engineering & RF/Circuit Design — working with scikit-rf for RF and microwave network analysis, S-parameter characterization, and transmission-line modeling, or ngspice for circuit simulation, operating point analysis, and frequency response characterization. Candidates should be comfortable designing problems that involve recovering circuit parameters from measurement data. — _Experience with other specialized software in this domain will also be considered._ — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $85 / hourOpen / Referral verified
Code / Remote

Computational Structural & Mechanical Engineering Expert

Computational Structural & Mechanical Engineering Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard computational scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience with open-source, domain-specific computational tools such as FEniCSx/DOLFINx, scikit-fem, OpenFOAM, deal.II, MFEM, MOOSE, CalculiX, Elmer FEM, Code\_Aster, SfePy, FiPy, Devito, Cantera, CoolProp, Pyomo, or SimPy, for finite-element analysis, computational mechanics, structural analysis, elasticity, CFD, multiphysics simulation, heat and mass transfer, thermodynamics, combustion, fluid mechanics, HVAC/thermal systems, manufacturing simulation, optimization, or thermophysical-property calculations. — Relevant work may include beam, plate, and shell analysis; linear or nonlinear elasticity; finite-element and variational formulations; mesh refinement and convergence studies; continuum and solid mechanics; computational fluid dynamics; coupled multiphysics problems; thermal-fluid simulation; structural or system optimization; reliability analysis; and related numerical engineering workflows. — Experience with underlying theories and numerical methods — such as Euler–Bernoulli and Timoshenko beam theory, continuum mechanics, finite-element methods, Galerkin/variational methods, finite-volume methods, PDE discretization, constitutive modeling, thermodynamics, numerical linear algebra, and nonlinear solution methods — is valuable. — Experience with other open-source computational structural or mechanical engineering software will also be considered, including scientific codes and solver frameworks built with Python, C, C++, or Fortran. — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $85 / hourOpen / Referral verified
STEM / Remote

Computational Seismology & Geophysics Expert

Computational Seismology & Geophysics Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience in: — Computational Seismology & Geophysics — hands-on experience with open-source, domain-specific computational tools such as SPECFEM, ObsPy, Pyrocko, SimPEG, pyGIMLi, SeisBench, EQcorrscan, or Fatiando a Terra, for seismic wave propagation and numerical simulation, synthetic seismogram generation, full-waveform inversion (FWI), seismic imaging, travel-time tomography, moment tensor inversion, event detection/location, or related computational geophysics workflows. — Experience with other open-source computational seismology or geophysics software will also be considered, including tools and scientific codes built with Python, C, C++, or Fortran. — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $85 / hourOpen / Referral verified
STEM / Remote

Computational Pharmacokinetics & Systems Biology Expert

Computational Pharmacokinetics & Systems Biology Expert — About the Project — We're building a large-scale benchmark to test how well advanced AI systems can solve hard scientific and engineering problems. As a task designer, you'll create challenging computational problems that check whether AI can use real scientific software to do research-level work — running simulations, interpreting results, designing experiments, and uncovering hidden information from data. — This isn't a typical data-labeling job. You'll design original, graduate-level problems based on real scientific workflows, test them against cutting-edge AI models, and fine-tune them until the difficulty is just right. — What You'll Do — You'll create problems that require skilled use of specialized scientific software. Some will ask the AI to compute exact answers from a fully defined setup — testing whether it can correctly carry out complex, multi-step workflows. Others will be harder: the AI must plan a series of queries or experiments to uncover information that isn't directly visible, which means thinking strategically about what to measure, how to read partial results, and how to narrow down the possibilities efficiently. — Each problem goes through a testing loop against state-of-the-art AI models, and you'll refine it until it hits the target difficulty. — Domains & Tools We're Hiring For — We're especially interested in experts with deep, hands-on experience in: — Pharmacokinetics & Systems Biology — working with libRoadRunner, Tellurium, or SBML-based tools for compartmental PK/PD modeling, enzyme kinetics, or systems biology simulations. — _Experience with other specialized software in this domain will also be considered._ — What Makes a Strong Candidate — You have graduate-level expertise (MS or PhD preferred) in the domain above, with real hands-on experience using these tools — not just theoretical knowledge. You've written code using these libraries to solve actual research problems, and you understand where they break, what their edge cases are, and what makes a problem genuinely hard rather than just complicated. — Beyond domain expertise, the best candidates think like puzzle designers: building problems where the challenge comes from smart reasoning rather than raw computation, where several approaches seem plausible but only careful analysis reveals the right one, and where surface-level pattern matching won't get you to the answer. — Requirements • Graduate-level training in a relevant STEM field (MS, PhD, or equivalent research experience) • Proven proficiency with at least one of the listed scientific software libraries, shown through research publications, open-source contributions, or professional work • Strong Python skills — you'll be writing problem setups, oracle functions, and solution validators • Ability to work independently and refine problem designs based on feedback • Comfortable working in a Linux/terminal environment with remote compute sandboxes • Available for at least 15–20 hours per week — Nice to Have • Experience across multiple listed domains or tools • Familiarity with benchmark or evaluation design • Background in scientific teaching or exam/problem-set design • Experience with computational reproducibility and containerized environments

$70 - $85 / hourOpen / Referral verified
Finance / Remote

Payment-posting & Reconciliation Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Cash Posting and Payment Reconciliation Managers to evaluate AI tools designed to automate electronic remittance processing, payment posting, and cash reconciliation workflows. Your expertise in ERA processing, EOB interpretation, and payment reconciliation will directly inform AI systems that improve cash posting accuracy and accelerate revenue cycle close. — Responsibilities • Oversee cash posting and payment reconciliation operations, including electronic remittance advice (ERA/835) processing, manual EOB posting, and lockbox reconciliation. • Evaluate AI-generated payment posting outputs, ERA matching recommendations, and reconciliation reports for accuracy and completeness. • Manage electronic and manual payment posting workflows across multiple payers and payment types. • Reconcile posted payments against bank deposits, lockbox reports, and payer remittances to ensure accuracy. • Identify and resolve posting errors, misapplied payments, and unapplied cash. • Monitor cash posting KPIs, including posting accuracy rates, days to post, unapplied cash balances, and reconciliation variance. • Collaborate with billing, A/R, and finance teams to resolve payment discrepancies and ensure timely cash close. • Ensure compliance with internal controls, HIPAA, and audit requirements for cash handling. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in cash posting, payment reconciliation, or revenue cycle operations, with at least 2 years in a management role. • Deep knowledge of ERA/835 electronic remittance processing, EOB interpretation, and lockbox reconciliation. • Strong understanding of payer payment methodologies and remittance adjustment reason codes. • Experience managing high-volume payment posting operations across multiple payers. • Proficiency with billing systems and payment posting platforms (Epic, Athenahealth, or equivalent). • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify posting errors and discrepancies in AI-generated payment outputs. — Preferred Qualifications • CRCR, CPC, or CHFP certification. • Experience with automated ERA posting platforms and RPA-driven cash posting solutions. • Background in hospital or physician group cash posting operations with multi-payer complexity. • Familiarity with AI tools and comfort evaluating AI-generated remittance and payment content. • Experience developing cash posting SOPs and internal control frameworks. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in payment reconciliation and revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$85 / hourOpen / Referral verified
Finance / Remote

Underpayment & Managed-care Contract Specialist

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Underpayment and Managed Care Contract Specialists to evaluate AI tools designed to detect payment variances and automate contract compliance monitoring. Your expertise in payer contract interpretation, payment variance analysis, and underpayment recovery will help AI systems optimise net revenue and identify systematic underpayments. — Responsibilities • Identify, analyse, and recover underpayments and payment variances across commercial, Medicare Advantage, and Medicaid managed care payer contracts. • Evaluate AI-generated underpayment detection alerts, contract compliance outputs, and payment variance analyses for accuracy. • Interpret payer contract terms including fee schedules, carve-outs, outlier provisions, and payment methodologies to validate claim payments. • Conduct contract modelling and payment reconciliation to identify systematic underpayment patterns. • Develop and submit underpayment claims and recovery correspondence to payers. • Collaborate with managed care contracting teams to identify contract language gaps and renegotiation opportunities. • Monitor underpayment recovery KPIs including identified underpayment dollars, recovery rates, and payer response rates. • Ensure compliance with payer contract terms, HIPAA, and timely filing requirements for underpayment recovery. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in underpayment recovery, payment variance analysis, or managed care contracting, with demonstrated expertise in contract interpretation. • Deep knowledge of payer payment methodologies including DRG-based, per diem, per cent-of-billed charges, and fee schedule reimbursement. • Strong experience with contract modelling, payment reconciliation, and payer dispute resolution. • Familiarity with managed care contract management systems and payment variance detection tools. • Proficiency with Excel-based financial analysis and revenue cycle analytics platforms. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify payment discrepancies in complex contract structures and AI-generated outputs. — Preferred Qualifications • CHFP, CRCR, or managed care contracting certification. • Experience with contract management software (e.g., Experian Health, Recondo, or similar). • Background in hospital, health system, or large physician group underpayment recovery operations. • Familiarity with AI tools and comfort evaluating AI-generated payment variance content. • Experience negotiating payer contracts and presenting underpayment analysis to executive leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in managed care contracting and the revenue cycle. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$85 / hourOpen / Referral verified
Code / Remote

Applied Computer Science Benchmark Specialist

Role Overview — We are seeking expert computer scientists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core computer science domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of computer science expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Computer Science Domains Covered — Accelerator / GPU Kernel Engineering, Formal Methods & Automated Reasoning, Computer Architecture & Accelerators, Distributed Systems, DevOps & Site Reliability, Data Engineering & Databases, Cloud & Infrastructure, OS & Systems Kernel, Machine Learning Engineering, Web & API Development, Embedded Systems Engineering, Computer Graphics & Game Development, Mobile Engineering. — Key Responsibilities • Author original computer science questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Computer Science, Electrical Engineering, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level CS theory, algorithms, systems design, and/or machine learning • Research publications, industry experience at top tech companies, or competitive programming background is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$66 - $84 / hourOpen / Referral verified
Code / Remote

Atomic Layer Deposition (ALD) Experts

Mercor is seeking experts in Atomic Layer Deposition (ALD) and thin-film processes to support a frontier AI research lab building models for semiconductors and the physical sciences. This is hands-on, expert-level work: you'll apply deep, specialized knowledge to generate, structure, and evaluate the scientific data these models learn from — and your input will directly shape how advanced models reason about deposition, materials, and semiconductor processes. — Key Responsibilities: • Contribute domain expertise across ALD process development, precursor chemistry, and thin-film characterization to build high-quality training and evaluation data. • Review and evaluate AI-generated scientific reasoning, catching errors and improving technical accuracy. • Design and solve challenging, expert-level problems in ALD and semiconductor processing. • Rate and rank model outputs against defined scientific criteria, with clear written reasoning. • Structure technical knowledge — process parameters, recipes, characterization results — into well-organized, model-ready data. • Deliver reliable, high-quality work within defined timelines. — You're a strong fit if you have: • Hands-on experience developing, optimizing, or troubleshooting ALD processes. • Deep knowledge of thin-film deposition for semiconductor or advanced-packaging applications. • A strong grasp of precursor chemistry and surface reaction mechanisms. • Experience with materials characterization (XRD, SEM, TEM, XPS, ellipsometry, etc.). • An advanced degree (PhD/MS) or equivalent hands-on experience in materials science, chemistry, chemical engineering, or physics. • Clear written English and the ability to explain technical reasoning precisely. — Role Details: • Type: Long-term, ongoing engagement • Engagement: Up to 40 hours/week (minimum 10) • Work arrangement: Remote (US-based)

$84 / hourOpen / Referral verified
Policy & Safety / Remote

AI Safety Red Teamer

We are seeking experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous ("grey-area") topics. — Responsibilities • Design adversarial prompts to stress-test frontier AI models. • Identify jailbreaks, unsafe behaviours, hallucinations, and policy failures. • Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains. • Document vulnerabilities and contribute to safety benchmarking and red-teaming reports. • Collaborate with AI researchers to improve model alignment, robustness, and safety. — Required Qualifications • Bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related discipline. • 5+ years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a related field. • Strong analytical reasoning, prompt design, and written communication skills. • Experience designing adversarial prompts or evaluating frontier AI systems. — Preferred Qualifications • Experience with AI Red Teaming, RLHF, SFT, AI Alignment, or Trust & Safety. • Familiarity with jailbreak testing, prompt engineering, or adversarial evaluation methodologies. • Expertise in one or more grey-area domains, including cyber, biosecurity, political content, misinformation, or scientific safety. — Why Join? • Help secure and strengthen the next generation of frontier AI models. • Work on cutting-edge adversarial testing alongside leading AI researchers and safety teams. • Influence how AI systems respond to complex, real-world safety challenges.

$70 - $84 / hourOpen / Referral verified
Medical / Remote

Clinical Documentation Integrity (CDI) Leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Clinical Documentation Integrity (CDI) managers and leaders to evaluate AI tools designed to enhance clinical documentation accuracy and coding integrity workflows. Your expertise in DRG optimisation, HCC risk adjustment, physician query management, and clinical documentation best practices will directly inform AI systems that improve documentation quality, compliance, and revenue integrity. — Responsibilities • Lead clinical documentation integrity programs for inpatient and/or outpatient settings, overseeing concurrent and retrospective review workflows. • Evaluate AI-generated clinical documentation improvement suggestions, physician queries, and coding recommendations for clinical accuracy and compliance. • Conduct and review clinical documentation to ensure accurate capture of diagnoses, procedures, severity of illness, and risk of mortality. • Develop and manage physician query processes in alignment with AHIMA and ACDIS guidelines. • Monitor CDI program KPIs including query response rates, CC/MCC capture rates, case mix index, and DRG accuracy. • Collaborate with coding, compliance, and clinical teams to address documentation gaps and improve query processes. • Provide education to physicians and clinical staff on documentation requirements and best practices. • Ensure compliance with Official Coding Guidelines, CMS regulations, and payer-specific requirements. • Annotate AI outputs and provide structured clinical feedback to support AI training datasets. — Requirements • 5+ years of experience in clinical documentation integrity or improvement, with at least 2 years in a manager or leadership role. • Deep knowledge of MS-DRG methodology, CC/MCC hierarchies, and ICD-10-CM/PCS coding guidelines. • Expertise in physician query management per AHIMA and ACDIS compliant query standards. • Strong clinical background with the ability to interpret medical records and clinical documentation. • Proficiency with CDI software platforms (3M, Nuance, Optum360, or equivalent) and EHR systems. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to critically evaluate clinical documentation and AI-generated outputs. • Comfortable working independently in a fully remote environment. — Preferred Qualifications • CCDS (Certified Clinical Documentation Specialist) or CDIP (Clinical Documentation Improvement Practitioner) credential. • Experience with HCC risk adjustment and outpatient CDI programs. • Background in RN, RHIA, CCS, or similar clinical or coding credential. • Familiarity with AI-assisted CDI tools (e.g., Nuance DAX, 3M M\*Modal) and comfort evaluating AI-generated clinical content. • Experience presenting CDI performance metrics to clinical and executive leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in clinical documentation and revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$84 / hourOpen / Referral verified
Code / Remote

ML Engineer (Coding Agent Experience)

About the Role — \- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. — \- Contributors help evaluate and improve frontier AI coding models through structured technical assessments. — \- The work focuses on realistic machine learning engineering workflows and model evaluation. — \- Spots are limited and filling quickly on a first come, first serve basis. — What You'll Do — \- Use frontier AI coding agents to complete and evaluate complex machine learning and AI engineering tasks. — \- Review model-generated implementations involving model training, inference systems, MLOps, and LLM applications. — \- Identify bugs, edge cases, performance issues, and failure modes. — \- Compare outputs from multiple frontier models and assess their strengths and weaknesses. — \- Apply professional engineering judgment to realistic ML engineering scenarios. — Time Commitment — \- Sprint based project that runs in 12-24 hour stretches based on client requirement. — Compensation — \- $400 per accepted task. — \- Typical tasks take approximately 2–3 hours after ramp-up. — \- Compensation is tied to accepted work. — Who Should Apply — \- 2+ years of professional machine learning engineering experience. — \- Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products. — \- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. — \- Ability to evaluate model-generated machine learning implementations and technical tradeoffs. — \- Experience deploying ML systems to production is preferred.

$85 / hourOpen / Referral verified
Code / Remote

DevOps / SRE / Cloud Engineer (Coding Agent Experience)

About the Role — \- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. — \- Contributors help evaluate and improve frontier AI coding models through structured technical assessments. — \- The work focuses on realistic infrastructure engineering workflows and model evaluation. — \- Spots are limited and filling quickly on a first come, first serve basis. — What You'll Do — \- Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. — \- Review model-generated implementations involving cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation. — \- Identify bugs, edge cases, reliability issues, and failure modes. — \- Compare outputs from multiple frontier models and assess their strengths and weaknesses. — \- Apply professional engineering judgment to realistic infrastructure engineering scenarios. — Time Commitment — \- Sprint based project that runs in 12-24 hour stretches based on client requirement. — Compensation — \- $400 per accepted task. — \- Typical tasks take approximately 2–3 hours after ramp-up. — \- Compensation is tied to accepted work. — Who Should Apply — \- 2+ years of professional DevOps, SRE, or Cloud Engineering experience. — \- Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling. — \- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. — \- Ability to evaluate model-generated infrastructure and reliability engineering solutions. — \- Experience supporting production-scale systems is preferred.

$85 / hourOpen / Referral verified
Code / Remote

Cybersecurity Expert

Role Overview • Mercor is seeking senior cybersecurity professionals to build evaluation tasks for AI systems operating in security operations, incident response, and risk management contexts. • The workflows are calibrated to the threat sophistication, business risk stakes, and scope of major enterprise security programs. • This role builds worlds on two tracks: a US track (NIST Cybersecurity Framework, SOC 2) and an International track (ISO 27001, EU NIS2 Directive). Experts qualified in either or both tracks are encouraged to apply. • Contributors design cybersecurity scenarios, draft reference outputs, and write rubrics that capture how senior security leaders think. — Key Responsibilities • Construct cybersecurity scenarios spanning security operations center monitoring, incident response and forensics, vulnerability management, and security architecture design. • Build tasks across security operations and threat detection, incident response and digital forensics, vulnerability and penetration testing, security architecture and engineering, and governance/risk/compliance. • Develop scenarios involving tools such as SIEM platforms (Splunk, Microsoft Sentinel), EDR tools (CrowdStrike), vulnerability scanners (Tenable, Qualys), and GRC platforms used at major enterprises. • Apply cybersecurity methodologies (threat modeling, incident response playbooks, risk assessment) to the standards track a world targets (US: NIST CSF, SOC 2; International: ISO 27001, NIS2), and produce reference incident reports, security assessments, architecture designs, and compliance documentation. • Author rubrics that distinguish authentic security judgment from generic textbook or certification-exam-level recall. — Ideal Qualifications • 5+ years working as a security engineer or CISO at a major company or security firm (Mandiant, CrowdStrike, or an in-house CISO/security lead). • Direct ownership of incident response programs, security architecture, or compliance initiatives. • Fluency in security tooling, plus understanding of regulatory frameworks and the current threat landscape. • A recognized professional credential is strongly preferred (CISSP, CISM, or an international equivalent); prior rubric or training authorship is a plus.

$80 - $90 / hourOpen / Referral verified
Business / Remote

Sales & Marketing Experts

Role Overview • Mercor is seeking senior sales and marketing professionals to build evaluation tasks for AI systems operating in Fortune 500 go-to-market contexts. • The workflows are calibrated to the deal sizes, stakeholder complexity, and brand stakes of Fortune 500 and large public companies. • Contributors design enterprise GTM scenarios, draft reference outputs, and write rubrics that capture how senior F500 operators think. — Key Responsibilities • Construct enterprise sales scenarios spanning $1M+ ACV deals, multi-stakeholder buying committees, and complex procurement cycles at F500 accounts. • Build marketing tasks across F500 brand strategy, enterprise ABM, demand generation at scale, lifecycle, and category positioning. • Develop RevOps and GTM scenarios involving Salesforce Enterprise, Marketo, 6sense, Gong, and Outreach in F500 stacks. • Apply enterprise sales methodologies (MEDDIC, Challenger, Force Management) and produce reference deal strategies, account plans, and executive narratives. • Author rubrics that distinguish authentic enterprise GTM judgment from generic playbook recall. — Ideal Qualifications • 5+ years selling, marketing, or running RevOps at a Fortune 500 enterprise software vendor (Salesforce, Oracle, ServiceNow, SAP, Workday, Microsoft, AWS) or inside an F500 brand or marketing organization (P&G, JPMorgan, Unilever, Microsoft, PepsiCo). • Direct ownership of F500 accounts, F500 brand campaigns, or F500 demand programs. • Fluency in enterprise GTM tooling and methodologies, plus understanding of how F500 budgets, procurement, and legal review actually work. • Prior rubric, sales-enablement curriculum, or training-content authorship is a plus. — Compensation Note • Hourly Pay: $65 to $90 per hour, set by Mercor based on demonstrated expertise. • Minimum Commitment: 20 hours per week. • Onboarding via the Mercor Rubric Academy, a paid program that calibrates contributors to the quality bar before live work. • Advancement: strong contributors move into reviewer, lead, and domain SME roles with elevated rates.

$65 - $90 / hourOpen / Referral verified
Business / Remote

Product Management Expert

Role Overview • Mercor is seeking senior product management professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise product contexts. • The workflows are calibrated to the product scale, stakeholder complexity, and market stakes of Fortune 500 and large public company product organizations. • Contributors design enterprise product scenarios, draft reference outputs, and write rubrics that capture how senior F500 product leaders think. — Key Responsibilities • Construct enterprise product scenarios spanning multi-quarter roadmap planning, cross-functional stakeholder alignment, and complex go-to-market or platform decisions at F500 scale. • Build product tasks across F500 product strategy, discovery and user research, pricing and packaging, platform/API product management, and product-led growth. • Develop product operations scenarios involving tools such as Jira/Productboard, Amplitude/Mixpanel, Figma, and enterprise experimentation platforms in F500 stacks. • Apply enterprise product methodologies (JTBD, RICE/opportunity scoring, dual-track agile, OKRs) and produce reference PRDs, roadmap strategies, and executive-level product narratives. • Author rubrics that distinguish authentic enterprise product judgment from generic framework or bootcamp-level recall. — Ideal Qualifications • 5+ years as a product manager or product leader at a Fortune 500 technology or enterprise organization (Google, Amazon, Microsoft, Salesforce, Adobe) or inside an F500 product organization at a non-tech enterprise (JPMorgan, UPS, Unilever, PepsiCo). • Direct ownership of F500-scale products, platforms, or product lines with measurable business impact. • Fluency in enterprise product tooling and methodologies, plus understanding of how F500 budgets, cross-functional governance, and executive stakeholder alignment actually work. • Prior rubric, product-training curriculum, or PRD/documentation authorship is a plus.

$80 - $90 / hourOpen / Referral verified
Business / United States Remote

Marketing Specialist

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Marketing subject-matter experts (SMEs) with deep domain expertise and hands-on experience evaluating AI model outputs against rubrics to bring rigor and real-world brand/growth judgment to our AI training data. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps in brand strategy, growth marketing, and campaign-reasoning tasks. • Design challenging, domain-relevant marketing tasks and write accurate, well-reasoned solutions grounded in real marketing practice. • Evaluate AI model outputs against structured rubrics and provide clear, written feedback on correctness, judgment, and reasoning quality. • Develop and refine evaluation guidelines and scoring rubrics specific to marketing tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 8+ years of dedicated professional experience in marketing (e.g., brand strategy, growth marketing, performance marketing) at a recognized, top-tier organization (e.g., P&G, Unilever, Nike, Ogilvy, WPP, Omnicom, or equivalent). • Prior hands-on experience evaluating LLM/AI model outputs against rubrics or structured scoring criteria — mandatory; please describe this experience in your application. • Demonstrable career progression (e.g., Marketing Manager → Senior Manager → Director/VP of Marketing). • Ability to engage reliably for at least 35 hours/week during weekdays. • Verbal and written communication skills, problem-solving skills, and interpersonal skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$60 - $80 / hourOpen / Referral verified
Medical / United States Remote

Insurance Specialist

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Insurance subject-matter experts (SMEs) with deep domain expertise and hands-on experience evaluating AI model outputs against rubrics to bring rigor and real-world underwriting judgment to our AI training data. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps in underwriting, claims, and risk-assessment reasoning. • Design challenging, domain-relevant insurance tasks and write accurate, well-reasoned solutions grounded in real underwriting/claims practice. • Evaluate AI model outputs against structured rubrics and provide clear, written feedback on correctness, judgment, and reasoning quality. • Develop and refine evaluation guidelines and scoring rubrics specific to insurance tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 8+ years of dedicated professional experience in insurance (e.g., underwriting, claims, actuarial, risk management) at a recognized, top-tier organization (e.g., AIG, Chubb, Allstate, Progressive, MetLife, Marsh McLennan, or equivalent). • Prior hands-on experience evaluating LLM/AI model outputs against rubrics or structured scoring criteria — mandatory; please describe this experience in your application. • Demonstrable career progression (e.g., Underwriter → Senior Underwriter → VP of Underwriting). • Ability to engage reliably for at least 35 hours/week during weekdays. • Verbal and written communication skills, problem-solving skills, and interpersonal skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$60 - $80 / hourOpen / Referral verified
Medical / United States Remote

Retail Specialist

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Retail subject-matter experts (SMEs) with deep domain expertise and hands-on experience evaluating AI model outputs against rubrics to bring rigor and real-world retail judgment to our AI training data. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps in retail merchandising, category management, and operations reasoning. • Design challenging, domain-relevant retail tasks and write accurate, well-reasoned solutions grounded in real retail practice. • Evaluate AI model outputs against structured rubrics and provide clear, written feedback on correctness, judgment, and reasoning quality. • Develop and refine evaluation guidelines and scoring rubrics specific to retail tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 8+ years of dedicated professional experience in retail (e.g., merchandising, category management, retail operations, buying/planning) at a recognized, top-tier organization (e.g., Amazon, Walmart, Target, Nike, Costco, Home Depot, or equivalent). • Prior hands-on experience evaluating LLM/AI model outputs against rubrics or structured scoring criteria — mandatory; please describe this experience in your application. • Demonstrable career progression (e.g., Category Manager → Senior Manager → Director of Merchandising). • Ability to engage reliably for at least 35 hours/week during weekdays. • Verbal and written communication skills, problem-solving skills, and interpersonal skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$60 - $80 / hourOpen / Referral verified
Medical / Remote

Medical Billing Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Billing and Claims Managers to support the evaluation of AI tools designed to automate and improve medical billing and claims submission workflows. Your expertise in claims processing, payer requirements, EDI transactions, and billing compliance will directly inform AI systems that improve first-pass claim acceptance rates and accelerate revenue cycle performance. — Responsibilities • Oversee end-to-end medical billing and claims submission operations across professional fee and/or facility billing environments. • Evaluate AI-generated billing outputs, claim edits, and coding validations for accuracy and payer compliance. • Manage claims submission workflows including electronic claim generation, clearinghouse edits, and payer-specific billing requirements. • Monitor clean claim rates, rejection rates, and first-pass acceptance rates and develop improvement strategies. • Coordinate with coding, CDI, and collections teams to resolve billing edits and claim rejections. • Ensure compliance with CMS billing guidelines, HIPAA 837 transaction standards, and payer-specific billing rules. • Manage billing staff workload, productivity, and quality performance metrics. • Develop and maintain billing SOPs and payer-specific billing reference guides. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in medical billing and claims management, with at least 2 years in a management role. • Deep knowledge of professional fee (CMS-1500/837P) and/or facility (UB-04/837I) billing requirements. • Expertise in HIPAA 837 transaction standards, clearinghouse operations, and payer-specific billing rules. • Strong understanding of Medicare, Medicaid, and commercial payer billing requirements. • Proficiency with billing platforms (Epic, Athenahealth, AdvancedMD, or equivalent) and clearinghouse tools. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify billing errors and compliance issues in AI-generated outputs. — Preferred Qualifications • CPC, CCS, CHFP, or CRCR certification. • Experience with automated billing platforms and RCM technology implementations. • Background in multi-speciality physician group, hospital, or health system billing operations. • Familiarity with AI tools and comfort evaluating AI-generated billing content. • Experience with payer contract interpretation and billing compliance program management. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in medical billing and revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$80 / hourOpen / Referral verified
Code / Remote

Coding Manager / HIM Coding Leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Coding Managers and HIM Coding leaders to evaluate AI-powered coding solutions and help train next-generation autonomous coding systems. Your expertise in professional fee (profee) and/or inpatient facility coding, ICD-10-CM/PCS, CPT/HCPCS, and HIM operations will directly inform AI tools designed to improve coding accuracy, productivity, and compliance across healthcare settings. — Responsibilities • Oversee professional fee and/or facility inpatient coding operations, ensuring accuracy, productivity, and compliance with coding guidelines. • Evaluate AI-generated coding assignments, including ICD-10-CM/PCS diagnoses, procedure codes, CPT/HCPCS codes, and DRG assignments, for accuracy and compliance. • Conduct coding quality audits and provide targeted feedback to coding staff and AI systems. • Monitor coding KPIs including coder productivity, accuracy rates, unbilled accounts, and claim denial rates attributable to coding errors. • Manage coding workflow queues, work distribution, and turnaround time compliance. • Ensure adherence to Official Coding Guidelines, CMS regulations, and payer-specific coding requirements. • Provide ongoing coding education and compliance training to coding staff. • Collaborate with CDI, billing, and compliance teams to address coding-related revenue integrity issues. • Annotate AI-generated coding outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in medical coding, with at least 2 years in a coding manager or HIM leadership role. • Expert knowledge of ICD-10-CM/PCS, CPT/HCPCS, and Official Coding Guidelines. • Proficiency in professional fee (profee) coding and/or facility inpatient coding with DRG assignment experience. • Experience conducting coding audits and developing coding quality improvement programs. • Proficiency with coding software (3M, Nuance, Optum360, TruCode) and EHR platforms. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify coding errors, compliance risks, and AI output inaccuracies. — Preferred Qualifications • CPC (Certified Professional Coder), CCS (Certified Coding Specialist), RHIA, or RHIT credential. • Experience with computer-assisted coding (CAC) tools and NLP-based coding platforms. • Background in inpatient facility coding with DRG optimisation experience. • Familiarity with AI coding tools and comfort evaluating AI-generated coding assignments. • Experience presenting coding performance data and quality metrics to leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in medical coding and health information management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$80 / hourOpen / Referral verified
Medical / Remote

Patient Access Leader

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Patient Access leaders to support the evaluation and improvement of AI tools designed for healthcare revenue cycle workflows. This is an opportunity to apply your expertise in patient registration, pre-registration, scheduling, and intake coordination to help shape AI systems that will transform front-end healthcare operations. — Responsibilities • Lead end-to-end patient access operations including pre-registration, registration, scheduling, and intake workflows across inpatient and outpatient settings. • Evaluate AI-generated outputs related to patient access processes, identifying errors and providing structured, actionable feedback. • Develop and enforce policies and procedures for registration accuracy, demographic capture, and point-of-service collections. • Monitor KPIs including registration error rates, pre-registration completion rates, insurance verification accuracy, and front-end denial rates. • Ensure clean claim submission from point of entry by identifying and resolving front-end deficiencies. • Collaborate with billing, clinical, and IT teams to streamline patient access workflows and reduce downstream claim errors. • Train and mentor staff on registration best practices, compliance requirements, and system usage. • Ensure compliance with HIPAA, CMS guidelines, and payer-specific requirements at the point of service. • Document observations, annotate AI outputs, and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in patient access, patient registration, or front-end revenue cycle management, with at least 2 years in a manager or director-level role. • Deep knowledge of pre-registration, insurance verification, scheduling intake, and point-of-service collections. • Strong understanding of HIPAA, CMS regulations, and commercial and government payer requirements. • Proficiency with Epic, Cerner, Meditech, or similar EHR and registration systems. • Exceptional written and verbal English communication skills. • High attention to detail and ability to identify inconsistencies in workflows, data, or AI-generated outputs. • Comfortable working independently in a fully remote environment. • Strong organisational skills with the ability to manage multiple priorities simultaneously. — Preferred Qualifications • NAHAM Certified Healthcare Access Manager (CHAM) or Certified Healthcare Access Associate (CHAA) credential. • Experience with denial prevention strategies at the point of registration. • Familiarity with AI tools such as ChatGPT, Claude, or similar systems. • Background in health system, hospital, or large physician group settings. • Experience developing SOPs and training materials for patient access teams. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in healthcare revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$80 / hourOpen / Referral verified
Finance / Remote

FP&A Expert

Role Overview • Mercor is seeking senior FP&A professionals to build evaluation tasks for AI systems operating in financial planning, forecasting, and business analysis contexts. • The workflows are calibrated to the forecasting complexity, decision stakes, and scope of major corporate financial planning functions. • This role builds worlds on two tracks: a US track (US GAAP-based reporting and corporate finance conventions) and an International track (IFRS-based reporting and international corporate finance conventions). Experts qualified in either or both tracks are encouraged to apply. • Contributors design FP&A scenarios, draft reference outputs, and write rubrics that capture how senior FP&A leaders think. — Key Responsibilities • Construct FP&A scenarios spanning budgeting and forecasting cycles, variance analysis and business partnering, and capital allocation and strategic planning decisions. • Build tasks across financial planning and forecasting, variance analysis and reporting, business partnering, capital budgeting and investment analysis, and financial modeling. • Develop scenarios involving tools such as FP&A platforms (Anaplan, Adaptive Insights, Planful), BI tools (Tableau, Power BI), and ERP systems used at large companies. • Apply FP&A methodologies (driver-based forecasting, variance analysis, scenario and sensitivity modeling) to the reporting basis a world targets (US GAAP-based; IFRS-based), and produce reference financial models, forecast decks, variance reports, and business case analyses. • Author rubrics that distinguish authentic FP&A judgment from generic textbook or spreadsheet-template-level work. — Ideal Qualifications • 5+ years working as an FP&A leader, Director of FP&A, or CFO at a major company. • Direct ownership of budgeting cycles, forecasting models, or capital allocation decisions. • Fluency in FP&A tooling, plus understanding of corporate finance and financial modeling best practices. • A recognized professional credential is a plus (CFA, MBA, or CPA); prior rubric or training authorship is a plus.

$80 - $90 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Swedish

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Swedish and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Swedish, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Swedish music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Swedish genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Swedish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$42 - $78 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Norwegian

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Norwegian and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Norwegian, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Norwegian music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Norwegian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Norwegian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$42 - $78 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Norwegian

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Norwegian and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Norwegian, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Norwegian genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Norwegian • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$42 - $78 / hourOpen / Referral verified
Code / Remote

Data Engineer (Coding Agent Experience)

About the Role — \- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. — \- Contributors help evaluate and improve frontier AI coding models through structured technical assessments. — \- The work focuses on realistic data engineering workflows and model evaluation. — \- Spots are limited and filling quickly on a first come, first serve basis. — What You'll Do — \- Use frontier AI coding agents to complete and evaluate complex data engineering tasks. — \- Review model-generated implementations involving ETL pipelines, data warehouses, analytics platforms, and distributed data systems. — \- Identify bugs, edge cases, scalability issues, and failure modes. — \- Compare outputs from multiple frontier models and assess their strengths and weaknesses. — \- Apply professional engineering judgment to realistic data engineering scenarios. — Time Commitment — \- Sprint based project that runs in 12-24 hour stretches based on client requirement. — Compensation — \- $400 per accepted task. — \- Typical tasks take approximately 2–3 hours after ramp-up. — \- Compensation is tied to accepted work. — Who Should Apply — \- 2+ years of professional data engineering experience. — \- Experience building ETL pipelines, data warehouses, analytics platforms, or distributed data systems. — \- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. — \- Ability to evaluate model-generated data infrastructure and pipeline implementations. — \- Experience operating large-scale data platforms is preferred.

$80 / hourOpen / Referral verified
Code / Remote

Applied Engineering Benchmark Specialist

Role Overview — We are seeking expert engineers to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core engineering domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of engineering expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Engineering Domains Covered — Semiconductor Design & Manufacturing (VLSI), Control Science and Engineering, Mechatronics, Reactor, Plant Design & Separations, Reservoir Engineering & Maintenance, Bioinstrumentation & Biotechnology. — Key Responsibilities • Author original engineering questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Engineering or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level engineering principles, applied mathematics, and domain-specific standards • Professional engineering licensure (PE) or industry experience is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$61 - $77 / hourOpen / Referral verified
Math / Remote

Applied Mathematics Benchmark Specialist

Role Overview — We are seeking expert mathematicians to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core mathematics domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of mathematical expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Mathematics Domains Covered — Signal Processing, Financial Mathematics & Actuarial Science, Mathematical Economics, Mathematical Modeling of Ecological & Biological Systems, Mathematical Programming & Combinatorial Optimization, Geomathematics & Climate Modeling. — Key Responsibilities • Author original math questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Mathematics, Applied Mathematics, Statistics, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level mathematical concepts and formal proof writing • Experience with rigorous academic problem design or mathematical competition writing is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$61 - $77 / hourOpen / Referral verified
STEM / Remote

Applied Chemistry Benchmark Specialist

Role Overview — We are seeking expert chemists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core chemistry domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of chemistry expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Chemistry Domains Covered — Materials, Polymer & Electronic Chemistry, Industrial & Process Chemistry, Energy Storage & Environmental Chemistry, Pharmaceutical & Agrochemical Chemistry, Consumer, Food & Specialty Chemicals. — Key Responsibilities • Author original chemistry questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Chemistry, Biochemistry, Chemical Engineering, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level chemistry concepts, reaction mechanisms, and quantitative analysis • Experience with rigorous academic problem design or chemistry olympiad writing is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$61 - $77 / hourOpen / Referral verified
STEM / Remote

Applied Physics Benchmark Specialist

Role Overview — We are seeking expert physicists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core physics domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of physics expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Physics Domains Covered — Semiconductor Physics, Nanoelectronics & Spintronics, Photonics, Quantum Optics & Ultrafast, Quantum Sensing & Metrology, Plasma Physics & Fusion Energy, Nonlinear Dynamics & Turbulence, Geophysics & Reservoir Simulation. — Key Responsibilities • Author original physics questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Physics, Applied Physics, Astrophysics, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level physics concepts and mathematical formalism • Experience with rigorous academic problem design or physics olympiad writing is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$61 - $77 / hourOpen / Referral verified
STEM / Remote

Chess Expert (1800+ ELO) — Game Reconstruction & Transcript Correction

We're hiring strong chess players to help build a high-quality chess dataset for an AI research partner: — 1. Transcript Correction — Watch a video of a strong player narrating a game and fix speech-to-text errors in the commentary transcript (piece names, squares, moves). The form syncs the transcript to the video (click a word to jump there). 2. Game Reconstruction (the main task) — Working ONLY from the corrected transcript (no video, no searching for the source game), rebuild the full game on an analysis board (lichess/chess.com) and paste the PGN. Then determine whether the transcript uniquely specifies the complete game; if not, itemize the minimal missing information needed. — You're a fit if you: • Are rated 1800+ ELO (FIDE or online equivalent) • Read algebraic notation fluently and are comfortable using an analysis board • Have a strong ear for spoken chess commentary and careful attention to detail

$90 / hourOpen / Referral verified
Code / Remote

Civil Engineering Expert

Role Overview • Mercor is seeking senior civil engineering professionals to build evaluation tasks for AI systems operating in large-scale infrastructure and construction contexts. • The workflows are calibrated to the design complexity, public safety stakes, and regulatory scope of major infrastructure projects and large public/private construction programs. • This role builds worlds on two standards tracks: an International track (Eurocodes and ISO) and a US track (ASCE, ACI, AISC, AASHTO). Experts qualified in either or both tracks are encouraged to apply. • Contributors design civil engineering scenarios, draft reference outputs, and write rubrics that capture how senior civil engineering leaders think. — Key Responsibilities • Construct civil engineering scenarios spanning large-scale infrastructure design cycles, multi-stakeholder permitting and regulatory review, and complex construction management or public works processes. • Build tasks across structural design and analysis, transportation and highway engineering, water resources and environmental engineering, geotechnical engineering, and construction project management. • Develop engineering scenarios involving tools such as AutoCAD Civil 3D, Revit, STAAD.Pro/ETABS, HEC-RAS, and enterprise project management platforms used on major infrastructure programs. • Apply civil engineering methodologies (structural load analysis, geotechnical site assessment, hydrology/hydraulic modeling) to the standards track a world targets (US: ASCE 7, ACI 318, AISC 360, AASHTO LRFD; International: Eurocodes EN 1990-1998 and ISO), and produce reference design plans, engineering calculations, and stakeholder/regulatory-facing technical narratives. • Author rubrics that distinguish authentic civil engineering judgment from generic textbook or coursework-level recall. — Ideal Qualifications • 5+ years working as a civil engineer or engineering lead • Direct ownership of large-scale infrastructure designs, permitting processes, or construction project delivery. • Fluency in civil engineering tooling and methodologies, plus understanding of how regulatory approvals, environmental review (NEPA or an international equivalent), and public agency oversight actually work. • A recognized professional engineering credential is strongly preferred (US PE, or an international equivalent such as CEng, EUR ING, or P.Eng); prior rubric, engineering-training curriculum, or design documentation authorship is a plus.

$70 - $80 / hourOpen / Referral verified
STEM / Remote

Applied Biology Benchmark Specialist

Role Overview — We are seeking expert biologists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core biology domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of biology expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Biology Domains Covered — Pharmaceutical Manufacturing, Industrial and Synthetic Biology, Medical Research & Drug Discovery, Agricultural, Environmental & Food Biology. — Key Responsibilities • Author original biology questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Biology, Molecular Biology, Biochemistry, Neuroscience, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level biological concepts, experimental design, and data interpretation • Research publications or laboratory experience in biological sciences is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$60 - $75 / hourOpen / Referral verified
Finance / Remote

A/R Follow-up Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced A/R Follow-Up Managers to evaluate AI tools designed to automate accounts receivable follow-up and payer collections workflows. Your expertise in claim status follow-up, payer correspondence, and A/R management will directly inform AI systems that reduce days in accounts receivable and improve revenue recovery. — Responsibilities • Lead A/R follow-up operations across commercial, Medicare, Medicaid, and managed care payers, ensuring timely resolution of outstanding claims. • Evaluate AI-generated A/R follow-up recommendations, claim status inquiry outputs, and payer correspondence drafts for accuracy and effectiveness. • Manage claim status follow-up workflows including electronic claim status inquiries (276/277 EDI), payer portal follow-up, and phone-based resolution. • Prioritise A/R queues by ageing bucket, payer, and dollar value to maximise revenue recovery. • Identify and resolve claim payment discrepancies, payer processing errors, and underpayments. • Monitor A/R KPIs including days in A/R, ageing bucket distribution, collection rates, and write-off rates. • Develop and implement payer-specific follow-up strategies to accelerate claim resolution. • Ensure compliance with FDCPA, HIPAA, and payer-specific follow-up and timely filing requirements. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in A/R follow-up, payer collections, or revenue cycle operations, with at least 2 years in a management role. • Deep knowledge of claim status follow-up workflows, EDI 276/277 transactions, and payer-specific collections processes. • Strong understanding of Medicare, Medicaid, and commercial payer claims processing and payment timelines. • Experience prioritising and managing high-volume A/R queues across multiple payers. • Proficiency with billing systems and A/R management platforms. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to identify payment errors and discrepancies in AI-generated A/R content. — Preferred Qualifications • CRCR, CPC, or CHFP certification. • Experience with RCM technology platforms featuring automated A/R follow-up capabilities. • Background in multi-payer follow-up operations in hospital or physician group settings. • Familiarity with AI tools and comfort evaluating AI-generated A/R follow-up content. • Experience developing A/R reduction action plans and presenting performance to leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in accounts receivable and revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$75 / hourOpen / Referral verified
Medical / Remote

Pharmacy Prior Authorization & Specialty-Medication Access Specialist

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Medication and Pharmacy Prior Authorisation specialists with expertise in speciality medication access to evaluate AI tools designed to automate pharmacy benefit management and speciality drug authorisation workflows. Your deep knowledge of drug formularies, step therapy, and payer PA criteria will directly shape AI systems that improve patient access to critical therapies. — Responsibilities • Process and manage prior authorisation requests for speciality medications, biologics, and high-cost pharmaceuticals across commercial, Medicare Part D, and Medicaid payers. • Review clinical documentation and pharmacy benefit criteria to determine medical necessity for speciality drug authorisations. • Evaluate AI-generated pharmacy prior authorisation recommendations for accuracy, completeness, and payer compliance. • Coordinate with prescribers, speciality pharmacies, and payers to obtain timely medication authorisations. • Manage appeals for denied pharmacy authorisations, including peer-to-peer requests and exception processes. • Navigate payer formularies, step therapy requirements, and speciality drug coverage policies. • Track authorisation status, approval rates, and denial trends to identify process improvement opportunities. • Ensure compliance with payer-specific pharmacy benefit requirements and CMS Part D guidelines. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in pharmacy prior authorisation, speciality medication access, or pharmacy benefit management. • Deep knowledge of speciality drug formularies, step therapy protocols, and payer PA criteria across commercial and government payers. • Experience with speciality pharmacy platforms and pharmacy benefit management (PBM) systems. • Familiarity with Medicare Part D coverage determination and exception processes. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to evaluate clinical documentation and AI-generated outputs. — Preferred Qualifications • Certified Pharmacy Technician (CPhT) or Certified Prior Authorisation Professional (CPAP) credential. • Experience with hub services and patient assistance programs for speciality medications. • Background in speciality pharmacy, infusion services, or oncology/rheumatology drug access. • Familiarity with AI tools and comfort evaluating AI-generated pharmacy content. • Experience with manufacturer copay assistance and patient access programs. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in healthcare revenue cycle management. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$75 / hourOpen / Referral verified
Business / Remote

HR Expert

Role Overview • Mercor is seeking senior HR professionals to build evaluation tasks for AI systems operating in talent management, employee relations, and organizational design contexts. • The workflows are calibrated to the organizational scale, compliance stakes, and scope of major workforce programs at large companies. • This role builds worlds on two tracks: a US track (Title VII, FLSA, ADA, and state employment law) and an International track (UK Employment Rights Act, EU labor directives). Experts qualified in either or both tracks are encouraged to apply. • Contributors design HR scenarios, draft reference outputs, and write rubrics that capture how senior HR leaders think. — Key Responsibilities • Construct HR scenarios spanning talent acquisition and workforce planning, employee relations and performance management, and compensation and benefits design. • Build tasks across recruiting and talent acquisition, employee and labor relations, compensation and benefits, HR compliance and employment law, and organizational development. • Develop scenarios involving tools such as HRIS platforms (Workday, SAP SuccessFactors), applicant tracking systems, and compensation benchmarking tools (Radford, Mercer) used at large organizations. • Apply HR methodologies (workforce planning, compensation benchmarking, employee relations investigation) to the standards track a world targets (US: Title VII, FLSA, ADA, state law; International: Employment Rights Act, EU labor directives), and produce reference HR policies, investigation reports, compensation analyses, and organizational design proposals. • Author rubrics that distinguish authentic HR judgment from generic textbook or certification-exam-level recall. — Ideal Qualifications • 5+ years working as an HR leader or Chief People Officer at a major company or HR consulting firm (Mercer, Korn Ferry, or in-house VP/CHRO at a large company). • Direct ownership of talent programs, employee relations matters, or compensation design. • Fluency in HR tooling and methodologies, plus understanding of how employment law compliance actually works. • A recognized professional credential is strongly preferred (SHRM-SCP, SPHR, or an international equivalent); prior rubric or training authorship is a plus.

$70 - $80 / hourOpen / Referral verified
Language / Remote

Design Expert

Role Overview • Mercor is seeking senior product design and AI-assisted development professionals to build evaluation tasks for AI systems operating in rapid prototyping, product design, and human-AI collaborative coding contexts. • The workflows are calibrated to the design complexity, user experience stakes, and scope of major product development and design systems work. • This role builds worlds across two subdomains: a Vibecoding track (AI-assisted, natural-language-driven software development and rapid prototyping) and a Design track (product, UX/UI, and visual design systems). Experts qualified in either or both subdomains are encouraged to apply. • Contributors design scenarios, draft reference outputs, and write rubrics that capture how senior product builders and designers think. — Key Responsibilities • Construct scenarios spanning rapid prototyping and iteration cycles, design system development, and cross-functional product-design-engineering collaboration. • Build tasks across AI-assisted application development, natural-language-to-code workflows, and rapid prototyping (Vibecoding); and UX/UI design, design systems and component libraries, user research, and usability testing (Design). • Develop scenarios involving tools such as AI coding assistants (Cursor, Claude Code, v0, Replit) for the Vibecoding subdomain, and Figma, prototyping tools (Framer), and design systems tooling for the Design subdomain. • Apply subdomain-specific methodologies (iterative prompt-driven development for Vibecoding; human-centered design principles and heuristic evaluation for Design), and produce reference prototypes, code artifacts, design specs, and usability evaluation reports. • Author rubrics that distinguish authentic senior-level product judgment from generic tutorial-level or template-driven work. — Ideal Qualifications • 5+ years working as a product designer, design lead, or AI-native engineer/vibecoder at a major product- or design-forward company (Figma, Airbnb, Notion, Linear, or a notable independent/startup product builder). • Direct ownership of product design systems or AI-assisted development workflows shipped to production. • Fluency in vibecoding or design tooling, plus understanding of user-centered design principles and rapid prototyping practices. • A portfolio of shipped work is strongly preferred in place of a formal credential; prior rubric, training, or design critique authorship is a plus.

$70 - $80 / hourOpen / Referral verified
Finance / Remote

Accounting Expert

Role Overview • Mercor is seeking senior accounting professionals to build evaluation tasks for AI systems operating in financial reporting, audit, and technical accounting contexts. • The workflows are calibrated to the reporting complexity, materiality stakes, and scope of major public company and enterprise accounting functions. • This role builds worlds on two tracks: a US track (US GAAP, PCAOB standards) and an International track (IFRS, International Standards on Auditing). Experts qualified in either or both tracks are encouraged to apply. • Contributors design accounting scenarios, draft reference outputs, and write rubrics that capture how senior accountants and auditors think. — Key Responsibilities • Construct accounting scenarios spanning financial statement preparation, technical accounting research, external and internal audit processes, and complex transaction accounting (M&A, revenue recognition, leases). • Build tasks across financial reporting, technical accounting and research, audit and assurance, complex transaction accounting, and internal controls/SOX compliance. • Develop scenarios involving tools such as ERP systems (SAP, Oracle), consolidation software, audit management platforms, and research tools (Bloomberg Tax, RIA Checkpoint) used at major firms. • Apply accounting methodologies (revenue recognition, lease accounting, consolidation) to the standards track a world targets (US: US GAAP/ASC codification, PCAOB standards; International: IFRS, ISA), and produce reference financial statements, technical accounting memos, audit workpapers, and internal control assessments. • Author rubrics that distinguish authentic accounting judgment from generic textbook or CPA exam-level recall. — Ideal Qualifications • 5+ years working as an accountant, controller, or audit partner at a major accounting firm or corporation (Big Four, or a corporate controller/CFO). • Direct ownership of financial reporting, audit engagements, or technical accounting matters. • Fluency in accounting tooling, plus understanding of regulatory reporting (SEC filings) and audit standards. • A recognized professional credential is strongly preferred (CPA, or an international equivalent such as ACCA or CA); prior rubric or training authorship is a plus.

$70 - $80 / hourOpen / Referral verified
Legal / US Remote

Legal Expert

Join a growing network of industry experts supporting AI research and development. — Overview — We're building a select group of experienced legal professionals to help evaluate and improve how AI systems understand and reason about legal topics. You'll bring your real-world expertise to review content for accuracy, answer domain-specific questions, and share feedback that helps make AI tools more reliable and useful in legal contexts. — This is a flexible, part-time engagement — approximately 10 hours per week, fully remote, with the opportunity to continue and grow over time. — What You'll Do • Review and evaluate written content and AI-generated responses in your area of legal expertise for accuracy and quality. • Answer questions and provide clear, well-reasoned feedback based on your professional experience. • Help identify useful reference material relevant to your field. • Share insight into how legal professionals actually work — the tools, workflows, and standards you rely on day to day. • Collaborate with a small group of other subject-matter experts across different fields. — What We're Looking For • Admitted to a United States state bar. A background as a legal librarian is a plus, though not required. • 2+ years of professional legal experience, ideally spanning more than one practice area (e.g. contracts, litigation, regulatory/compliance, corporate, IP). • Comfortable researching and writing using standard legal tools and references (case law databases, statutes, treatises). • Some experience writing or publishing legal content — memos, briefs, CLE materials, or similar — is a plus. • Comfortable using everyday AI tools (e.g. ChatGPT or similar). • Based in the United States. • Professional working English. • Able to commit reliably to about 10 hours per week. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Business / Remote

Wealth Management & Asset Management Expert

Role Overview • Mercor is seeking senior wealth management and asset management professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise investment and advisory contexts. • The workflows are calibrated to the portfolio complexity, regulatory stakes, and client sophistication of Fortune 500 asset managers and large private wealth institutions. • Contributors design enterprise wealth and asset management scenarios, draft reference outputs, and write rubrics that capture how senior F500 investment professionals think. — Key Responsibilities • Construct enterprise wealth management scenarios spanning high-net-worth portfolio construction, multi-stakeholder trust and estate planning, and complex fiduciary or regulatory review cycles. • Build asset management tasks across F500 portfolio strategy, alternative investments, institutional client servicing, risk and performance analytics, and manager due diligence. • Develop investment operations scenarios involving tools such as Bloomberg Terminal, Aladdin, Morningstar Direct, and enterprise portfolio/order management systems (OMS) in F500 stacks. • Apply enterprise investment methodologies (modern portfolio theory, factor investing, fiduciary duty standards, ESG integration) and produce reference investment policy statements, portfolio strategies, and client-facing executive narratives. • Author rubrics that distinguish authentic enterprise wealth and asset management judgment from generic textbook or CFA-exam-level recall. — Ideal Qualifications • 5+ years working as a portfolio manager, wealth advisor, or investment professional at a Fortune 500 asset manager, private bank, or wirehouse (BlackRock, Vanguard, Morgan Stanley, Goldman Sachs, JPMorgan). • Direct ownership of F500-scale client portfolios, institutional mandates, or investment strategy decisions. • Fluency in enterprise investment tooling and methodologies, plus understanding of how F500 regulatory compliance (SEC, FINRA), fiduciary standards, and client reporting actually work. • Prior rubric, investment-training curriculum, or portfolio documentation authorship is a plus.

$70 - $80 / hourOpen / Referral verified
Finance / US Remote

Finance Specialist — CFA/ACA/ACCA/CPA Required

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking Finance subject-matter experts (SMEs) with solid domain expertise to bring rigor and real-world financial judgment to our AI training data. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps in financial reasoning, analysis, and decision-making. • Design challenging, domain-relevant finance tasks and write accurate, well-reasoned solutions grounded in real financial practice. • Evaluate AI model outputs against structured rubrics and provide clear, written feedback on correctness, judgment, and reasoning quality. • Develop and refine evaluation guidelines and scoring rubrics specific to finance tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 2+ years of dedicated professional experience in finance (e.g., investment banking, asset management, corporate finance, accounting/audit, corporate treasury, financial planning & analysis) — not a generalist role that only touches finance peripherally. • Professional finance credential required: CFA, ACA, ACCA, or CPA. • Ability to engage reliably for at least 10 hours/week during weekdays. • Verbal and written communication skills, problem-solving skills, and interpersonal skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$65 - $90 / hourOpen / Referral verified
Business / Remote

Denials Management & Appeals Manager

Mercor is working with a leading AI research lab to improve the capabilities of next-generation AI systems. We are seeking experienced Denials Management and Appeals Managers to evaluate AI tools designed to automate denial prevention, appeal writing, and root cause analysis workflows. Your expertise in payer denial patterns, clinical and technical appeals, and denial management analytics will directly inform AI systems that reduce denial rates and maximise revenue recovery. — Responsibilities • Lead denials management and appeals operations, overseeing the identification, categorisation, and resolution of claim denials. • Evaluate AI-generated appeal letters, denial root cause analyses, and denial prevention recommendations for accuracy and effectiveness. • Analyse denial trends by payer, denial code (CARC/RARC), and denial category to identify systemic root causes. • Develop and manage clinical and technical appeal strategies across multiple payer types. • Coordinate with clinical, coding, billing, and compliance teams to implement denial prevention initiatives. • Monitor denial management KPIs including denial rates, appeal overturn rates, revenue recovery, and days in A/R. • Manage the appeals calendar to ensure timely submission within payer and regulatory deadlines. • Ensure compliance with payer appeal requirements, CMS regulations, and timely filing deadlines. • Annotate AI outputs and provide structured feedback to support AI training datasets. — Requirements • 5+ years of experience in denials management, appeals, or revenue cycle operations, with at least 2 years in a management role. • Deep knowledge of CARC/RARC denial codes, payer denial patterns, and appeal strategies across commercial, Medicare, and Medicaid payers. • Strong understanding of clinical and technical appeal processes including peer-to-peer reviews and external reviews. • Experience with denial analytics platforms and revenue cycle reporting tools. • Proficiency with EHR systems and billing platforms. • Exceptional written and verbal English communication skills. • High attention to detail with the ability to evaluate appeal quality and identify errors in AI-generated denial management content. — Preferred Qualifications • CPC, CCS, CRCR, or CHFP certification. • Experience with AI-assisted denial management platforms (e.g., Waystar, Experian Health, Nthrive). • Background in complex clinical appeals including medical necessity, experimental/investigational, and level of care denials. • Familiarity with AI tools and comfort evaluating AI-generated appeal and denial content. • Experience presenting denial management performance to revenue cycle leadership. — Why Join? • Contribute to the development of frontier AI systems in healthcare. • Collaborate with a world-class AI research organisation. • Gain exposure to cutting-edge AI workflows in denials management and revenue cycle. • Opportunity to work on high-impact projects shaping the future of healthcare AI.

$70 - $93 / hourOpen / Referral verified
Language / Remote

Mechanical Engineering Writer - Engineering FRQ

We're looking for a senior mechanical engineering expert to help build a benchmark of the hardest reasoning questions in the field — questions specifically designed to be difficult enough that today's most capable AI models still get them wrong. You'll author original, free-response engineering problems grounded in real industry scenarios, write a complete expert-level solution for each, and validate difficulty by running it against three frontier language models until at least one fails. — Requirements: PhD in Mechanical Engineering, plus 8+ years of hands-on professional/industry experience in design, analysis, or R&D. Candidates from well-known/blue-chip employers preferred.

$85 / hourOpen / Referral verified
Language / Remote

Generalist Expert (UK/Europe)

We’re looking for UK or Europe-based generalist experts to help train and evaluate frontier AI models. In this role, you’ll review and assess a wide range of everyday professional content — documents, slides, spreadsheets, and other written materials — judging them for quality, accuracy, clarity, and completeness. Your feedback directly shapes how AI systems reason about and produce real-world work product. — This is a great fit if you’re a sharp, detail-oriented generalist who’s comfortable moving across different formats and subject areas. You don’t need deep specialization in any one field — strong judgment, careful reading, and clear written reasoning matter most. — What you’ll do • Review and evaluate documents, slides, spreadsheets, and similar materials for quality and correctness • Provide clear, well-reasoned written feedback and ratings • Compare and rank AI-generated outputs against defined criteria • Flag errors, inconsistencies, and gaps in reasoning or formatting — What we’re looking for • Bachelor’s degree (minimum requirement) • Based in the UK or Europe • Excellent reading comprehension and written communication in English • Strong attention to detail and sound judgment across varied subject matter • Comfort working independently across common productivity tools (docs, slides, sheets) — No prior AI or machine-learning experience is required — we’ll provide the guidelines and context you need to succeed.

$50 - $70 / hourOpen / Referral verified
Code / Remote

Network Engineer - Data for Autonomous Systems annotation

Are you a Level 3 / Tier 3 network support engineer interested in data science and autonomous infrastructure? Our client is building vertically integrated networking systems and using the data they generate to power the next generation of AI-driven infrastructure. They're looking for engineers experienced in final escalations, packet analysis, troubleshooting, and RCA workflows to help label, annotate, and structure networking data from real production systems. — This is a hands-on role that blends your L3 troubleshooting and incident-response experience with a growing understanding of how data pipelines are built and used in AI systems. — In this role, you'll: • Review real-world data from deployed networks: logs, configs, telemetry, event streams • Label and classify network behaviors, issues, anomalies, and incident patterns • Help define schemas and structure for large-scale data pipelines that downstream ML models will train on — You're a strong fit if you: • Work today as a Level 3 / Tier 3 / Principal Support Engineer keeping existing enterprise infrastructure online and stable — on-call rotation, RCAs, final escalations, troubleshooting outages — Must Have • Have hands-on experience with end-customer enterprise networks (switches, APs, firewalls in retail, healthcare, financial, manufacturing, university, hospitality, etc.) — Must Have • Bring hands-on Wi-Fi/wireless proficiency — enterprise WLAN controllers (Cisco WLC, Aruba, or Meraki), 802.1X/RADIUS, and wireless troubleshooting — Must Have • Do packet-level troubleshooting yourself — Wireshark, tcpdump, SPAN captures • Are curious about how raw infra data becomes machine learning input — This is a maintainer role — likely not the right fit if your current work is mainly network design/architecture, cloud/SRE, security/SOC, or IT helpdesk. — Your work will directly feed into the pipelines that power client's AI models, and help shape how intelligent systems reason about networks in the real world. — Here are more details about the role: • You will interface directly with the client team. • You are expected to work 30-40 hours/week, with your hours overlapping the Pacific (PT) business day. • This is an individual 1099 contract paid to a personal account — no corp-to-corp or agency billing. • You must be authorized to work in the US or Canada without sponsorship.

$50 - $70 / hourOpen / Referral verified
Business / United States Remote

Project Coordinator - AI & Data Projects

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — Are you ready to help shape the future of artificial intelligence? Join a leading AI lab's cutting-edge GenAI team, where you'll be at the forefront of building groundbreaking AI models. We're seeking talented Project Coordinators to support and accelerate world-class AI research and data operations — acting as the operational bridge between program leadership and a growing team of domain experts. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. What You'll Do • Act as a day-to-day point of contact for AI data projects, helping keep workflows and operations running smoothly. • Work closely with program leads and domain experts — take inputs from leadership, turn them into clear guidelines, and share them with expert teams. • Help onboard and support new experts as the program grows. • Join weekly business reviews, keep notes and action items organized, and lead calls when needed. • Help spot quality issues, understand root causes, and share findings across teams. • Contribute to creating and refining project guidelines with research and product partners. • Flag potential roadblocks early and help get them resolved. • Track project deliverables using Google Sheets and Excel. • Stay flexible and adapt as project needs evolve. — 3. Qualifications • Location: Must be based in the USA. • Education: STEM background strongly preferred. • Experience: 3+ years in project coordination or project management, with the ability to coordinate large, cross-functional teams. • AI fluency: hands-on experience on AI training-data or human-data projects — as a project coordinator, team lead, EPM, or expert contributor. • Ideal: prior project coordination on AI/data programs at organizations like Meta, TikTok, or Amazon, or EPM/project-lead experience on Mercor projects or other AI tranining projects. • Skills: strong data management experience (Excel, SQL) and working knowledge of coding. • Leadership: prior people management experience or demonstrated ability to lead large groups effectively. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$45 - $70 / hourOpen / Referral verified
Policy & Safety / Remote

AI Safety Practitioner

We are seeking experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback. — Responsibilities • Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality. • Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains. • Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking. • Identify unsafe outputs, hallucinations, reasoning failures, and policy violations. • Provide structured feedback to improve model alignment and safety performance. • Collaborate with AI researchers and safety teams on ongoing evaluation initiatives. — Required Qualifications • Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline. • 5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field. • Excellent written English, critical thinking, and analytical reasoning skills. • Ability to consistently evaluate nuanced and policy-sensitive scenarios. — Preferred Qualifications • Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation. • Familiarity with safety policies, content moderation, or evaluation rubric development. • Experience reviewing complex, high-risk, or ambiguous content. — Why Join? • Shape the safety and behaviour of frontier AI models used by millions worldwide. • Work on challenging, real-world safety evaluations across nuanced and high-impact domains. • Collaborate with leading AI researchers, engineers, and safety teams.

$60 - $70 / hourOpen / Referral verified
Business / Remote

IT Services & Consulting Expert

Role Overview • Mercor is seeking senior IT services and consulting professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise consulting and technology delivery contexts. • The workflows are calibrated to the engagement scale, stakeholder complexity, and delivery stakes of Fortune 500 and large public company consulting engagements. • Contributors design enterprise consulting scenarios, draft reference outputs, and write rubrics that capture how senior F500 consulting and delivery leaders think. — Key Responsibilities • Construct enterprise consulting scenarios spanning $1M+ engagement values, multi-stakeholder steering committees, and complex statement-of-work (SOW) and procurement cycles at F500 accounts. • Build consulting tasks across F500 digital transformation strategy, systems integration, managed services delivery, IT advisory, and technology change management. • Develop delivery and engagement scenarios involving tools such as Jira/Confluence, ServiceNow, SAP, Salesforce, and enterprise PMO platforms in F500 client environments. • Apply enterprise consulting methodologies (Agile/SAFe delivery, business case development, RACI/governance models) and produce reference engagement plans, account strategies, and executive-level deliverables. • Author rubrics that distinguish authentic enterprise consulting judgment from generic framework recall. — Ideal Qualifications • 5+ years in IT consulting, systems integration, or managed services delivery at a Fortune 500 technology or consulting firm (Accenture, Deloitte, IBM, Capgemini, Cognizant, TCS) or inside an F500 enterprise IT/transformation organization (JPMorgan, UPS, Unilever, Microsoft, PepsiCo). • Direct ownership of F500 client engagements, F500 delivery programs, or F500 transformation initiatives. • Fluency in enterprise consulting methodologies and delivery tooling, plus understanding of how F500 budgets, procurement, and legal/SOW review actually work. • Prior rubric, consulting-enablement curriculum, or training-content authorship is a plus.

$70 - $80 / hourOpen / Referral verified
Language / Remote

Generalist (Macbook User)

Role Overview: — Mercor is looking for detail-oriented individuals with 2-3 years of experience in STEM, non-STEM fields, or currently enrolled postgrad college students to support a research project with a leading AI lab. You will help benchmark and improve cutting-edge AI models. — Qualifications: • Required: All experts on this project must have a Macbook device with an ‘M’ series chip to perform tasks on. • Currently pursuing or recently completed Masters studies, or have 2-3 years of relevant experience + Bachelors from a prestigious institution • Strong online research and communication skills • Ability to gather and clearly summarize information from diverse sources • Excellent written communication skills • Interdisciplinary degree is a nice to have • Must have a Macbook with one of the following requirements: • Apple Silicon (ARM) Macs • M‑series MacBook Pro, or • MacBook Air models running macOS 15 or higher — Job Details: • Part-time commitment of approximately 10-20 hours per week • Responsibilities include creating high-quality research questions, answers, and evaluation materials to train advanced language models • Clearly structured responsibilities without the high pressure typically found in internships or full-time consulting roles • Immediate start preferred — Application and Onboarding Process: • Submit your resume, followed by a brief (15-minute) conversation with our AI interviewer to assess research and reasoning skills • Complete a brief paid assessment to further evaluate your fit for the role • You will receive follow-up communication within a few days regarding your application status and next steps

$50 - $60 / hourOpen / Referral verified
Code / Remote — US-based

Atomistic & Surface Modeling Experts (Computational Materials & Catalysis)

Mercor is seeking computational scientists specializing in atomistic and surface modeling to support a frontier AI research lab building models for materials science and the physical sciences. This is hands-on, expert-level work: you'll apply deep, specialized knowledge to generate, structure, and evaluate the scientific data these models learn from — and your input will directly shape how advanced models reason about materials, surfaces, and chemical processes. — Key Responsibilities: • Contribute domain expertise across first-principles and molecular simulation — electronic structure, surface and interface modeling, adsorption, and reaction energetics — to build high-quality training and evaluation data. • Review and evaluate AI-generated scientific reasoning, catching errors and improving technical accuracy. • Design and solve challenging, expert-level problems in atomistic and surface modeling. • Rate and rank model outputs against defined scientific criteria, with clear written reasoning. • Structure technical knowledge — simulation setups, methods, and results — into well-organized, model-ready data. • Deliver reliable, high-quality work within defined timelines. — You're a strong fit if you have: • Hands-on experience with atomistic modeling using first-principles or molecular methods (DFT, ab initio molecular dynamics, classical MD, or Monte Carlo). • Experience modeling surfaces, interfaces, and adsorption or reaction phenomena (slab models, surface reconstructions, transition states, NEB, microkinetics). • Experience modeling semiconductor-relevant materials, or a background in computational (heterogeneous) catalysis. • Proficiency with standard tooling (e.g., VASP, Quantum ESPRESSO, CP2K, GPAW, LAMMPS, ASE, pymatgen). • A PhD in materials science, chemistry, physics, chemical engineering, or a related field, ideally with several years of research experience beyond the PhD. • Clear written English and the ability to explain technical reasoning concisely. — Role Details: • Type: Long-term, ongoing engagement • Engagement: Up to 40 hours/week (minimum 10) • Work arrangement: Remote (US-based)

$84 / hourOpen / Referral verified
Code / United States Remote

Software Engineer, Full Stack (Python, Java, Rust, C#, C++)

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Full-Stack Software Engineers with hands-on production experience across modern application stacks — Python, Java, Rust, C#, C++, and TypeScript — to build, integrate, and stress-test real software on top of frontier models before they ship. Engineers work directly with the AI lab's engineering manager on fast-moving projects and are expected to onboard quickly and contribute working code early. This is a full-time commitment of 40 hours per week. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Build and ship full-stack applications, services, and internal tools — in Python, Java, Rust, C#, C++, or TypeScript — that exercise frontier model capabilities end to end. • Integrate pre-release model APIs into working software, including the surrounding scaffolding: tool interfaces, evaluation harnesses, and telemetry. • Diagnose and document model and integration failure modes surfaced while building, translating engineering observations into clear written feedback for the research team. • Prototype quickly against shifting requirements, delivering working increments with limited up-front specification. • Collaborate with the engineering manager and other engineers to maintain consistency in code quality, architecture, and technical documentation. — 3. Core Qualifications • 3+ years of dedicated professional experience building and shipping production software at a recognized, top-tier organization. • Production experience in at least one of Python, Java, Rust, C#, or C++, plus working proficiency in a second language — this cohort is intentionally staffed across multiple language ecosystems. • Hands-on experience across the stack: backend services and APIs, a modern front-end framework (React or equivalent), relational or document databases, and cloud deployment. • Demonstrated ability to onboard onto an unfamiliar codebase and deliver working code quickly with minimal ramp-up. • Demonstrable career progression. • Ability to engage reliably for at least 40 hours/week during weekdays. • Strong written communication skills and the ability to explain complex technical decisions clearly. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$50 - $65 / hourOpen / Referral verified
Language / United States Remote

AI Rater Guidelines Writer (Linguist / Instructional Designer)

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Linguists and Instructional Designers with deep experience translating ambiguous program requirements into clear, unambiguous rater guidelines to bring precision and consistency to our AI training data — working across domains from finance and retail to insurance, legal, and sports. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide research and program teams to close the gap between ambiguous program requirements and rater-ready instructions across a range of subject-matter domains. • Design clear, non-contradictory rating guidelines and rubrics that raters can apply consistently, including under edge cases. • Evaluate draft guidelines for ambiguity, internal contradiction, and coverage gaps, and revise until raters can apply them without escalation. • Translate program specifications from domains such as finance, retail, insurance, legal, and sports into precise, discipline-specific rater instructions. • Collaborate with other subject matter experts and program leads to ensure consistency and accuracy across guideline sets. — 3. Core Qualifications • 3+ years of professional experience in linguistics, instructional design, technical writing, or a closely related field, with direct experience writing or refining guidelines/rubrics for human raters in a GenAI/RLHF context. • Demonstrated ability to work across multiple subject-matter domains (e.g., finance, retail, insurance, legal, sports) and translate domain-specific nuance into clear, unambiguous instructions. • Strong track record of resolving ambiguity and contradiction in written specifications — able to point to concrete before/after examples. • Demonstrable career progression. • Ability to engage reliably for at least 35 hours/week during weekdays. • Strong written communication skills and the ability to explain complex or nuanced guidance clearly and precisely. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$45 - $65 / hourOpen / Referral verified
Code / Remote

Inorganic Materials, Semiconductor & Superconductor Experts

Mercor is seeking experimental scientists and engineers across inorganic synthesis, characterization, superconductors, and semiconductors (including advanced packaging) to support a frontier AI research lab building models for materials science and the physical sciences. This is hands-on, expert-level work: you'll apply deep, specialized knowledge to generate, structure, and evaluate the scientific data these models learn from — and your input will directly shape how advanced models reason about materials, devices, and processes. — Key Responsibilities: • Contribute domain expertise across synthesis, characterization, fabrication, and device physics to build high-quality training and evaluation data. • Review and evaluate AI-generated scientific reasoning, catching errors and improving technical accuracy. • Design and solve challenging, expert-level problems in your area of specialization. • Rate and rank model outputs against defined scientific criteria, with clear written reasoning. • Structure technical knowledge — experimental procedures, characterization results, process data — into well-organized, model-ready data. • Deliver reliable, high-quality work within defined timelines. — You're a strong fit if you have: • Hands-on experimental experience in one or more of: inorganic synthesis (solid-state, solution, solvothermal, sol-gel), superconducting materials, or semiconductors and advanced packaging. • Strong materials or device characterization skills (XRD, SEM, TEM, spectroscopy, electrical/transport measurements). • Experience with thin-film growth or device fabrication (MBE/epitaxy, MOCVD, CVD, sputtering, IBAD, etch, clean-room microfabrication) — a plus. • An advanced degree (PhD/MS) or equivalent hands-on experience in materials science, chemistry, physics, or a related engineering field. • Clear written English and the ability to explain technical reasoning concisely. — Role Details: • Type: Long-term, ongoing engagement • Engagement: Up to 40 hours/week (minimum 10) • Work arrangement: Remote (US-based)

$84 / hourOpen / Referral verified
Medical / Remote

Applied Psychology Benchmark Specialist

Role Overview — We are seeking expert psychologists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core psychology domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of psychology expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Psychology Domains Covered — Neuromorphic Engineering, Human Factors and Engineering Psychology, Consumer and Market Psychology, Psychometrics, Digital Health, Psychedelic-assisted Therapy. — Key Responsibilities • Author original psychology questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise reasoning in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD, PsyD, or doctoral candidate in Psychology or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level psychological theory, research methodology, and empirical literature • Clinical licensure or research publications in psychology is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$50 - $63 / hourOpen / Referral verified
STEM / Remote

Applied Philosophy Benchmark Specialist

Role Overview — We are seeking expert philosophers to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core philosophy domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of philosophy expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Philosophy Domains Covered — Formal Ontology & Knowledge Representation, AI Ethics, Applied Epistemology, Philosophy of Technology & Robotics, Philosophy of Science. — Key Responsibilities • Author original philosophy questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise reasoning in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Philosophy or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of philosophical argumentation, formal logic, and canonical texts across traditions • Research publications or teaching experience in philosophy is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$50 - $63 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Korean

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Korean and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Korean, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Korean music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Korean genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Korean • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$35 - $62 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Japanese

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Japanese and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Japanese, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Japanese music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Japanese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Japanese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$35 - $62 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Korean

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Korean and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Korean, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Korean genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Korean • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$35 - $62 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Norwegian

Location: Remote — Fluent Language Skills Required: English & Norwegian. Native fluency in English and Norwegian is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$48 - $62 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Danish

Location: Remote — Fluent Language Skills Required: English & Danish. Native fluency in English and Danish is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$48 - $62 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Finnish

Location: Remote — Fluent Language Skills Required: English & Finnish. Native fluency in English and Finnish is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$48 - $62 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Swedish

Location: Remote — Fluent Language Skills Required: English & Swedish. Native fluency in English and Swedish is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$48 - $62 / hourOpen / Referral verified
Language / Remote

AI Safety Experts — English & Dutch

Location: Remote — Fluent Language Skills Required: English & Dutch. Native fluency in English and Dutch is required for this position. — Why This Role Exists — At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers. — This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated. — What You’ll Do • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent • Document reproducibly: produce reports, datasets, and attack cases customers can act on — Who You Are • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing) • You’re curious and adversarial: you instinctively push systems to breaking points • You’re structured: you use frameworks or benchmarks, not just random hacks • You’re communicative: you explain risks clearly to technical and non-technical stakeholders • You’re adaptable: thrive on moving across projects and customers — Nice-to-Have Specialties • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction • Cybersecurity: penetration testing, exploit development, reverse engineering • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing • Creative probing: psychology, acting, writing for unconventional adversarial thinking — What Success Looks Like • You uncover vulnerabilities automated tests miss • You deliver reproducible artifacts that strengthen customer AI systems • Evaluation coverage expands: more scenarios tested, fewer surprises in production • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary — Why Join Mercor • Build experience in human data-driven AI red teaming at the frontier of safety • Play a direct role in making AI systems more robust, safe, and trustworthy

$48 - $62 / hourOpen / Referral verified
Finance / Remote

Fraud Detection Experts

Role Overview • Mercor is seeking senior fraud detection professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise fraud and financial crimes contexts. • The workflows are calibrated to the transaction scale, adversarial sophistication, and financial-loss stakes of Fortune 500 banks, payment companies, and large enterprises. • Contributors design enterprise fraud detection scenarios, draft reference outputs, and write rubrics that capture how senior F500 fraud and financial crimes operators think. — Key Responsibilities • Construct enterprise fraud scenarios spanning large-scale transaction monitoring, multi-stakeholder investigation workflows, and complex regulatory reporting cycles (SARs, CTRs) at F500 institutions. • Build tasks across F500 payment fraud, account takeover and identity theft, anti-money laundering (AML), synthetic identity fraud, and merchant/card network fraud. • Develop fraud operations scenarios involving tools such as SAS Fraud Management, Actimize, Feedzai, and enterprise case management/transaction monitoring platforms in F500 stacks. • Apply enterprise fraud detection methodologies (rules-based and ML-driven detection models, network/link analysis, KYC/AML frameworks, BSA compliance) and produce reference investigation reports, risk models, and executive-level fraud narratives. • Author rubrics that distinguish authentic enterprise fraud investigation judgment from generic textbook or certification-exam-level recall. — Ideal Qualifications • 5+ years working in fraud detection, financial crimes investigation, or AML compliance at a Fortune 500 bank, payments company, or financial institution (JPMorgan Chase, Visa, PayPal, American Express, Wells Fargo). • Direct ownership of F500-scale fraud detection programs, investigation caseloads, or transaction monitoring systems. • Fluency in enterprise fraud detection tooling and methodologies, plus understanding of how F500 regulatory reporting (FinCEN, BSA/AML), law enforcement coordination, and case escalation actually work. • Prior rubric, fraud-investigation training curriculum, or case documentation authorship is a plus.

$60 - $70 / hourOpen / Referral verified
Finance / Remote

Retail Banking Expert

Role Overview • Mercor is seeking senior retail banking professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise consumer banking contexts. • The workflows are calibrated to the customer volume, regulatory complexity, and operational stakes of Fortune 500 and large national retail banks. • Contributors design enterprise retail banking scenarios, draft reference outputs, and write rubrics that capture how senior F500 retail banking operators think. — Key Responsibilities • Construct enterprise retail banking scenarios spanning large-scale deposit and lending operations, multi-stakeholder branch network decisions, and complex regulatory examination or compliance review cycles. • Build tasks across F500 consumer lending (mortgage, auto, personal), branch and digital banking operations, credit risk and underwriting, fraud prevention, and customer experience/retention strategy. • Develop retail banking operations scenarios involving tools such as core banking platforms (FIS, Fiserv, Temenos), loan origination systems (LOS), CRM platforms, and fraud detection/AML monitoring tools in F500 stacks. • Apply enterprise retail banking frameworks (credit risk scoring models, CFPB/regulatory compliance standards, omnichannel service design, KYC/AML protocols) and produce reference lending policies, risk assessments, and executive-level operational narratives. • Author rubrics that distinguish authentic enterprise retail banking judgment from generic textbook or licensing-exam-level recall. — Ideal Qualifications • 5+ years working in consumer lending, branch operations, credit risk, or compliance at a Fortune 500 retail bank (JPMorgan Chase, Bank of America, Wells Fargo, Citi, U.S. Bank). • Direct ownership of F500-scale lending portfolios, branch/digital banking operations, or regulatory compliance programs. • Fluency in enterprise retail banking tooling and frameworks, plus understanding of how F500 regulatory examinations (OCC, CFPB, FDIC), KYC/AML compliance, and credit risk governance actually work. • Prior rubric, banking-training curriculum, or policy/risk documentation authorship is a plus.

$60 - $70 / hourOpen / Referral verified
STEM / Remote

Education Expert - Sourcing Funnel (Private)

Role Overview — Mercor is collaborating with a leading AI lab to engage experienced educators across all areas of practice — including K-12 and higher-education teaching, curriculum development and instructional design, assessment and psychometrics, special education, tutoring and academic support, and educational technology. Contributors help build AI systems that reason about real educational work by translating everyday teaching, assessment, and instructional-design workflows, judgments, and decision-making into structured, high-quality training data. — Key Responsibilities — \- Design realistic educational scenarios and tasks drawn from your day-to-day work (e.g., lesson planning, assessment and item writing, grading against a rubric, differentiation and intervention, curriculum and unit design, IEPs and accommodations, student feedback) — \- Review and compare AI-generated educational outputs for accuracy, standards alignment (e.g., Common Core, NGSS, state or discipline standards), grade-level appropriateness, and sound pedagogical judgment — \- Create structured examples that reflect how educators actually reason through problems — \- Provide clear written feedback that improves how AI performs teaching and instructional tasks — \- Collaborate asynchronously with the research team — Ideal Qualifications — \- 3+ years of professional experience in education (K-12 or higher-ed teaching, curriculum/instructional design, assessment, special education, tutoring, or a related field) — \- A state teaching license/certification, National Board Certification, or an advanced degree (MEd/EdD/PhD or a subject master's) preferred, but not required — \- Bachelor's degree in Education or a subject-matter field — \- Comfortable with common classroom and instructional tools (e.g., Google Classroom, Canvas, an SIS such as PowerSchool, assessment and curriculum platforms) — \- Strong written communication and attention to detail — More About the Opportunity — \- Open to all education specialties and grade bands — contribute where your expertise is strongest — \- Work spans task design, evaluation, and structured feedback on AI educational outputs — \- Strong contributors advance into reviewer, lead, and domain-expert roles — Application Process — \- Submit a resume or a short summary of your teaching or education experience — \- Complete a short form on your field, specialties, and credentials — \- Selected applicants may complete a brief sample task — \- Follow-up typically provided within a few days

$45 - $60 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Dutch

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Dutch and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Dutch, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Dutch music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Dutch genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Dutch • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$28 - $60 / hourOpen / Referral verified
Business / Remote

Google Workspace & Business Profile Owners

Participant Qualifications — To ensure the integrity and quality of our insights, we are seeking contributors who meet the following professional criteria: — Core Requirement: Profile Ownership — Applicants must be active owners or administrators of a verified Google Business Profile (GBP). We are looking for individuals who engage with the platform regularly to manage their online presence, respond to inquiries, or update business information. — Experience & Tenure • Technical Expertise: No specialized technical background is required beyond a functional understanding of managing your own business listing. • Account Maturity: To provide the most valuable data, we prefer accounts with at least one year of active history. However, we welcome applications from newer business owners who demonstrate consistent profile engagement. — Industry Representation — We aim to cultivate a diverse rater pool that reflects the current marketplace. We are actively seeking representatives from the following sectors: • Hospitality & Lodging: Hotels, B&Bs, and specialized accommodations. • Food & Beverage: Restaurants, cafes, and catering services. • Retail & Shopping: Boutique storefronts and specialized commerce. • Professional Services: Trade services, consulting, and consumer-facing agencies. — _Note: While we strive for a broad distribution across these verticals, there is no strict quota per industry. All eligible business owners are encouraged to apply._

$0 - $60 / hourOpen / Referral verified
Finance / United States Remote

Finance Program Coordinator — AI Training Data Operations

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking a talented Program Coordinator with a finance background to bring hands-on coordination and organizational rigor to our AI training data program, keeping finance-specific tasks and rater teams on track and consistent. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. Key Responsibilities • Guide day-to-day coordination between finance-domain raters and the broader training-data program to close execution gaps and keep work on schedule. • Design and maintain task-tracking, escalation, and quality-check workflows specific to finance evaluation tasks. • Evaluate rater throughput and quality signals on finance tasks and provide clear, written status updates and feedback. • Triage and resolve finance-specific rater questions, escalating ambiguous cases to subject matter experts. • Collaborate with subject matter experts and program leads to ensure consistency and accuracy across finance training data. — 3. Core Qualifications • 5–10 years of professional experience in finance or finance operations, with a track record of coordinating cross-functional teams or projects. • Direct experience managing or coordinating a team of contributors/reviewers (raters, analysts, or similar) against deadlines and quality bars. • Demonstrable career progression (e.g., Analyst → Senior Analyst → Program/Project Coordinator). • Ability to engage reliably for at least 35 hours/week during weekdays. • Strong verbal and written communication skills, organizational skills, and problem-solving skills. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$40 - $60 / hourOpen / Referral verified
Finance / Remote

Expert Project Manager

We're hiring an Expert Project Manager to help run projects that push the frontier of LLM browsing capabilities. Day to day: tracking project and annotator performance in Google Sheets, reviewing annotator quality, handling contributor communications, and more. We're looking for someone with high agency, strong organization, an analytical mindset, and real investment in seeing LLM training projects succeed. Ops or project management background preferred, but the bigger thing is being able to pick up an unfamiliar problem and own it start to finish.

$50 - $60 / hourOpen / Referral verified
Finance / Remote

Supply Chain Expert

Role Overview • Mercor is seeking senior supply chain professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise supply chain contexts. • The workflows are calibrated to the network complexity, volume scale, and operational stakes of Fortune 500 and large public company supply chain operations. • Contributors design enterprise supply chain scenarios, draft reference outputs, and write rubrics that capture how senior F500 supply chain operators think. — Key Responsibilities • Construct enterprise supply chain scenarios spanning large-scale demand planning, multi-tier supplier networks, and complex logistics or procurement cycles at F500 accounts. • Build supply chain tasks across F500 sourcing and procurement strategy, inventory optimization, distribution and logistics network design, supplier risk management, and S&OP (sales and operations planning). • Develop supply chain and operations scenarios involving tools such as SAP, Oracle SCM, Blue Yonder, Kinaxis, and enterprise TMS/WMS platforms in F500 stacks. • Apply enterprise supply chain methodologies (Lean/Six Sigma, network optimization modeling, risk-adjusted sourcing strategies) and produce reference sourcing plans, network strategies, and executive-level supply chain briefings. • Author rubrics that distinguish authentic enterprise supply chain judgment from generic textbook or framework recall. — Ideal Qualifications • 5+ years working in supply chain, procurement, or logistics at a Fortune 500 manufacturer, retailer, or logistics provider (Amazon, Walmart, UPS, Procter & Gamble, Unilever) or inside an F500 supply chain/operations organization. • Direct ownership of F500 supplier relationships, F500 distribution networks, or F500 procurement programs. • Fluency in enterprise supply chain tooling and methodologies, plus understanding of how F500 sourcing, procurement, and vendor contract review actually work. • Prior rubric, supply chain training curriculum, or operations documentation authorship is a plus.

$60 - $70 / hourOpen / Referral verified
Math / Remote

Data Science and Analytics Experts

Role Overview • Mercor is seeking senior data science and analytics professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise data and analytics contexts. • The workflows are calibrated to the data scale, model complexity, and business-critical stakes of Fortune 500 and large public company data operations. • Contributors design enterprise data science scenarios, draft reference outputs, and write rubrics that capture how senior F500 data leaders think. — Key Responsibilities • Construct enterprise data science scenarios spanning large-scale predictive modeling, multi-stakeholder analytics governance, and complex data infrastructure decisions at F500 accounts. • Build analytics tasks across F500 machine learning model development, enterprise data pipelines, business intelligence at scale, experimentation/causal inference, and data strategy. • Develop data and MLOps scenarios involving tools such as Snowflake, Databricks, Python/R, SQL, Tableau/Power BI, and enterprise ML platforms (SageMaker, Vertex AI, MLflow) in F500 stacks. • Apply enterprise data science methodologies (statistical rigor, A/B testing frameworks, model validation, MLOps best practices) and produce reference analyses, model documentation, and executive-level insights. • Author rubrics that distinguish authentic enterprise data science judgment from generic textbook or tutorial-level recall. — Ideal Qualifications • 5+ years working as a data scientist, analytics leader, or ML engineer at a Fortune 500 technology or enterprise organization (Google, Meta, Amazon, Microsoft, Netflix) or inside an F500 data/analytics organization (JPMorgan, UPS, Unilever, PepsiCo, Walmart). • Direct ownership of F500 data products, F500 analytics initiatives, or F500 machine learning systems in production. • Fluency in enterprise data science tooling and methodologies, plus understanding of how F500 data governance, privacy compliance, and cross-functional stakeholder alignment actually work. • Prior rubric, technical curriculum, or model documentation authorship is a plus.

$60 - $70 / hourOpen / Referral verified
Business / Remote

Higher Education Expert

Role Overview • Mercor is seeking senior higher education professionals to build evaluation tasks for AI systems operating in large university and higher education institution contexts. • The workflows are calibrated to the academic complexity, stakeholder diversity, and institutional stakes of large research universities and higher education systems. • Contributors design higher education scenarios, draft reference outputs, and write rubrics that capture how senior faculty and university administrators think. — Key Responsibilities • Construct higher education scenarios spanning curriculum and program design, multi-stakeholder faculty senate or accreditation review, and complex institutional budget or enrollment planning cycles. • Build education tasks across academic program development, student affairs and retention, research administration and grant compliance, faculty governance, and institutional advancement/fundraising. • Develop higher-ed operations scenarios involving tools such as Banner/Workday Student, Canvas/Blackboard, CRM platforms for admissions and advancement, and research administration systems (Cayuse, InfoEd). • Apply higher education frameworks (accreditation standards, shared governance models, learning outcomes assessment, enrollment management strategy) and produce reference academic plans, accreditation documentation, and administrator/board-level communications. • Author rubrics that distinguish authentic higher education judgment from generic academic or policy-manual recall. — Ideal Qualifications • 5+ years working as faculty, an academic administrator, or in a senior operational role at a college, university, or higher education system. • Direct ownership of academic programs, enrollment/retention initiatives, research administration, or institutional strategic priorities. • Fluency in higher education tooling and frameworks, plus understanding of how university budgets, accreditation compliance, and shared governance actually work. • Prior rubric, curriculum-design, or faculty-training content authorship is a plus.

$60 - $70 / hourOpen / Referral verified
Code / Remote

Moonlight MCQA - Norwegian Generalist

Write original general-knowledge multiple-choice questions in Norwegian for an AI evaluation dataset. — You will draft exam-style questions from your own expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: master's or PhD, 2 to 6 years of professional experience, educated or based in Norway, native-level Norwegian.

$48.51 - $59.29 / hourOpen / Referral verified
Legal / Remote

Moonlight MCQA - Norwegian Law

Write original multiple-choice questions on Norwegian law for an AI evaluation dataset. — You will draft exam-style questions from your own professional expertise, each with ten answer options and a worked solution. Remote, flexible hours. — Requirements: law degree or Norwegian equivalent, ideally 2+ years practising, educated or based in Norway, native-level Norwegian.

$48.51 - $59.29 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - German

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in German and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in German, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the German music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary German genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in German • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$30 - $58 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Swedish

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Swedish and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Swedish, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Swedish genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Swedish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$42 - $78 / hourOpen / Referral verified
Business / Remote

Sales and Marketing Expert

Role Overview • Mercor is seeking senior sales and marketing professionals to build evaluation tasks for AI systems operating in Fortune 500 go-to-market contexts. • The workflows are calibrated to the deal sizes, stakeholder complexity, and brand stakes of Fortune 500 and large public companies. • Contributors design enterprise GTM scenarios, draft reference outputs, and write rubrics that capture how senior F500 operators think. — Key Responsibilities • Construct enterprise sales scenarios spanning $1M+ ACV deals, multi-stakeholder buying committees, and complex procurement cycles at F500 accounts. • Build marketing tasks across F500 brand strategy, enterprise ABM, demand generation at scale, lifecycle, and category positioning. • Develop RevOps and GTM scenarios involving Salesforce Enterprise, Marketo, 6sense, Gong, and Outreach in F500 stacks. • Apply enterprise sales methodologies (MEDDIC, Challenger, Force Management) and produce reference deal strategies, account plans, and executive narratives. • Author rubrics that distinguish authentic enterprise GTM judgment from generic playbook recall. — Ideal Qualifications • 5+ years selling, marketing, or running RevOps at a Fortune 500 enterprise software vendor (Salesforce, Oracle, ServiceNow, SAP, Workday, Microsoft, AWS) or inside an F500 brand or marketing organization (P&G, JPMorgan, Unilever, Microsoft, PepsiCo). • Direct ownership of F500 accounts, F500 brand campaigns, or F500 demand programs. • Fluency in enterprise GTM tooling and methodologies, plus understanding of how F500 budgets, procurement, and legal review actually work. • Prior rubric, sales-enablement curriculum, or training-content authorship is a plus.

$60 - $70 / hourOpen / Referral verified
STEM / Remote

Applied History & Political Science Benchmark Specialist

Role Overview — We are seeking experts in history and political science to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core history and political science domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — History & Political Science Domains Covered — National Security, Public Policy, Business History, Environmental History, Latin American History. — Key Responsibilities • Author original history and political science questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise reasoning in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in History, Political Science, International Relations, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of historiographical methods, political theory, and comparative analysis • Research publications or policy experience is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$44 - $56 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Spanish (MX)

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Spanish (MX) and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Spanish (MX), and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Spanish (MX) music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Spanish (MX) genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Spanish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$13 - $54 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Spanish (ES)

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Spanish (ES) and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Spanish (ES), and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Spanish (ES) music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Spanish (ES) genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Spanish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$39 - $54 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - Mandarin Chinese

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Mandarin Chinese and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in Mandarin Chinese, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the Mandarin Chinese music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary Mandarin Chinese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Mandarin Chinese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$18 - $54 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - French

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in French and English. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Native or near-native proficiency in French, and the ability to follow detailed written instructions in English • Experience as a songwriter, lyricist, performer, composer, or music journalist, with knowledge of the French music scene • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary French genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in French • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $54 / hourOpen / Referral verified
Language / Remote

Music & Lyrics Expert - English (US/UK/Australia)

Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards. — Key Responsibilities • Compare AI-generated lyrics with published songs to identify similarities • Rate lyrics based on quality, creativity, prompt adherence, and originality • Evaluate whether lyrics sound natural, including word choice, slang, and regional expressions. — Requirements • Experience as a songwriter, lyricist, performer, composer, or music journalist • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or published songwriting work • Broad familiarity with contemporary genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Domain Expert Interview • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$23 - $54 / hourOpen / Referral verified
Language / Remote

Music Production Expert - French

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in French and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in French, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary French genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in French • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$17 - $54 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Spanish (MX)

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Spanish (MX) and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Spanish (MX), and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Spanish (MX) genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Spanish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$13 - $54 / hourOpen / Referral verified
Medical / Remote

Healthcare Administrative Specialist

About the Role — We're looking for professionals with deep, hands-on experience in the administrative and operational work that keeps healthcare organizations running — the "back office" of a medical practice, hospital, or health plan. This spans the full administrative lifecycle: getting patients registered and covered, securing authorizations, moving claims through to payment, managing provider and payer relationships, and keeping records and compliance in order. You'll help us understand the precise steps these workflows take in the real tools you use every day. — Core Workflows (you should actively own one or more) • Patient Access & Front-End: registration, scheduling, insurance/eligibility verification, benefits interpretation, intake, and referral management • Prior Authorizations & Utilization Support: initiating, submitting, and tracking authorizations through payer portals • Claims & Revenue Cycle: claim submission, status tracking, denials and appeals, A/R follow-up, and resolution of rejections • Remittance & Payment Posting: electronic remittance advice (ERA), payment posting, and reconciliation against clinical charges • Provider & Payer Operations: credentialing, provider enrollment, payer configuration, appeals & grievances, and member/provider services • Health Information & Practice Administration: medical records administration, release of information, and day-to-day practice/office operations — Requirements • 3+ years of daily, hands-on experience in a healthcare administrative / back-office role • Direct experience with the systems this work runs on — payer portals (Availity, Optum/Change Healthcare, Waystar, Office Ally, or individual payer hubs like UHC Link, BCBS, Aetna), EHR / practice-management platforms, and/or clearinghouses • Currently working in a healthcare administrative role • Based in the United States — Ideal Backgrounds • Practice Administrator / Medical Office Manager • Revenue Cycle Specialist / RCM Analyst • Medical Billing Coordinator / Full-Cycle Biller • Patient Access Lead / Intake / Referral Coordinator • Prior Authorization Specialist • Denials / Claims Resolution Specialist • Credentialing / Provider Enrollment Specialist • Health Plan / Payer Operations (claims, appeals & grievances, provider configuration) • Health Information Management (HIM) / Medical Records administrator — Note: We are looking for administrative and operational workflow expertise across healthcare — not medical coding or clinical care roles.

$40 - $50 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Canadian French)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Canada-based voice actors with native Canadian French fluency and international French (neutral, accent-free) delivery to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. — Your voice recordings will be used exclusively for the client's internal CX AI Agent. They will not be sold, licensed, or reused for any other product, dataset, or purpose. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native Canadian French speaker currently based in Canada, with the ability to speak international French — neutral, without strong regional accent • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability for 5-10 hours per week over the project duration • \[IMP\]: Your voice may be cloned for the client's CX AI Agent so please only apply if you are okay with voice cloning. Your recordings will not be sold or reused outside of this specific use case. — Preferred Qualifications • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 / hourOpen / Referral verified
Code / Remote

Audio Engineer (Speech / TTS Audio Specialist) – German

Overview We're looking for audio engineers with strong experience editing and quality-checking voice recordings for speech and text-to-speech (TTS) applications, specifically for German-language audio. This role focuses on preparing high-quality audio datasets for machine learning systems, ensuring clarity, consistency, and technical compliance across large volumes of German speech data. Native or near-native German fluency is required to accurately assess speech quality, pronunciation, and naturalness. • * * — What You'll Do • Edit and process German voice recordings for use in TTS and speech AI systems • Perform audio cleanup including silence trimming, breath reduction/removal, de-clicking, de-essing, and noise reduction • Conduct detailed quality assurance checks for: • Clipping and distortion • Background noise and artifacts • Loudness consistency and level balancing • Ensure audio meets strict technical specifications for ML training pipelines • Evaluate recordings for German pronunciation accuracy, naturalness, and fluency • Work with large batches of recordings, maintaining consistency and throughput • Collaborate with teams working on German speech datasets, voice talent recordings, and model training workflows • * * — Basic Qualifications • Native or near-native German speaker with strong listening intuition for German speech nuances • Experience editing speech or voice recordings (not just music production) • Strong understanding of speech audio quality standards and common issues in recorded dialogue • Hands-on experience with tools like iZotope RX, Pro Tools, or similar audio restoration software • Ability to identify and fix artifacts, inconsistencies, and technical defects in audio • Experience working with high-volume audio datasets or structured workflows • * * — Preferred Qualifications • Experience working with TTS vendors or platforms • Prior involvement in speech ML, voice AI, or dataset creation/annotation workflows • Familiarity with audio QA pipelines, labeling, or evaluation processes • Experience directing or editing German voice talent recordings • * * — Ideal Candidate Profile • Detail-oriented with strong critical listening skills in German • Comfortable working with repetitive, high-precision tasks at scale • Able to balance speed and quality in production environments • Familiar with the nuances of human German speech vs synthetic voice requirements

$50 / hourOpen / Referral verified
Multimodal / Remote

Voice Actor: CX Agent Voice Cloning (Indonesian Bahasa - Female)

Mercor is partnering with a leading AI research lab to build next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking Indonesia-based female voice actors who are native Bahasa Indonesia speakers to contribute high-quality voice recordings for training and evaluating cutting-edge speech models. Fluency in English is a plus. This role is ideal for professionals with experience in voice acting, narration, or broadcast who can deliver consistent, expressive, and clean audio across a variety of scripts. • * * — Key Responsibilities • Record high-quality voice samples across diverse scripts (conversational, narrative, instructional, etc.) • Deliver clear, natural, and expressive speech with strong control over tone, pacing, and pronunciation • Maintain consistency in voice, accent, and delivery across recording sessions • Follow detailed recording guidelines (environment, microphone setup, file formatting) • Perform multiple takes with variation in emotion, emphasis, and style when required — Requirements • Native Bahasa Indonesia speaker currently based in Indonesia • Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting • Access to a professional or near-professional recording setup (quality microphone, quiet environment, pop filter, etc.) • Strong command of intonation, diction, and emotional range • Ability to follow scripts precisely while maintaining natural delivery • Reliable availability over the project duration • \[IMP\]: Your voice may be cloned for the clients CX AI Agent so please only apply if you are okay with voice cloning — Preferred Qualifications • Fluency in English in addition to Bahasa Indonesia • Experience recording for TTS, audiobooks, IVR systems, or AI voice datasets • Familiarity with audio editing tools (e.g., Audacity, Adobe Audition, Reaper) • Ability to deliver multiple vocal styles (e.g., conversational, corporate, energetic, calm) • * *

$50 / hourOpen / Referral verified
Code / Remote

Audio Engineer (Speech / TTS Audio Specialist) – French

Overview We're looking for audio engineers with strong experience editing and quality-checking voice recordings for speech and text-to-speech (TTS) applications, specifically for French-language audio. This role focuses on preparing high-quality audio datasets for machine learning systems, ensuring clarity, consistency, and technical compliance across large volumes of French speech data. Native or near-native French fluency is required to accurately assess speech quality, pronunciation, and naturalness. • * * — What You'll Do • Edit and process French voice recordings for use in TTS and speech AI systems • Perform audio cleanup including silence trimming, breath reduction/removal, de-clicking, de-essing, and noise reduction • Conduct detailed quality assurance checks for: • Clipping and distortion • Background noise and artifacts • Loudness consistency and level balancing • Ensure audio meets strict technical specifications for ML training pipelines • Evaluate recordings for French pronunciation accuracy, naturalness, and fluency • Work with large batches of recordings, maintaining consistency and throughput • Collaborate with teams working on French speech datasets, voice talent recordings, and model training workflows • * * — Basic Qualifications • Native or near-native French speaker with strong listening intuition for French speech nuances • Experience editing speech or voice recordings (not just music production) • Strong understanding of speech audio quality standards and common issues in recorded dialogue • Hands-on experience with tools like iZotope RX, Pro Tools, or similar audio restoration software • Ability to identify and fix artifacts, inconsistencies, and technical defects in audio • Experience working with high-volume audio datasets or structured workflows • * * — Preferred Qualifications • Experience working with TTS vendors or platforms • Prior involvement in speech ML, voice AI, or dataset creation/annotation workflows • Familiarity with audio QA pipelines, labeling, or evaluation processes • Experience directing or editing French voice talent recordings • * * — Ideal Candidate Profile • Detail-oriented with strong critical listening skills in French • Comfortable working with repetitive, high-precision tasks at scale • Able to balance speed and quality in production environments • Familiar with the nuances of human French speech vs synthetic voice requirements

$50 / hourOpen / Referral verified
Multimodal / Remote

Russian Audio Generalist Evaluator Expert (San Francisco Bay Area)

Mercor is seeking a Russian Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project with a leading research lab. In this role, you will work on transcription, annotation, and evaluation tasks that help train and benchmark advanced language models. This is a short-term, structured engagement ideal for candidates with strong academic or analytical backgrounds who are fluent in Russian and English and are based in the San Francisco Bay Area, enabling occasional in-person collaboration if required. — Job Responsibilities — Transcribe and Optimise Audio & Video • Listen to, analyse, and transcribe audio and video content in Russian, following detailed constraints and instructions. • Produce high-quality written outputs in Russian, with supporting work in English when required. • Ensure clarity, accuracy, and strict adherence to formatting and stylistic guidelines. • Capture nuances such as tone, intent, formal vs. informal register, regional expressions, dialectal variations, and contemporary Russian usage where relevant. — Define and Document Evaluation Standards • Establish clear expectations for correct and high-quality responses in general consumer audio contexts. • Develop detailed evaluation rubrics and grading guidelines in Russian and English. • Document standards to ensure consistency across reviewers and model evaluations. • Identify linguistic nuances, grammatical complexities, colloquialisms, and edge cases specific to Russian. — Conduct Model Testing and Grading • Run prompts through language models and assess generated outputs. • Evaluate responses against predefined criteria for accuracy, completeness, fluency, and instructional clarity. • Provide structured feedback to improve model performance in Russian audio tasks. — Support Benchmarking and Quality Assurance • Participate in QA and review cycles to ensure tasks, rubrics, and outputs meet Mercor’s quality bar. • Maintain consistency and reliability before datasets are integrated into official benchmarks. • Collaborate with project leads to resolve ambiguities and improve task design. — Minimum Qualifications • Strong writing, editing, and critical thinking skills. • Ability to work independently, manage time effectively, and meet deadlines. • Native or near-native fluency in Russian (spoken and written) and professional fluency in English. • Strong familiarity with spoken Russian, regional vocabulary, dialects, and contemporary language usage. • Ability to accurately transcribe and analyse Russian audio content across general consumer contexts. • Must be based in the San Francisco Bay Area. • Available to commit 10–20 hours per week. — Preferred Qualifications • College students or recent graduates. • Background in linguistics, humanities, social sciences, journalism, translation/localization, or technical disciplines. • Prior experience with transcription, annotation, localisation, evaluation, or research workflows in Russian. • Familiarity with regional variations of Russian and contemporary digital language usage. • Interest in AI, language models, or applied research environments. — Application & Onboarding Process • Complete a short AI-led interview (approximately 15 minutes). • If selected, you will be onboarded and invited to begin project work.

$50 / hourOpen / Referral verified
Business / Remote

Insurance Experts

Role Overview • Mercor is seeking senior insurance professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise insurance and risk contexts. • The workflows are calibrated to the underwriting complexity, regulatory stakes, and claims scale of Fortune 500 and large public insurance carriers. • Contributors design enterprise insurance scenarios, draft reference outputs, and write rubrics that capture how senior F500 insurance operators think. — Key Responsibilities • Construct enterprise insurance scenarios spanning complex commercial underwriting, multi-stakeholder claims adjudication, and regulatory filing or compliance review cycles at F500 carriers. • Build insurance tasks across F500 underwriting and risk assessment, claims management, actuarial pricing and reserving, reinsurance strategy, and regulatory/compliance affairs. • Develop insurance operations scenarios involving tools such as Guidewire, Duck Creek, actuarial modeling platforms (SAS, R), and enterprise claims/policy administration systems in F500 stacks. • Apply enterprise insurance methodologies (actuarial risk modeling, loss reserving, ERM frameworks, NAIC compliance standards) and produce reference underwriting guidelines, claims strategies, and executive-level risk narratives. • Author rubrics that distinguish authentic enterprise insurance judgment from generic textbook or licensing-exam recall. — Ideal Qualifications • 5+ years working in underwriting, claims, actuarial, or risk management at a Fortune 500 insurance carrier or reinsurer (State Farm, Allstate, Chubb, AIG, Berkshire Hathaway) or inside an F500 enterprise risk/insurance organization. • Direct ownership of F500-scale underwriting portfolios, claims operations, or actuarial pricing models. • Fluency in enterprise insurance tooling and methodologies, plus understanding of how F500 regulatory compliance, reinsurance treaties, and reserving practices actually work. • Prior rubric, actuarial/underwriting training curriculum, or claims documentation authorship is a plus.

$50 - $60 / hourOpen / Referral verified
Legal / Remote

Senior Civil Legal Paralegal / Accredited Representative

Mercor is seeking experienced Senior Civil Legal Paralegals and Accredited Representatives to review AI-generated legal guidance involving civil legal services workflows. Experts will assess procedural accuracy, documentation quality, and practical usefulness across common civil legal matters. • * * — Responsibilities • Review AI-generated legal guidance. • Evaluate procedural correctness and documentation. • Assess housing, family, consumer debt, and bankruptcy scenarios. • Identify filing errors and missing procedural steps. • Provide structured written feedback. • * * — Required Qualifications • Significant experience as a Senior Paralegal or Accredited Representative. • Strong knowledge of civil legal procedures. • Experience supporting housing, family, consumer, or debt matters. • Excellent written communication skills. • * * — Preferred Qualifications • Legal aid or nonprofit legal services experience. • Familiarity with jurisdiction-specific filing procedures. • Experience supporting underserved populations. • Experience mentoring legal support staff. • * * — Why Join Mercor? • Bring practical legal operations expertise into AI development. • Help improve AI-generated legal guidance for real-world users. • Competitive consulting rates. • Opportunity to contribute to next-generation legal technology.

$70 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Dutch

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Dutch and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Dutch, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Dutch genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Dutch • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$28 - $60 / hourOpen / Referral verified
Medical / Remote

Adult Inpatient Nurses (RN)

We're hiring experienced Adult Inpatient Nurses (RNs) to help train and evaluate AI systems used in clinical and healthcare settings. This role is ideal for nurses who want to apply their frontline experience to improve the accuracy, safety, and reliability of medical AI tools. — You'll work on projects that require deep clinical judgment, attention to detail, and the ability to translate real-world bedside documentation practices into structured feedback for AI systems. — Key Responsibilities • Review and evaluate AI-generated clinical outputs based on nursing flowsheet documentation • Validate accuracy, completeness, and adherence to current documentation standards • Annotate and structure inpatient nursing assessment data for AI training datasets • Provide expert feedback on nursing assessments and documentation practices • Identify gaps, inconsistencies, or risks in AI-generated responses • Ask clarifying questions when annotation guidance is ambiguous, and contribute to guideline refinement • Collaborate with technical teams to improve model performance • Contribute to the development of high-quality clinical benchmarks — Basic Qualifications • Active RN license (U.S., outside California) • Recent adult acute care inpatient bedside experience, ideally in a non-procedural area of specialty • Experience within the last 5–10 years, reflecting familiarity with current workflows and documentation practices (current documentation experience is prioritized over total years of nursing experience) • Comfortable performing and documenting comprehensive nursing assessments • Comfortable following detailed annotation guidelines consistently • Excellent written communication skills and responsiveness to feedback • Ability to work independently and meet deadlines — Preferred Qualifications • Epic EHR experience (most common platform, closely aligns with many workflows) • Comfortable learning new annotation tools and web-based platforms • Able to navigate transcripts efficiently and use AI-assisted tools appropriately while still verifying outputs • Experience reviewing charts for quality improvement, utilization review, CDI, or clinical informatics • Previous annotation, chart abstraction, or healthcare AI experience

$55 - $65 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Japanese

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Japanese and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Japanese, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Japanese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Japanese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$35 - $62 / hourOpen / Referral verified
Medical / Remote

Social Work Expert

Role Overview • Mercor is seeking senior social work professionals to build evaluation tasks for AI systems operating in large human services and clinical social work contexts. • The workflows are calibrated to the case complexity, stakeholder sensitivity, and client-outcome stakes of large public agencies, healthcare systems, and community-based organizations. • Contributors design social work scenarios, draft reference outputs, and write rubrics that capture how senior clinicians and case management leaders think. — Key Responsibilities • Construct social work scenarios spanning complex case assessment, multi-stakeholder care coordination, and crisis intervention or mandated reporting cycles. • Build tasks across clinical social work practice, child and family welfare, healthcare/medical social work, substance use and mental health services, and community-based case management. • Develop practice scenarios involving tools such as electronic health/case record systems (EHR, CCWIS), risk and needs assessment instruments, and interagency referral platforms. • Apply social work frameworks (biopsychosocial assessment, trauma-informed care, strengths-based practice, NASW Code of Ethics) and produce reference case plans, clinical documentation, and interdisciplinary team communications. • Author rubrics that distinguish authentic clinical and ethical judgment from generic textbook or policy-manual recall. — Ideal Qualifications • 5+ years working as a licensed clinical social worker (LCSW/LMSW) or senior case management professional at a hospital system, child welfare agency, behavioral health organization, or large nonprofit. • Direct ownership of clinical caseloads, family/child welfare cases, or program-level service delivery. • Fluency in social work practice standards and case management tooling, plus understanding of how mandated reporting, interagency coordination, and ethical/legal compliance actually work. • Prior rubric, clinical training curriculum, or case documentation authorship is a plus.

$40 - $50 / hourOpen / Referral verified
Business / Remote

K-12 Education Expert

Role Overview • Mercor is seeking senior K-12 education professionals to build evaluation tasks for AI systems operating in large school district and public education contexts. • The workflows are calibrated to the instructional complexity, stakeholder diversity, and student-outcome stakes of large public school districts and state education systems. • Contributors design K-12 education scenarios, draft reference outputs, and write rubrics that capture how senior educators and district leaders think. — Key Responsibilities • Construct K-12 scenarios spanning curriculum design and adoption, multi-stakeholder IEP/504 planning, and complex district budget or policy review cycles. • Build education tasks across instructional design, classroom differentiation and special education, assessment and standards alignment, student support services, and family/community engagement. • Develop education operations scenarios involving tools such as PowerSchool, Canvas/Google Classroom, state assessment platforms, and district-level MTSS/RTI systems. • Apply K-12 pedagogical frameworks (Universal Design for Learning, standards-based grading, data-driven instruction, restorative practices) and produce reference lesson plans, IEP documentation, and administrator-level communications. • Author rubrics that distinguish authentic K-12 educator judgment from generic pedagogical or textbook recall. — Ideal Qualifications • 5+ years teaching, administering, or leading instruction at a public or private K-12 school, district office, or state education agency. • Direct ownership of classroom instruction, IEP/special education caseloads, curriculum programs, or school/district-level initiatives. • Fluency in K-12 education tooling and frameworks, plus understanding of how school budgets, state standards compliance, and family/community engagement actually work. • Prior rubric, curriculum-writing, or teacher-training content authorship is a plus.

$40 - $50 / hourOpen / Referral verified
Language / Remote

Music Production Expert - German

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in German and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in German, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary German genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in German • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$30 - $58 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Spanish (ES)

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Spanish (ES) and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Spanish (ES), and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Spanish (ES) genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Spanish • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$39 - $54 / hourOpen / Referral verified
Language / Remote

Music Production Expert - Mandarin Chinese

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Mandarin Chinese and English. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • Native or near-native proficiency in Mandarin Chinese, and the ability to follow detailed written instructions in English • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary Mandarin Chinese genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Bilingual Competency Interview in Mandarin Chinese • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$18 - $54 / hourOpen / Referral verified
Business / United States Remote

Project Coordinator

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — Are you ready to help shape the future of artificial intelligence? Join a leading AI lab's cutting-edge GenAI team, where you'll be at the forefront of building groundbreaking AI models. We're seeking talented Project Coordinators to support and accelerate world-class AI research and data operations — acting as the operational bridge between program leadership and a growing team of domain experts. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. — 2. What You'll Do • Act as a day-to-day point of contact for AI data projects, helping keep workflows and operations running smoothly. • Work closely with program leads and domain experts — take inputs from leadership, turn them into clear guidelines, and share them with expert teams. • Help onboard and support new experts as the program grows. • Join weekly business reviews, keep notes and action items organized, and lead calls when needed. • Help spot quality issues, understand root causes, and share findings across teams. • Contribute to creating and refining project guidelines with research and product partners. • Flag potential roadblocks early and help get them resolved. • Track project deliverables using Google Sheets and Excel. • Stay flexible and adapt as project needs evolve. — 3. Qualifications • Location: Must be based in the USA. • Education: STEM background strongly preferred. • Experience: 3+ years in project coordination or project management, with the ability to coordinate large, cross-functional teams. • AI fluency: hands-on experience on AI training-data or human-data projects — as a project coordinator, team lead, EPM, or expert contributor. • Ideal: prior project coordination on AI/data programs at organizations like Meta, TikTok, or Amazon, or EPM/project-lead experience on Mercor projects or other AI tranining projects. • Skills: strong data management experience (Excel, SQL) and working knowledge of coding. • Leadership: prior people management experience or demonstrated ability to lead large groups effectively. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$45 - $55 / hourOpen / Referral verified
Language / Remote

Music Production Expert - English (US/UK/Australia)

Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. — You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards. — Key Responsibilities • Compare AI-generated songs head-to-head on musicality, creativity, prompt adherence, vocal quality, and production/mix quality • Label songs by genre, structure, instruments, vocals, and other musical characteristics • Review AI-generated lyrics and vocals against reference material for quality and accuracy — Requirements • 2+ years of experience as a music producer, audio engineer, or mixing engineer • Headphones or studio monitors suitable for critical listening — Nice to have • Formal training in music performance, theory, or composition • Credited or commercially released production, engineering, or mixing work • Broad familiarity with contemporary genres, sub-genres, and artists — Project Timeline • Start date: immediate • Duration: up to 6 months • Commitment: flexible. Most experts work around 20 hours per week; there is no cap — Application & Onboarding • Upload your resume and complete the application steps, including the Domain Expert Interview • Receive next steps and onboarding details within a few days — Apply today and put your musical expertise to work. — We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

$23 - $54 / hourOpen / Referral verified