Role directory

Remote — United States jobs

12 current, verified opportunities. Compare the requirements here, then complete applications on the partner platform.

Current opportunities

Verified listing details with no HumanitApp account or application fee.

Code / Remote — United States

AI Developer Trace Task Auditor

Evaluate the quality and correctness of AI-assisted software-development traces used to train and evaluate a frontier AI lab's models. You'll assess end-to-end coding sessions produced with AI-assisted developer tools — judging...

$70 - $90 / hourVerified / Application completed externally
Code / Remote — United States

AWS Serverless & Infrastructure-as-Code Task Auditor

Evaluate the quality, correctness, and cloud-architecture soundness of AWS serverless and infrastructure-as-code tasks used to train and evaluate a frontier AI lab's models. You'll assess multi-service serverless designs, IaC fidelity, and...

$70 - $90 / hourVerified / Application completed externally
Other / Remote (United States)

Baseball Fan – Live MLB Game AI Evaluator (US)

Role overview — Mercor is looking for US-based people with strong baseball knowledge to evaluate AI assistants live, while real MLB games are being played. On scheduled game days, you'll ask AI assistants the questions you'd naturally ask...

$100 - $120 / per-taskVerified / Application completed externally
Code / Remote — United States

CVE Vulnerability Expert

Evaluate the quality, fidelity, and completeness of vulnerability-reproduction and remediation tasks used to train and evaluate a frontier AI lab's models. You'll assess whether CVE reproductions are faithful, fixes are sound, verification...

$70 - $90 / hourVerified / Application completed externally
Language / Remote — United States

GPU Kernel Expert

Evaluate the quality, correctness, and completeness of GPU/accelerator kernel development tasks used to train and evaluate a frontier AI lab's models. You'll assess numerical correctness, performance-benchmarking fairness, task scoping,...

$70 - $90 / hourVerified / Application completed externally
Code / Remote — United States

Kubernetes Task Auditor

Evaluate the quality, correctness, and production-readiness of Kubernetes tasks used to train and evaluate a frontier AI lab's models. You'll assess cluster-operations scenarios, manifest correctness, and failure-mode troubleshooting — and...

$70 - $90 / hourVerified / Application completed externally
Medical / Remote (United States)

Medicare Advantage Brokers Survey – Agent & Enrollment Insights

Mercor is conducting a paid research study in collaboration with a leading AI research lab focused on improving healthcare and member experiences. We are seeking licensed health insurance agents and brokers who actively sell Medicare...

$140 / one-timeVerified / Application completed externally
Code / Remote — United States

ML Challenge Task Auditor

Evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate a frontier AI lab's models. You'll assess experiment design, model-selection reasoning, and evaluation methodology —...

$70 - $90 / hourVerified / Application completed externally
STEM / Remote (United States)

Stanford Youth Study — US Parents & Kids (Age 9–13)

Mercor is recruiting participants for the Youth Well-being Project, a research study led by Stanford University in collaboration with Bocconi University, focused on how smartphones and digital tools affect children and families. — The...

$110 / one-timeVerified / Application completed externally
Code / Remote — United States

SWE-Bench Task Auditor

Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading...

$70 - $90 / hourVerified / Application completed externally
Business / Remote — United States

Trainium (NKI) Kernel Expert

Evaluate the quality, correctness, and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks used to train and evaluate a frontier AI lab's models. You'll assess CUDA→NKI migration fidelity, Trainium-specific...

$70 - $90 / hourVerified / Application completed externally
Medical / Remote (United States)

Urologists & Medical Oncologists (Prostate Cancer) – Paid Research Study

Mercor is conducting a paid research study in collaboration with a leading AI research lab focused on improving healthcare. We are seeking board-certified urologists and medical oncologists in the United States who manage patients with...

$175 / one-timeVerified / Application completed externally