Role directory

Machine learning jobs

11 active, referral-verified opportunities.

Code / Remote

Data Scientist Talent Network

Mercor is building a network of experienced data scientists for potential future projects with leading AI research organizations. These projects may focus on evaluating how effectively AI systems perform real-world data science work. — There is no immediate project opening, but qualified applicants may be contacted as relevant opportunities become available. — 2. Potential Responsibilities — Future projects may involve: • Designing precise, task-specific grading criteria for data science deliverables, including exploratory data analyses, statistical modeling work, machine learning pipelines, experimentation and A/B test write-ups, feature engineering, and technical reports or notebooks • Evaluating AI-generated or human-created work against established criteria • Providing detailed written justifications for evaluations and scores • Applying consistent, evidence-based judgment so that assessments are reproducible and defensible • Incorporating structured feedback from senior reviewers and iterating on submitted work — Specific responsibilities will vary depending on the project. — 3. Ideal Qualifications • 1+ years of professional data science experience • Experience at a leading technology, research, or quantitative firm (such as top FAANG, AI labs, top-tier quant funds, or equivalent) • Strong command of Python, SQL, statistical modeling, machine learning, experimentation and causal inference, and translating messy real-world data into rigorous analyses • Exceptional written communication skills, including the ability to convey technical findings clearly • A detail-oriented and consistent approach to evaluating complex work • Comfort receiving feedback and calibrating judgment against established standards

$100 - $150 / hourOpen / Referral verified
Code / Remote

LLM Research Scientist (Pre-training & Computer Vision & Adversarial Robustness)

We're looking for experienced machine learning researchers with hands-on experience training and improving deep learning models end-to-end, across vision and language. You'll work on well-scoped empirical open-ended ML research problems. — Responsibilities • Train image classifiers and generative image models from scratch, and fine-tune open-weight language models. • Get the most out of limited data, compute, and model-size budgets. • Make models robust — to adversarial inputs and to adversarial conversations. • Compress models to meet hard size and latency constraints without sacrificing accuracy. • Diagnose and resolve training issues. — Requirements — We are looking for candidates with strong expertise in one or more of the following areas: — Adversarial Robustness — Experience with: • Adversarial training of image classifiers (e.g. PGD-based training, TRADES). • Evaluating robust accuracy under standard threat models (e.g. L∞ attacks, AutoAttack) and avoiding gradient-masking pitfalls. • Managing the robustness–accuracy trade-off and robust overfitting. — Efficient Computer Vision — Experience with: • Training image classifiers end-to-end, especially for fine-grained recognition (many visually similar classes, few examples per class). • Model compression: quantization, pruning, and knowledge distillation from large teachers into small students. • Deploying models under hard size or latency budgets (on-device, edge, or embedded settings). — Generative Image Modeling — Experience with: • Training image generative models from scratch: diffusion models, GANs, VAEs, or flow-based models. • Iterating against sample-quality metrics such as FID. • Training-efficiency tricks that produce good generators quickly and at small parameter counts. — LLM Post-Training & Behavioral Robustness — Hands-on experience with one or more of: • Supervised fine-tuning and preference optimisation (DPO, RLHF, RLAIF) of open-weight language models, including building your own datasets via synthetic generation, noisy or weak supervision, and rejection sampling. • Shaping conversational behaviour over multiple turns: resistance to persuasion and sycophancy, calibrated confidence, and knowing when to accept corrections. • Alignment-style fine-tuning that changes a specific behaviour while preserving general capability. — Multilingual Pre-training — Experience with: • Training multilingual or low-resource-language models from scratch. • Tokenizer design across scripts and typologically diverse languages. • Balancing highly unequal per-language data (sampling temperatures, cross-lingual transfer) in data-constrained regimes. — Additional Areas of Interest — Experience in any of the following is a plus: • Scaling laws and training-efficiency research. • Curriculum learning and data ordering. • Model evaluation: benchmark construction, contamination control, statistically sound comparisons. • Uncertainty estimation and model calibration. • Data augmentation and synthetic data for robustness. — General Qualifications • 3+ years of machine learning research experience (PhD research counts toward this requirement). • Strong experience with PyTorch, JAX, TensorFlow, or similar ML frameworks. • Degree from a top-100 university, experience at a FAANG or comparable AI company, or an equivalent research track record through publications or impactful open-source contributions. — Why Join • Work on cutting-edge machine learning research. • Collaborate with leading AI researchers on challenging, high-impact projects. • Flexible, project-based work with competitive compensation.

$100 - $120 / hourOpen / Referral verified
Code / United States Remote

MLOps Engineer (JAX, PyTorch, Pallas/Triton)

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented MLOps Engineers with deep, hands-on expertise in modern ML frameworks — specifically JAX, PyTorch, and kernel-level programming (Pallas/Triton). This role involves AI model training and evaluation work, including writing and assessing MLOps tasks and solutions to generate high-quality training data for frontier AI systems. — This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. This is a 40-hour full-time engagement, with no conflicts/no other engagements. — 2. Key Responsibilities • Guide research and engineering teams to close knowledge gaps and improve AI model performance in MLOps, training infrastructure, and ML framework-level topics. • Design challenging, domain-relevant tasks, and write accurate and well-structured solutions to MLOps and ML systems problems. • Evaluate MLOps tasks and solutions and provide clear, written technical feedback. • Develop guidelines and detailed rubrics/evaluation frameworks to assess training pipeline design, distributed systems reasoning, and kernel-level optimization across tasks. • Collaborate with other subject matter experts to ensure consistency and accuracy in training data. — 3. Core Qualifications • 2+ years of dedicated professional experience in ML infrastructure, MLOps, or ML systems engineering at a recognized, top-tier organization. • Hands-on production experience with JAX and/or PyTorch at scale. • Experience writing or optimizing custom GPU kernels using Pallas (JAX) or Triton. • Demonstrable career progression. • Ability to engage reliably for at least 40 hours/week during weekdays. • Strong written communication skills and the ability to explain complex technical decisions clearly. — About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

$70 - $110 / hourOpen / Referral verified
Code / United States Remote

Machine Learning Engineer — Model Evaluation & Experimentation

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation. You will author complex, multi-step ML tasks — for example, taking a vague research idea like "modify how an RL reward is computed," implementing the change, running the training experiment, and analyzing the results to determine success — and verify exactly where frontier models fall short. — Each task represents one to two days of continuous, focused effort and spans multiple technical skills: implementation, experiment setup and execution, and rigorous analysis. You will work in a tight feedback loop with the lab's researchers. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design tasks: Turn real ML research ideas — like tweaking how an RL reward is computed — into well-defined, multi-step tasks. • Run experiments: Implement changes, run training experiments, and analyze the results to show what a correct solution looks like. • Explore RL ideas: Build some of your tasks around reinforcement-learning basics such as reward functions and training behavior. • Evaluate models: See how frontier models handle your tasks, and note where and why they fall short. • Work as a team: Compare notes with researchers and fellow experts so tasks stay consistent, rigorous, and fair. — 3. Core Qualifications • MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy domain. • 1+ years of experience in a research or research-engineering role. • Hands-on experience training and evaluating ML models and running experiments end-to-end — experiment setup, execution, and analysis — not just using ML libraries superficially. • Strong familiarity with large language models: their capabilities, limitations, and evaluation techniques. • Working proficiency in Python and Git, with comfort in both scripting and notebook environments. • Basic understanding of reinforcement learning (reward functions, policy training) is preferred. • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / Remote

Applied Computer Science Benchmark Specialist

Role Overview — We are seeking expert computer scientists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core computer science domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. — You will be assigned one of two task types: • Question Authoring — Create original, challenging multiple-choice questions in your area of computer science expertise, rate their difficulty, and submit them for review. • Question Verification — Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made. — Computer Science Domains Covered — Accelerator / GPU Kernel Engineering, Formal Methods & Automated Reasoning, Computer Architecture & Accelerators, Distributed Systems, DevOps & Site Reliability, Data Engineering & Databases, Cloud & Infrastructure, OS & Systems Kernel, Machine Learning Engineering, Web & API Development, Embedded Systems Engineering, Computer Graphics & Game Development, Mobile Engineering. — Key Responsibilities • Author original computer science questions that test deep conceptual understanding, not surface-level recall • Ensure questions are unambiguous, self-contained, and precisely defined — all necessary information must be in the problem statement • Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above) • Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers • Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format • Supply 1–5 academic references per question from reputable sources (peer-reviewed journals, university repositories) • For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made — Ideal Qualifications • PhD or doctoral candidate in Computer Science, Electrical Engineering, or a closely related field • Master's degree considered for candidates with exceptional depth in a specific subdomain • Strong command of graduate-level CS theory, algorithms, systems design, and/or machine learning • Research publications, industry experience at top tech companies, or competitive programming background is a strong plus • Excellent written English and ability to express complex ideas clearly and concisely — More About the Opportunity • Expected commitment: 10+ hours/week • Asynchronous, fully remote work

$66 - $84 / hourOpen / Referral verified
Code / Remote

ML Engineer (Coding Agent Experience)

About the Role — \- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. — \- Contributors help evaluate and improve frontier AI coding models through structured technical assessments. — \- The work focuses on realistic machine learning engineering workflows and model evaluation. — \- Spots are limited and filling quickly on a first come, first serve basis. — What You'll Do — \- Use frontier AI coding agents to complete and evaluate complex machine learning and AI engineering tasks. — \- Review model-generated implementations involving model training, inference systems, MLOps, and LLM applications. — \- Identify bugs, edge cases, performance issues, and failure modes. — \- Compare outputs from multiple frontier models and assess their strengths and weaknesses. — \- Apply professional engineering judgment to realistic ML engineering scenarios. — Time Commitment — \- Sprint based project that runs in 12-24 hour stretches based on client requirement. — Compensation — \- $400 per accepted task. — \- Typical tasks take approximately 2–3 hours after ramp-up. — \- Compensation is tied to accepted work. — Who Should Apply — \- 2+ years of professional machine learning engineering experience. — \- Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products. — \- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. — \- Ability to evaluate model-generated machine learning implementations and technical tradeoffs. — \- Experience deploying ML systems to production is preferred.

$85 / hourOpen / Referral verified
Code / Remote

ML Research PhD Experts (ICML / NeurIPS / ICLR Publications)

Overview — We're looking to rapidly assemble a small group of world-class machine learning researchers for an initial pilot. This is a high-priority engagement with an accelerated timeline, so both exceptional candidate quality and fast turnaround are critical. We're specifically seeking researchers with demonstrated contributions to frontier ML research, particularly those driving algorithmic innovation rather than applied analytics. — Candidate Requirements — Required Qualifications • PhD in Machine Learning, Computer Science, AI, or a closely related field • Published at least one main conference paper at ICML, NeurIPS, or ICLR • Strong preference for candidates with 2+ publications at these venues • Demonstrated experience conducting original ML research — Preferred Research Areas • Reinforcement Learning (RL) • Meta-Learning • Recursive Self-Improvement • AI for Science (e.g. weather forecasting, protein modeling, scientific discovery) — Timeline — Priority: Urgent — Expected Schedule • Initial data/results: Monday / Tuesday 27th July • * * — Success Criteria — Candidates should have a proven record of advancing state-of-the-art machine learning through publications at top-tier conferences and possess deep expertise in frontier ML research. Speed of sourcing is important, but quality should not be compromised.

$60 - $100 / hourOpen / Referral verified
Code / Remote

Network Engineer - Data for Autonomous Systems annotation

Are you a Level 3 / Tier 3 network support engineer interested in data science and autonomous infrastructure? Our client is building vertically integrated networking systems and using the data they generate to power the next generation of AI-driven infrastructure. They're looking for engineers experienced in final escalations, packet analysis, troubleshooting, and RCA workflows to help label, annotate, and structure networking data from real production systems. — This is a hands-on role that blends your L3 troubleshooting and incident-response experience with a growing understanding of how data pipelines are built and used in AI systems. — In this role, you'll: • Review real-world data from deployed networks: logs, configs, telemetry, event streams • Label and classify network behaviors, issues, anomalies, and incident patterns • Help define schemas and structure for large-scale data pipelines that downstream ML models will train on — You're a strong fit if you: • Work today as a Level 3 / Tier 3 / Principal Support Engineer keeping existing enterprise infrastructure online and stable — on-call rotation, RCAs, final escalations, troubleshooting outages — Must Have • Have hands-on experience with end-customer enterprise networks (switches, APs, firewalls in retail, healthcare, financial, manufacturing, university, hospitality, etc.) — Must Have • Bring hands-on Wi-Fi/wireless proficiency — enterprise WLAN controllers (Cisco WLC, Aruba, or Meraki), 802.1X/RADIUS, and wireless troubleshooting — Must Have • Do packet-level troubleshooting yourself — Wireshark, tcpdump, SPAN captures • Are curious about how raw infra data becomes machine learning input — This is a maintainer role — likely not the right fit if your current work is mainly network design/architecture, cloud/SRE, security/SOC, or IT helpdesk. — Your work will directly feed into the pipelines that power client's AI models, and help shape how intelligent systems reason about networks in the real world. — Here are more details about the role: • You will interface directly with the client team. • You are expected to work 30-40 hours/week, with your hours overlapping the Pacific (PT) business day. • This is an individual 1099 contract paid to a personal account — no corp-to-corp or agency billing. • You must be authorized to work in the US or Canada without sponsorship.

$50 - $70 / hourOpen / Referral verified
Math / Remote

Data Science and Analytics Experts

Role Overview • Mercor is seeking senior data science and analytics professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise data and analytics contexts. • The workflows are calibrated to the data scale, model complexity, and business-critical stakes of Fortune 500 and large public company data operations. • Contributors design enterprise data science scenarios, draft reference outputs, and write rubrics that capture how senior F500 data leaders think. — Key Responsibilities • Construct enterprise data science scenarios spanning large-scale predictive modeling, multi-stakeholder analytics governance, and complex data infrastructure decisions at F500 accounts. • Build analytics tasks across F500 machine learning model development, enterprise data pipelines, business intelligence at scale, experimentation/causal inference, and data strategy. • Develop data and MLOps scenarios involving tools such as Snowflake, Databricks, Python/R, SQL, Tableau/Power BI, and enterprise ML platforms (SageMaker, Vertex AI, MLflow) in F500 stacks. • Apply enterprise data science methodologies (statistical rigor, A/B testing frameworks, model validation, MLOps best practices) and produce reference analyses, model documentation, and executive-level insights. • Author rubrics that distinguish authentic enterprise data science judgment from generic textbook or tutorial-level recall. — Ideal Qualifications • 5+ years working as a data scientist, analytics leader, or ML engineer at a Fortune 500 technology or enterprise organization (Google, Meta, Amazon, Microsoft, Netflix) or inside an F500 data/analytics organization (JPMorgan, UPS, Unilever, PepsiCo, Walmart). • Direct ownership of F500 data products, F500 analytics initiatives, or F500 machine learning systems in production. • Fluency in enterprise data science tooling and methodologies, plus understanding of how F500 data governance, privacy compliance, and cross-functional stakeholder alignment actually work. • Prior rubric, technical curriculum, or model documentation authorship is a plus.

$60 - $70 / hourOpen / Referral verified
Code / Remote

Audio Engineer (Speech / TTS Audio Specialist) – German

Overview We're looking for audio engineers with strong experience editing and quality-checking voice recordings for speech and text-to-speech (TTS) applications, specifically for German-language audio. This role focuses on preparing high-quality audio datasets for machine learning systems, ensuring clarity, consistency, and technical compliance across large volumes of German speech data. Native or near-native German fluency is required to accurately assess speech quality, pronunciation, and naturalness. • * * — What You'll Do • Edit and process German voice recordings for use in TTS and speech AI systems • Perform audio cleanup including silence trimming, breath reduction/removal, de-clicking, de-essing, and noise reduction • Conduct detailed quality assurance checks for: • Clipping and distortion • Background noise and artifacts • Loudness consistency and level balancing • Ensure audio meets strict technical specifications for ML training pipelines • Evaluate recordings for German pronunciation accuracy, naturalness, and fluency • Work with large batches of recordings, maintaining consistency and throughput • Collaborate with teams working on German speech datasets, voice talent recordings, and model training workflows • * * — Basic Qualifications • Native or near-native German speaker with strong listening intuition for German speech nuances • Experience editing speech or voice recordings (not just music production) • Strong understanding of speech audio quality standards and common issues in recorded dialogue • Hands-on experience with tools like iZotope RX, Pro Tools, or similar audio restoration software • Ability to identify and fix artifacts, inconsistencies, and technical defects in audio • Experience working with high-volume audio datasets or structured workflows • * * — Preferred Qualifications • Experience working with TTS vendors or platforms • Prior involvement in speech ML, voice AI, or dataset creation/annotation workflows • Familiarity with audio QA pipelines, labeling, or evaluation processes • Experience directing or editing German voice talent recordings • * * — Ideal Candidate Profile • Detail-oriented with strong critical listening skills in German • Comfortable working with repetitive, high-precision tasks at scale • Able to balance speed and quality in production environments • Familiar with the nuances of human German speech vs synthetic voice requirements

$50 / hourOpen / Referral verified
Code / Remote

Audio Engineer (Speech / TTS Audio Specialist) – French

Overview We're looking for audio engineers with strong experience editing and quality-checking voice recordings for speech and text-to-speech (TTS) applications, specifically for French-language audio. This role focuses on preparing high-quality audio datasets for machine learning systems, ensuring clarity, consistency, and technical compliance across large volumes of French speech data. Native or near-native French fluency is required to accurately assess speech quality, pronunciation, and naturalness. • * * — What You'll Do • Edit and process French voice recordings for use in TTS and speech AI systems • Perform audio cleanup including silence trimming, breath reduction/removal, de-clicking, de-essing, and noise reduction • Conduct detailed quality assurance checks for: • Clipping and distortion • Background noise and artifacts • Loudness consistency and level balancing • Ensure audio meets strict technical specifications for ML training pipelines • Evaluate recordings for French pronunciation accuracy, naturalness, and fluency • Work with large batches of recordings, maintaining consistency and throughput • Collaborate with teams working on French speech datasets, voice talent recordings, and model training workflows • * * — Basic Qualifications • Native or near-native French speaker with strong listening intuition for French speech nuances • Experience editing speech or voice recordings (not just music production) • Strong understanding of speech audio quality standards and common issues in recorded dialogue • Hands-on experience with tools like iZotope RX, Pro Tools, or similar audio restoration software • Ability to identify and fix artifacts, inconsistencies, and technical defects in audio • Experience working with high-volume audio datasets or structured workflows • * * — Preferred Qualifications • Experience working with TTS vendors or platforms • Prior involvement in speech ML, voice AI, or dataset creation/annotation workflows • Familiarity with audio QA pipelines, labeling, or evaluation processes • Experience directing or editing French voice talent recordings • * * — Ideal Candidate Profile • Detail-oriented with strong critical listening skills in French • Comfortable working with repetitive, high-precision tasks at scale • Able to balance speed and quality in production environments • Familiar with the nuances of human French speech vs synthetic voice requirements

$50 / hourOpen / Referral verified