Role directory

Software engineering jobs

4 active, referral-verified opportunities.

Code / Remote

Software Engineer — Agentic Search Systems

We're looking for engineers who have shipped production search systems (especially agentic ones now that we're in the era of LLMs and agents) and are thinking hard about the requirements and optimizations of those systems. — You've probably: • Owned relevance or retrieval on a system real users depended on • Decided how to quantify impact and improvements with these systems. • Built out systems related to agentic search. • Scaled data infrastructure to power modern search systems. — This is a 25-minute conversational interview. No coding, no take-home. We want to hear how you actually think about search quality, evaluation, and the tradeoffs you've made in real systems. Bring war stories — the messier and more specific, the better. — If your interview stands out, we'll follow up with a paid 30-minute live conversation with our team — $200 for your time, paid on completion of the call. — If that sounds like you, apply and complete the interview. We review every submission.

$80 - $150 / per-taskOpen / Referral verified
Code / United States Remote

QA/Test Engineer

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models, and complex multi-step tasks are only useful if they are airtight: unambiguous, correctly graded, and robust to shortcuts. We are seeking experienced QA and test engineers to be the quality backbone of this benchmark — designing the test cases and review processes that guarantee every task measures what it claims to measure. — Each task under review represents one to two days of expert effort and spans multiple technical skills, so quality review here means genuinely understanding the task: running it, probing its edge cases, and debugging its environment. You will work in a tight feedback loop with the lab's researchers and task authors. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design checks: Create test cases that confirm each task works as intended — including the tricky edge cases. • Review tasks: Give tasks and reference solutions a careful read before they're finalized, catching ambiguity and gaps early. • Debug: Roll up your sleeves in Python when a task or its checks don't behave the way they should. • Shape the process: Help build simple, repeatable quality checklists, and share feedback authors can act on right away. • Protect the results: Watch for shortcuts and grading gaps in AI agent runs so benchmark scores stay trustworthy. — 3. Core Qualifications • MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain. • 1+ years of experience in test engineering, quality assurance, or a research/software engineering role with strong quality ownership. • Demonstrated skill designing test cases and quality-review processes, and debugging complex systems end-to-end. • Working proficiency in Python and Git, and comfort navigating unfamiliar codebases and environments. • Exceptional attention to detail and clear written documentation habits. • Past experience in AI training, model evaluation, or quality review of AI-generated work is preferred. • A perfectionist mindset: creativity in finding what others missed, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / United States Remote

Software Engineering Expert

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models. — 1. Overview — A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced software engineers to act as task curators. You will design, implement, and review complex, multi-step engineering tasks that simulate the real-world challenges research engineers face — realistic, genuinely hard problems that today's best AI coding agents cannot yet solve reliably. — Each task represents one to two days of continuous, focused effort and spans multiple technical skills: Python implementation, environment and tooling setup, debugging, and clear documentation. You will work in a tight feedback loop with the lab's researchers, verifying exactly where and why frontier models fail on your tasks. — This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. — 2. Key Responsibilities • Design tasks: Create realistic, multi-step software engineering challenges — the kind of work you'd actually do on the job — that push the limits of today's best AI coding agents. • Build reference solutions: Solve your own tasks in Python, with the setup and checks needed so each task has a clear, verifiable answer. • Work with AI tools: Use AI coding assistants as part of your everyday workflow, and observe where they help and where they fall short. • Review and refine: Look over tasks built by fellow experts and share feedback on clarity, correctness, and difficulty. • Learn from failures: See how AI agents attempted your tasks and help the research team understand what tripped them up. — 3. Core Qualifications • MSc or PhD in computer science or another STEM field, or equivalent practical experience in a research-heavy domain requiring significant coding and data analysis. • 1+ years of experience in a research, research-engineering, or software engineering role. • Strong hands-on Python scripting and debugging skills, with clean-code habits and attention to readability. • Everyday fluency with version control (Git), IDEs, and standard software development workflows. • Experience with AI coding assistants, prompt engineering, or agent workflows is preferred. • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred. • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems. • Ability to engage reliably for approximately 35 hours per week. — About Cincinnatus LLC — Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives. — Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows. — Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC. — Equal Employment Opportunity — Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic. — Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

$60 - $90 / hourOpen / Referral verified
Code / India Remote

Software Engineer, Full Stack — India

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models. — 1. Overview — Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking talented Full-Stack Software Engineers with hands-on production experience across modern application stacks — Python, Java, Rust, C#, C++, and TypeScript — to build, integrate, and stress-test real software on top of frontier models before they ship. Engineers work directly with the AI lab's engineering manager on fast-moving projects and are expected to onboard quickly and contribute working code early. — This is a full-time engagement of 40 hours per week. — 2. Key Responsibilities • Build and ship full-stack applications, services, and internal tools — in Python, Java, Rust, C#, C++, or TypeScript — that exercise frontier model capabilities end to end. • Integrate pre-release model APIs into working software, including the surrounding scaffolding: tool interfaces, evaluation harnesses, and telemetry. • Diagnose and document model and integration failure modes surfaced while building, translating engineering observations into clear written feedback for the research team. • Prototype quickly against shifting requirements, delivering working increments with limited up-front specification. • Collaborate with the engineering manager and other engineers to maintain consistency in code quality, architecture, and technical documentation. — 3. Core Qualifications • 3+ years of dedicated professional experience building and shipping production software at a recognized, top-tier organization. • Production experience in at least one of Python, Java, Rust, C#, or C++, plus working proficiency in a second language — this cohort is intentionally staffed across multiple language ecosystems. • Hands-on experience across the stack: backend services and APIs, a modern front-end framework (React or equivalent), relational or document databases, and cloud deployment. • Demonstrated ability to onboard onto an unfamiliar codebase and deliver working code quickly with minimal ramp-up. • Demonstrable career progression. • Ability to engage reliably for at least 40 hours/week during weekdays. • Strong written communication skills and the ability to explain complex technical decisions clearly.

$25 - $30 / hourOpen / Referral verified