
Real job — pulled straight from METR’s careers page · Verified August 10, 2026 · No reposts.
Job description
METR is hiring a Task Development Engineer — a contract, remote role ($150k–$300k). Apply directly on METR's careers page below.
Task Development Engineer
Team: Engineering & Research
Location: Remote
Commitment: Contractor
Workplace Type: remote
About METR
We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment.
METR has consistently set precedents for catastrophic AI risk evaluations, including the first independent safety evaluations (working informally with Anthropic and OpenAI in 2022), the first loss-of-control evaluations and first agentic dangerous capability evaluations, the first evaluations using finetuning (mentioned briefly here), the first independent evaluations using internal information about training, the first review partnership for company risk analysis, the first embedded redteaming, and the first evaluations of internal deployments.
We’ve been consulted and/or favorably referenced by groups on opposite ends of various spectra, including a16z, Khosla, Gary Marcus, Obama, and Dean Ball, and are known for producing one of the most positive results on AI capabilities (the time horizon trend) and the most negative (our downlift study). We’re generally referenced as the canonical third party assessor, e.g. as the obvious candidate to verify conditional pause agreements, and are trusted with AI incident investigations by frontier labs and governments.
We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.
About the role
- As part of informing the world about risk from frontier AI systems, METR often runs and publishes evaluations of frontier models.
-
Time Horizons is a central tool the world uses to understand AI progress. Our methodology has been included in system cards, called an "obsession" by the NYT, has wide reach online, and is used by governments to inform national policy. It is essential to our broader risk assessment work to have good capability evaluations.
-
Task Development Engineers contribute to METR’s expanding ambition of our evaluations with high quality tasks supporting the Time Horizons methodology. We expect our results to be seen by policymakers, frontier labs, national security stakeholders, and other key decisionmakers influencing society’s response to AI progress.
What this role looks like
-
(Primarily, and most importantly) Developing difficult, novel tasks for models. You will build well-scoped tasks that remain challenging as model time horizons grow, potentially to hundreds of hours.
-
Quality assurance for existing tasks. Once a task has been developed, you will verify that it's actually solvable as specified, and that the model is given (only) the information it needs.
-
Baselining and scoring tasks. Where helpful, you may be asked to baseline tasks within your domain of expertise, and/or score task completions from AIs or human baseliners.
-
Improving task development infrastructure. We're always improving our processes. Strong candidates will notice when existing workflows are inefficient or produce low-quality output, and take responsibility for improving them.
Skills we're looking for
-
Software engineering: You have several years of experience working on complex projects and codebases.
-
Evaluations: You have experience building hard (ideally agent-based) AI evaluations (e.g. RE-Bench, HCAST, SWE-bench Verified, Cybench, GPQA), ideally using the Inspect framework.
-
High attention to detail: You read closely, spot misspecifications and ambiguity, and pay attention to fiddly minutiae.
-
(Nice to have) Familiarity with METR infrastructure: Prior experience with Hawk, and familiarity with the methodology behind our Time Horizons work, is a plus.
Job details and compensation
- Location: Remote (worldwide)
- Hours: 20-40 hours per week (flexible schedule determined by you)
- Timezone Requirements: A minimum of 1 hour (and ideally 4 hours) of overlap with the Pacific Coast Time workday, but you determine your exact work schedule.
- Employment type: Contract / freelance
- You decide the manner in which you complete your tasks to a standard that matches other professionals in this field.
- Compensation: $150-300/hour.
- Top of this range is reserved for exceptional candidates.
- Individuals who contribute >80 hours will be acknowledged in the final research output (if desired).
Get Task Development Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Internship Digital Twin and Task Simulation for Humanoid Robot
Frequently asked questions
What is the salary for Task Development Engineer at METR?
The estimated salary range for Task Development Engineer at METR is $150,000 - $300,000 USD per year.
Is Task Development Engineer at METR a remote job?
Yes, Task Development Engineer at METR is a remote position. This role is open to remote candidates.
What is the seniority level for Task Development Engineer at METR?
Task Development Engineer at METR is a Mid Level / Senior level position.
How do I apply for Task Development Engineer at METR?
You can view the full description and apply for Task Development Engineer at METR on EchoJobs: https://echojobs.io/job/metr-task-development-engineer-gcm51.