Researcher, Evaluations

🏢 Epoch AI · all Epoch AI jobs
📍 United States
💰 USD 115,000 - 200,000 / annual
📅 Posted 2026-07-15 · via Himalayas
🏷 AI-Research,Evaluation-and-Benchmarking,AI-model-evaluation,Research-Analyst,AI-Evaluation-Specialist,Evaluation-Researcher,Research-Evaluation-Specialist,Research-and-Evaluation
Apply on original site ↗

Epoch AI is looking for a researcher to evaluate frontier AI models on hard-to-grade tasks drawn from real-world scenarios.

About the role

We’re seeking a Researcher to lead a new effort evaluating how well frontier models perform on the kinds of open-ended tasks that make up real office work. You will curate a suite of realistic tasks to serve as a benchmark, design the grading rubrics for AI performance, and run newly-released models through the suite, assessing their performance both quantitatively and qualitatively.

The focus is on how models handle messy, real-world work rather than on scientific knowledge or programming ability. The role makes heavy use of AI tools, but strong software engineering experience is not required. Comfort setting up AI-assisted automated workflows is a plus.

If this role sounds interesting, we are also looking for researchers on multiple other teams.

Applications are rolling .

Key Responsibilities

-
Create and curate an evaluation suite. Find real-world tasks that serve as challenging tests for practical AI capabilities, and update the tasks over time as AI capabilities evolve. Devise rubrics for evaluating AI performance.

-
Evaluate AI systems. Regularly evaluate new, notable AI models and products on the task suite. Update tasks and rubrics to reflect the changing landscape of AI capabilities.

-
Communicate your research . Create public-facing reports, blog posts, and data visualizations with your observations. Ensure the evaluations feed into our other research topics and help keep our team informed.

-
Conduct data analysis . Analyze evaluation results and compare models across tasks.

-
Improve the process . You might automate parts of the workflow, and build out parts of the evaluation into standalone benchmarks.

What we are looking for

-
Analytical thinking. You conduct experiments with rigor and care, making sure that findings are well-supported by evidence.

-
Grounded, skeptical mentality . You form your own well-reasoned view of what an AI system can do, distinguishing practical capabilities from hype.

-
Comfort with AI agents and tools. You have experience working with AI agents in the course of your own work, and are comfortable delegating tasks.

-
Familiarity with AI benchmarks and evaluations . You follow AI capabilities at least casually and have opinions on what benchmarks do and don’t tell us.

-
Research and data-analysis experience, including enough comfort with light coding to analyze your own results.

-
Strong written communication skills: You can convey nuanced observations clearly and precisely.

Nice to have

- Experience testing frontier models and writing assessments of their capabilities

- Coding skills, including python proficiency
If you don’t tick all these boxes but think you would be a great fit, please consider applying anyway!

Compensation & Benefits

- Annual salary between $115,000 – $200,000 USD , depending on location and experience.

- Salaries are not restricted to USD, and contracts and payments are usually in local currencies. Conversions are based on one-year average exchange rates.

- Fully remote environment, including flexible work hours.

- Competitive global benefits program, including a comprehensive health insurance program—including supplemental benefits specific to a local country, as available and mandated by local law—and life insurance and a pension plan, if applicable in your country.

- Generous paid time off (PTO), including no specific annual limit, with 30 days PTO per year protected, unlimited personal and sick leave, and 4 months paid parental leave for permanent staff with at least 12 months of tenure (prorated parental leave if less than 12 months).

- A flexible and generous expense policy for you to spend on equipment and a large range of productivity tools or learning/development opportunities, including unlimited spending on AI tools, subject to regulations and manager appro

← All remote jobs