SWE Task Evaluator - Fully Remote | Upto $90/hr
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .
Position: SWE-Bench Task Auditor
Type: Contract
Compensation: $70–$90/hour
Location: Remote
Role Responsibilities
- Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks .
- Assess repository-level tasks, reference patches, test harnesses, and grading integrity.
- Provide clear, rubric-based written feedback to improve AI model training .
- Audit reference patches, test runners, and Docker isolation to detect answer leakage and reward hacking.
- Work independently and asynchronously to meet deadlines and enhance AI model performance .
Qualifications
Must-Have
- 3+ years professional software engineering experience.
- Real open-source contribution or maintainer experience (merged PRs, committer/maintainer roles).
- Strong ability to audit reference patches, test runners, and Docker isolation.
- Fluency across common ecosystems ( Python and at least one of Java / Go / TypeScript / C++ ).
Preferred
- Familiarity with SWE-Bench (Verified) or similar repository benchmarks.
- Maintainer history on major Python OSS ( Django , Flask , scikit-learn , sympy , pytest , etc.).
- Prior code-review or task-grading experience.
Application Process (Takes 20–30 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to:
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
Originally posted on Himalayas