AI Code Evaluator - Fully Remote | Upto $90/hr
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .
Position: AI Developer Trace Task Auditor
Type: Contract
Compensation: $70–$90/hour
Location: Remote
Role Responsibilities
- Evaluate the quality and correctness of AI-assisted software-development traces to enhance model training and evaluation.
- Assess end-to-end coding sessions produced with AI-assisted developer tools for correctness, workflow soundness, and reasoning.
- Provide clear, rubric-based written feedback to improve AI model outputs .
- Review complex coding trajectories for alignment with best practices and correctness.
- Collaborate with AI research teams to ensure consistency and relevance in training data.
- Work independently and asynchronously to meet deadlines while enhancing AI model performance .
Qualifications
Must-Have
- 3+ years professional software development.
- Hands-on experience with AI-assisted coding tools and agentic/spec-driven workflows ( Cursor , GitHub Copilot , Claude Code , or similar).
- Strong code-reading and debugging skills across full-stack or backend systems.
- Ability to evaluate multi-step coding trajectories for correctness and best practice.
Preferred
- Experience with Kiro or Amazon CodeCatalyst .
- Prior work evaluating or grading AI-generated code .
- Contributions to developer tooling .
Application Process (Takes 20–30 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to:
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
Originally posted on Himalayas