LLM Model Response Evaluation
- Evaluating UI widgets, infographics, image factuality, side-by-side comparisons, and similar AI evaluation activities.
- The work may involve text, images, audio, video, HTML widgets, PDFs, or combinations of these modalities.
- The work is domain-agnostic and may cover topics across arts, culture, history, science, engineering, and more.
- Resources will be expected to independently research unfamiliar topics using trusted sources before making evaluation decisions.
- Each task will include detailed project guidelines within the evaluation platform.
- 3+ years of hands-on experience in LLM / GenAI data evaluation.
- Bachelor's Degree required
- Ability to research unfamiliar topics using trusted sources and make well-supported judgments.
- Comfortable evaluating content across multiple modalities
Flexible and remote work
Variable workload: Accept or decline tasks based on your availability
No guaranteed hours: Workload may vary weekly
Our client, a global technology company that helps businesses build, train, and manage AI systems is looking for experts to evaluate model-generated content against defined quality rubrics such as factuality, consistency, aesthetics, and other evaluation criteria.
Originally posted on Himalayas
This role requires you to be in the United States. If that means relocating or flying in, it is worth checking fares before you commit to a start date.
Compare flights and hotels →Get remote data science jobs like this by email
10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.
Similar for you
Get 10 hand-picked remote jobs like this one in your inbox every morning. One email a day, matched to what you browse. No spam, one-click unsubscribe.
No thanks — continue to the application ↗