AI Safety Specialist - Fully Remote | Upto $45/hr

๐Ÿข mercor ยท all mercor jobs
๐Ÿ“ South Africa
๐Ÿ’ฐ USD 29 - 45 / hourly
๐Ÿ“… Posted 2026-08-23 ยท via Himalayas
๐Ÿท AI-Safety-Specialist,Red-Teaming,Cybersecurity,AI-Research,Security-Engineering,AI-Safety-Analyst,AI-Safety-Expert,AI-Safety-Practitioner,AI-Safety-Researcher,AI-Safety-Evaluator,Remote-AI-Specialist
Apply on original site โ†—

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: AI Safety Experts โ€” English & Portuguese (global)
Type: Contract
Compensation: $29โ€“$45/hour
Location: Remote
Role Responsibilities

- Red team conversational AI models and agents. Focus on jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.

- Generate high-quality human data. Annotate failures, classify vulnerabilities, and flag systemic risks.

- Apply structure using taxonomies, benchmarks, and playbooks to maintain consistent testing.

- Document reproducibly. Produce reports, datasets, and attack cases that customers can act on.

- Work independently and asynchronously . Participate in optional higher-sensitivity projects with clear guidelines and wellness resources.

Qualifications

Must-Have

- Native fluency in English and Portuguese (global, excluding Brazilian Portuguese).

- Prior red teaming experience in AI adversarial work , cybersecurity , or socio-technical probing .

- Strong communication skills to explain risks to technical and non-technical stakeholders.

Preferred

- Experience in Adversarial ML : jailbreak datasets, prompt injection, RLHF/DPO attacks , model extraction.

- Cybersecurity skills: penetration testing, exploit development, reverse engineering.

- Socio-technical risk expertise: harassment/disinfo probing, abuse analysis, conversational AI testing.

- Creative probing skills in psychology, acting, or writing for unconventional adversarial thinking.

Application Process (Takes 20โ€“30 mins to complete)

- Upload resume

- AI interview based on your resume

- Submit form

Resources & Support

- For details about the interview process and platform information, please check:

- For any help or support, reach out to:

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Originally posted on Himalayas

โ† All remote jobs