AI Safety Expert - Red Team

๐Ÿข mercor ยท all mercor jobs
๐Ÿ“ South Africa
๐Ÿ’ฐ USD 29 - 45 / hourly
๐Ÿ“… Posted 2026-08-24 ยท via Himalayas
๐Ÿท AI-Safety,Red-Team-Operations,Cybersecurity,AI-Research,Adversarial-Machine-Learning,AI-Safety-Red-Teaming,AI-Red-Team-Specialist,AI-Safety-Expert,AI-Red-Team,AI-Red-Team-Security-Engineer,AI-Safety-Specialist,AI-Red-Team-Tester,AI-Red-Teaming,Freelance-AI-Red-Team-Specialist,AI-Safety-Practitioner
Apply on original site โ†—
About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts โ€” English & Portuguese (global) Type: Contract Compensation: $29โ€“$45/hour Location: Remote Role Responsibilities - Red team conversational AI models and agents. Focus on jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation. - Generate high-quality human data. Annotate failures, classify vulnerabilities, and flag systemic risks. - Apply structure using taxonomies, benchmarks, and playbooks to maintain consistent testing. - Document reproducibly. Produce reports, datasets, and attack cases that customers can act on. - Work independently and asynchronously . Participate in optional higher-sensitivity projects with clear guidelines and wellness resources. Qualifications Must-Have - Native fluency in English and Portuguese (global, excluding Brazilian Portuguese). - Prior red teaming experience in AI adversarial work , cybersecurity , or socio-technical probing . - Strong communication skills to explain risks to technical and non-technical stakeholders. Preferred - Experience in Adversarial ML : jailbreak datasets, prompt injection, RLHF/DPO attacks , model extraction. - Cybersecurity skills: penetration testing, exploit development, reverse engineering. - Socio-technical risk expertise: harassment/disinfo probing, abuse analysis, conversational AI testing. - Creative probing skills in psychology, acting, or writing for unconventional adversarial thinking. Application Process (Takes 20โ€“30 mins to complete) - Upload resume - AI interview based on your resume - Submit form Resources & Support - For details about the interview process and platform information, please check: - For any help or support, reach out to: PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity. Originally posted on Himalayas

โ† All remote jobs