About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: LLM Red Team Specialist — Failure Modes & Edge Cases Type: Contract Compensation: $60–$90/hour Location: Remote Commitment: 35 hours/week Role Responsibilities Evaluate frontier AI models on coding, ML , and analysis tasks to identify vulnerabilities and failure modes. Design complex tasks that challenge models and are fair for grading. Document findings with clear evidence and reproducible steps. Collaborate with task authors to close loopholes and improve grading. Share insights with researchers to enhance benchmark quality. Work independently and asynchronously to meet deadlines and improve AI model performance . Qualifications Must-Have MSc or PhD in a STEM field or equivalent experience. 1+ years in research, research-engineering, security, or AI evaluation. Experience identifying vulnerabilities in LLMs or ML systems . Proficiency in Python and Git . Familiarity with LLM capabilities and evaluation techniques. Ability to work 35 hours/week . Preferred Experience in AI training , model evaluation, or benchmark/task authoring. Resources & Support For details about the interview process and platform information, please check: For any help or support, reach out to: PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity. Originally posted on Himalayas
mercorActive opening
AI Vulnerability Expert - Fully Remote | Upto $90/hr
Apply to this job
Free credits included. Sign up to start applying with Jobfinder.
llmmachine learningpythongitsecurityvulnerability assessmentevaluationbenchmarking
mercorAI Vulnerability Expert - Fully Remote | Upto $90/hr
Apply to this job