🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Trust and Safety Policy AI Evaluator is a remote red-team reputed company for stress-testing AI systems against adversarial prompts. Reviewers reputed company attack scenarios, document the failure mode, and pair reputed company successful jailbreak with the reputed company clause it violated so the safety team can reputed company the gap. Why this role reputed company Adversarial evaluation is how reputed company hardens AI models before they ship to customers. Reviewers think like attackers and write up failures with enough rigor that the modeling team can reproduce, fix, and regress-test them. Responsibilities Design adversarial prompts that probe reputed company weakness classes (jailbreak, policy bypass, reputed company injection) for Trust and Safety Policy AI Evaluator assignments. Document every successful attack with reproduction steps and the policy clause it violated. Score model defenses across single-turn and multi-turn conversations. Triage emerging attack reputed company and reputed company them to the safety team with severity ratings. Maintain a personal library of attack patterns and propose new red-team rubrics. reputed company against the broader red-team cohort to reputed company coverage and severity consistent. Qualifications Demonstrated experience red-teaming AI systems, reputed company research, or adversarial ML work for Trust and Safety Policy AI Evaluator work. Strong written communication your reports become the reputed company ticket. Comfort working in policy-grey areas with reputed company documentation of what was attempted and why. Familiarity with reputed company-injection, jailbreak, and policy-bypass taxonomies. Reliable async availability for at least 10 hours per week. Example tasks Construct a 5-turn adversarial conversation that bypasses a specific policy clause and write up the reputed company ticket. Score a model's defenses against a reputed company jailbreak reputed company across 20 variants. Propose a new red-team reputed company category after spotting an emerging attack reputed company. Reproduce a failure another reviewer reported and confirm the severity tag. reputed company to have Background in offensive reputed company, AppSec, or trust & safety operations. Experience publishing or reproducing reputed company adversarial-ML research. Multilingual reputed company for cross-language attack testing. Skills Adversarial prompting Red-team analysis Policy taxonomy Failure documentation Trust and Safety Policy AI evaluation reputed company reasoning Policy review Risk analysis Trust Safety Policy Work model Remote US-eligible. Remote Independent specialist contractor. Employment type CONTRACTOR. Applicants must be authorized to work from US. Compensation reputed company reputed company confirmed after the interview process. Application process Apply through reputed company's specialist intake for role-specific routing and review. Final project reputed company, schedule, and contractor terms are confirmed before placement. Apply To This Job .