Jwumasay joo-ma
← All roles
AI safety and technical robustness · Remote

AI Red-Team STEM Evaluator

A AI Red-Team STEM Evaluator helps train and evaluate AI by applying expert ai safety and technical robustness knowledge to structured data, model outputs and quality review workflows.

RemoteRemote · contract, project or permanentEnglish
Role overview

Stress-test AI systems using STEM knowledge to discover hallucinations, overconfidence, reasoning failures and unsafe technical outputs.

Key responsibilities

  • Design adversarial but policy-compliant tests.
  • Document reproducible failures.
  • Score severity and exploitability within project frameworks.
  • Verify fixes through regression testing.

What we look for

  • Strong STEM background plus AI evaluation or analytical testing experience.

How success is measured

Unique failures found, reproducibility, severity calibration.

How this works on Jwuma

Candidate profile → skills evidence → domain assessment → calibration task → qualification → project matching → production → peer review → expert QA → performance feedback and progression. Corpshore AI service alignment Applicable across annotation and labeling, RLHF and preference data, model evaluation and red-teaming, multimodal datasets and specialized AI data operations, depending on project scope.

Common questions

What does a AI Red-Team STEM Evaluator do?

A AI Red-Team STEM Evaluator helps train and evaluate AI by applying expert ai safety and technical robustness knowledge to structured data, model outputs and quality review workflows.

Related

Similar roles

Apply once, get matched to what fits

One application covers every role you qualify for. A person reads it.

Apply now