Jwumasay joo-ma

Red teaming work, getting paid to find where a model breaks

· 2 min read · work types, evaluation

What is AI red teaming work and how do I get into it?

Red teaming means deliberately probing an AI model for unsafe, biased or manipulable behaviour, then documenting what you found so somebody else can reproduce it. The valuable output is not the fact that the model failed, it is a written account precise enough that an engineer can trigger the same failure deliberately.

The short version

  • Red teaming is adversarial testing of a model's safety behaviour.
  • A finding is only useful if another person can reproduce it from your write-up.
  • Findings are triaged by severity rather than counted.
  • Multilingual and culturally specific probing is significantly undersupplied.
  • The work requires documenting method, not just outcome.

Want to do this work?

Applications are read by a person and answered either way.

Apply to join
Common questions

People also ask

Read next

Related reading