Community / AI red teaming
The field guide.
A starting point for curious minds. Learn the language, find a place to practise, and make your findings useful to someone else.
An independent resource from NestCipher. Gray Swan and the other projects below are separate organisations.
01 / A research workflow
Make every attempt
teach you something.
Start with a question. Keep the evidence. Change one variable at a time.
Understand the failure mode.
Pick a risk and learn what success or failure looks like. Our OWASP explorer connects the definitions with examples and mitigations.
Open the OWASP explorerChoose a scoped challenge.
Find a suitable Arena challenge, read its rules, and note the target behaviour. Establish a baseline before testing your hypothesis.
Explore Gray Swan ArenaDocument the observation.
Record the exact input, output and test conditions. Repeat the attempt and distinguish a reproducible failure from a one-off response.
Open the Research Workbench Use the notes checklist
Repeat the loop ↻ Keep the scope, change one variable, and compare with your baseline.
02 / Practise with known evidence
AUTHORED EXERCISES / VERSION 1Read the record.
Question the conclusion.
Work through invented cases with inspectable evidence and written explanations. Then prepare an untested, private experiment in the Workbench. No model or tools run in these labs.
03 / Keep a useful record
A finding is more
than a screenshot.
A simple record makes an experiment easier to reproduce, compare and explain.
- target
- Model, version, date and permitted scope
- hypothesis
- The specific behaviour you are testing
- input
- Exact prompt and relevant context
- observation
- Actual response, proposed actions and completed effects
- reproduction
- Attempts, conditions and repeatability
- impact
- Why the observed failure matters
04 / Outside the nest
Good places
to go deeper.
Gray Swan Arena
app.grayswan.aiExplore structured AI red-teaming challenges and the community around them. Read each challenge’s rules and scope before you begin.
OWASP GenAI Security
genai.owasp.orgThe primary reference for the LLM Top 10, including prompt injection, sensitive information disclosure and excessive agency.
PortSwigger Research
portswigger.netFollow technical security research and see how researchers explain a vulnerability, its impact and a reproducible technique.
CyberChef
gchq.github.ioDecode, transform and inspect data with recipes you can save and share. Useful when an input or output needs a closer look.
These resources open on their own websites. Challenge availability and access are managed by their respective providers.
05 / Build the field guide
PUBLIC TEACHING MATERIALTeach a lesson.
Show your reasoning.
A useful contribution starts with an invented example, a precise question and evidence someone else can inspect. Use the template to prepare a teaching proposal for maintainer review.
- 01
Invent the scenario. Use made-up tasks, records and recipients. Leave private prompts, transcripts and challenge techniques out.
- 02
Expose the evidence. Distinguish a reply, a proposed action and a completed effect. State what remains unknown.
- 03
Explain the answer. Include a counterexample, the limits of the conclusion and a primary reference.
A Markdown template to work on locally. This site does not accept or publish submissions. These lessons are reviewed as authored teaching examples, not measurements of a model's behaviour.
Something missing from your research workflow?
Suggest a tool or resource