Red teaming is a deliberate attempt to make the system fail before a customer or an attacker does. People try to talk it into breaking its rules, leak data, or take an action it should refuse. The point is the list of failures you fix, not a certificate that it is safe.
Have someone who did not build it run the tests. Builders are too kind to their own instructions. Include the awkward versions: polite, rude, indirect, and hidden inside a long paste. Write down what got through, close those holes, and run the same tests again after the next model update, because a new model can reopen them.
Think of it this way: Red teaming is hiring someone to try to break into your building before an actual burglar does. You control the conditions, learn the vulnerabilities, and fix them before it matters.
Before launching a customer-facing AI assistant, a team runs three days of structured adversarial testing. They find the model will output a competitor's pricing if prompted a specific way. Fixed before launch.
Before launch, a tester asks a customer bot to ignore its rules and email them the last caller's account number. The instruction holds. The same request, buried inside a forwarded thread and phrased as a manager's order, gets a partial account number. That second case is the one they fix. The first case would have passed a happy-path demo.
Before any public or customer-facing AI launch, and after significant capability additions. Treat it as a required phase of any AI product build, not an optional extra. Red teaming finds the vulnerabilities you know to look for. It does not find every possible attack vector. Ongoing monitoring and incident response are the necessary complement.
RaftLabs treats this as part of the build: a source on the answer, a test set, and a record of what the system did. The launch is the start of that work, not the end. The related work on our side is AI governance.
This sits with the other reliability & risk terms on the glossary. Why a confident answer can still be wrong, and how you catch it. Worth reading next: Hallucination, Guardrails, and Evaluation (evals).