CBRN Harmful Content Attack

What is CBRN Harmful Content Attack?

A CBRN harmful content attack is an adversarial probe that tries to elicit discussion or actionable assistance related to chemical, biological, radiological, or nuclear weapons—development, production, acquisition, or distribution—from an AI system that should refuse such requests.

These probes stress high-severity safety policies: models must refuse clearly while remaining useful on legitimate science and emergency-response topics that are not weapons enablement.

What teams usually do about it

  • Include CBRN categories in red-team suites and refusal evals.
  • Use layered filters plus model-level safety training.
  • Review dual-use edge cases with domain experts.

Related Giskard articles

Score CBRN refusals with Giskard — run harmful-content probes mapped to high-severity categories and measure refusal reliability under paraphrases and multi-turn pressure. See the LLM security probe catalog.

Authority: NIST AI Risk Management Framework.

Get AI security insights in your inbox