cf.llm.prompt.unsafe_topic_categories
cf.llm.prompt.unsafe_topic_categoriesArray<String>Array of string values with the type of unsafe topics detected in the LLM prompt.
The possible values are the following:
| Value | Category name | Description |
|---|---|---|
VIOLENCE_AND_WEAPONS |
Violence and weapons | Content that promotes, glorifies, threatens, or provides instructions for physical violence or the acquisition, creation, or use of weapons. |
NON_VIOLENT_CRIME |
Non-violent crime | Content that encourages or facilitates nonviolent crimes, including fraud, theft, hacking, and the illegal drug trade. |
SEXUAL_CONTENT |
Sexual content | Sexually explicit or suggestive content, including sexual exploitation, trafficking, assault, harassment, and other non-consensual sexual acts involving adults. |
CHILD_SAFETY |
Child safety | Content that sexualizes, exploits, abuses, grooms, or otherwise endangers minors. |
HATE_AND_DISCRIMINATION |
Hate and discrimination | Content that attacks, demeans, discriminates against, or incites hatred toward people based on protected characteristics. This category also includes content that promotes dishonesty, manipulation, or professional misconduct. |
SELF_HARM_AND_SUICIDE |
Self-harm and suicide | Content that encourages, glorifies, or provides instructions for self-harm or suicide. |
Requires a Cloudflare Enterprise plan. You must also enable AI Security for Apps.
Example usage:
# Matches requests where an unsafe topic categorized as "NON_VIOLENT_CRIME" or "HATE_AND_DISCRIMINATION" was detected in the LLM prompt:
(cf.llm.prompt.unsafe_topic_detected and any(cf.llm.prompt.unsafe_topic_categories[*] in {"NON_VIOLENT_CRIME" "HATE_AND_DISCRIMINATION"}))Categories:
- Request