Skip to main content
Detect and filter unsafe content across multiple safety categories in both user inputs and assistant outputs. Use this detector to enforce community standards and regulatory policies.

What it detects

  • Hate and harassment
  • Self-harm and dangerous activities
  • Sexual and adult content
  • Criminal activity and weapons
  • Privacy, IP, elections, and safety-sensitive topics

Available models (versions)

  • moderation-v0 — general-purpose moderation across core categories

Detection Categories

The current model moderation-v0 provides comprehensive coverage across safety-sensitive categories, ensuring that your AI application remains compliant and secure.

Using the Moderation detector

Threshold levels

  • L1 (0.9): Confident
  • L2 (0.8): Very Likely
  • L3 (0.7): Likely
  • L4 (0.6): Less Likely

Notes

  • For stricter environments, use higher thresholds on sensitive categories.
  • Combine with Injection and PII detectors for comprehensive runtime safety.