Back to openai-moderation
OpenAI /

AI
text-moderation-stable

Closed SourceSpecializedtextUpdated July 14, 2026

text-moderation-stable: OpenAI's Stable Moderation Model

Model Overview

text-moderation-stable is OpenAI's pinned, stable version of their content moderation API model. It classifies text content into harmful categories — including violence, self-harm, sexual content, harassment, and hate speech — and returns per-category scores along with an overall flagged verdict. The -stable variant ensures consistent behavior over time without automatic updates, making it suitable for production applications that require reproducible moderation results.


✨ Key Features

Specification Table
FeatureDescription
Stable VersionPinned to a fixed version to ensure consistent behavior over time
Multi-CategoryClassifies across violence, self-harm, sexual, harassment, hate, and more
Per-Category ScoresReturns probability scores for each harm category individually
Free APIThe moderation endpoint is free to use via the OpenAI API
Fast InferenceOptimized for high-throughput content moderation at scale

🔗 Resources

Specification Table

📜 License & Access

Proprietary (Free) — Free to use via OpenAI API for moderation purposes.

You might also want to compare

Verified Sources

Model Specs

closed-source

Parameters

Unknown

Context Window

Unknown

License

Proprietary

Deployment

api-only

Resources & Links

Lineage

Model Family

Part of the openai-moderation family

Curator Notes

Bulk imported from OpenAI developer docs.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model