AItext-moderation
text-moderation: OpenAI's Latest Moderation Model
Model Overview
text-moderation (also referred to as text-moderation-latest) is OpenAI's most recent content moderation API model, automatically updated to the latest version over time. It classifies text content into harmful categories including violence, self-harm, sexual content, harassment, and hate speech, returning per-category probability scores and an overall flagged verdict. Unlike the -stable variant, this endpoint receives automatic improvements as OpenAI updates their moderation capabilities.
✨ Key Features
| Feature | Description |
|---|---|
| Auto-Updated | Always reflects the latest improvements to OpenAI's moderation capabilities |
| Multi-Category | Classifies across violence, self-harm, sexual, harassment, hate, and more |
| Per-Category Scores | Returns probability scores for each harm category individually |
| Free API | The moderation endpoint is free to use via the OpenAI API |
| Continuous Improvement | Benefits from ongoing OpenAI safety research without API changes |
🔗 Resources
| Resource | Link |
|---|---|
| API Docs | platform.openai.com/docs/guides/moderation |
📜 License & Access
Proprietary (Free) — Free to use via OpenAI API for moderation purposes.
You might also want to compare
Verified Sources
Model Specs
Parameters
Unknown
Context Window
Unknown
License
Proprietary
Deployment
Resources & Links
Lineage
Curator Notes
Bulk imported from OpenAI developer docs.
Compare Specs
Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.
Compare Model