Mistral Launches Text Moderation API With Category Scores and Configurable Thresholds
Mistral AI has released a text-focused Moderation API that classifies content across predefined safety categories including Sexual, Hate, Violence, PII, and Jailbreaking. The API returns category-level scores that developers can use to build custom guardrail workflows, with options to block, flag, or route content based on configurable thresholds. Two moderation endpoints are available — one for raw text and another for conversational content — to accommodate different application needs. The current supported model is mistral-moderation-2603, with older 2411 endpoints now deprecated. While the API provides flexible scoring infrastructure, organizations must still define their own escalation policies, audit processes, and monitoring practices to operationalize the outputs effectively.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in