ModerationClient

ModerationClient

OpenAI moderation client wrapper.

Constructor

new ModerationClient(moderationApiKey)

Initializes a new moderation client.
Source:
Parameters:
Name Type Description
moderationApiKey string OpenAI API key used for the moderation endpoint.

Methods

_createGuardrailLlm() → {object}

Create the OpenAI client adapter used by LLM-based guardrails. OpenAI Guardrails currently adds a temperature to GPT-5 requests. GPT-5.6 uses reasoning by default and does not accept that parameter, so the adapter removes it before forwarding the request.
Source:
Returns:
Type:
object
OpenAI-compatible client for guardrail execution.

_normalizeOpenAIError(error) → {Error}

Convert OpenAI and guardrail execution failures to a consistent error.
Source:
Parameters:
Name Type Description
error unknown Error raised by OpenAI or OpenAI Guardrails.
Returns:
Type:
Error
Normalized error.

(async) detectJailbreakAttempt(message, minimumJailbreakConfidenceScoreopt, conversationHistoryopt) → {Promise.<boolean>}

Use OpenAI Guardrails to detect a jailbreak attempt.
Source:
Parameters:
Name Type Attributes Default Description
message string Message text to inspect.
minimumJailbreakConfidenceScore number <optional>
0.7 Minimum confidence required to flag the message.
conversationHistory Array.<object> <optional>
[] Recent conversation messages.
Returns:
Type:
Promise.<boolean>
True when the jailbreak tripwire is triggered.

(async) moderate(message) → {Promise.<object>}

Use OpenAI's moderation API to check if the message violates content policy.
Source:
Parameters:
Name Type Description
message string Message text to moderate.
Returns:
Type:
Promise.<object>
Moderation result object.