anoman
Sign InGet API Key
Documentation home

Guardrails

Four pre-call checks and two post-call checks run on every request. Blocked requests return 403 — never enqueued, never billed.

Guardrails always run before routing

The guardrail pipeline runs before any routing decision. A request that fails a pre-call check returns 403 immediately — it never reaches the LLM and is never billed.

Auth + rate limit check
→ [PRE-CALL GUARDRAILS]
    1. Prompt injection detection (ML classifier, ~30ms)
    2. PII detection (~20ms)
    3. Content moderation (keyword + ML, ~5ms)
    4. Tool-call policy check (allowlist/denylist, ~10ms)
    ↳ blocked? Return 403. Request never reaches LLM. Never billed.
      (Tool-call policy blocks only in Block mode; under Monitor, the default, the
       request passes and the hit is recorded in Security Events.)
Response cache check (opt-in) → Routing → LLM call
→ [POST-CALL GUARDRAILS]
    5. Response content filter
    6. Response PII scan
Token metering → Response to client

Four pre-call layers

Prompt Injection

ML-based prompt-injection classifier. Confidence threshold configurable per policy group. ~30ms CPU latency. Cannot be disabled system-wide; threshold is adjustable.

PII Masking

PII detection engine. Default entities: EMAIL, CREDIT_CARD, SG NRIC, ID NIK, MY MyKad, TH National ID. Opt-in: PHONE, IP_ADDRESS, URL. Four modes: redact, tokenize, synthetic, block.

Content Moderation

Keyword + ML classifier. English + Bahasa Indonesia. ~5ms latency. Configurable per policy group.

Conversational Guardrails

YAML + flow-rule config per policy group. Conversational flow enforcement — blocks topic steering, jailbreaks, system-prompt disclosure.

Four ways to handle detected PII

Set pii_mode per API key or policy group. The default is redact.

redact (default)

// Input: "My email is [email protected]"
// Output: "My email is [REDACTED]"
 
// API key setting:
{ "pii_mode": "redact" }

tokenize

// Input: "My email is [email protected]"
// Output: "My email is <PII:EMAIL:tok_a1b2c3>"
 
// De-anonymized in the LLM response automatically.
// Reversible — original value recoverable via PIITokenMap.
{ "pii_mode": "tokenize" }

synthetic

// Input: "My email is [email protected]"
// Output: "My email is [email protected]"
 
// Generated substitute — realistic-looking, non-reversible.
{ "pii_mode": "synthetic" }

block

// Any detected PII → 403 Forbidden immediately.
// Request never reaches the LLM.
{ "pii_mode": "block" }

See PII protection for the full mode reference.

Override guardrail defaults per API key

An API key's guardrail_overrides JSONB can selectively bypass guardrails independent of its policy group. Only PII and content moderation can be overridden per key — injection detection is mandatory.

// PATCH /anoman/v1/keys/{id}
{
  "guardrail_overrides": {
    "piiEnabled": true,
    "contentModerationEnabled": false
  }
}
 
// Priority stack (highest → lowest):
// 1. System forced (ops emergency)
// 2. Per-key guardrail_overrides
// 3. Policy group defaults
// 4. System defaults

Conversational flow enforcement

Configure via the dashboard Guardrails tab. YAML config + flow rules are stored per policy group and cached by config hash for performance.

# Policy group conversational-guardrail config
models:
  - type: main
    engine: openai
    model: gpt-4o
 
rails:
  input:
    flows:
      - self check input
  output:
    flows:
      - self check output
define flow self check input
  $allowed = execute self_check_input
  if not $allowed
    bot refuse to respond

define bot refuse to respond
  "I'm not able to respond to that."

Full guardrails overview: Platform / Guardrails →

Secure every AI call with guardrails

Guardrails run on every request — real-time and batch — before the model is ever called.

On this page