anoman
Sign InGet API Key
Gateway

One API. Every provider.

OpenAI-compatible and Anthropic-compatible endpoints. Drop-in replacement — change base_url and api_key, nothing else.

API surface

Complete endpoint reference

OpenAI-compatible (/v1/)

  • POST/v1/chat/completionsMain completion. Returns 202 for batch-eligible requests.
  • POST/v1/completionsLegacy text completions.
  • GET/v1/modelsModel catalog with tier and batch support flag.
  • POST/v1/embeddingsEmbeddings passthrough.

Anthropic-compatible (/anthropic/)

  • POST/anthropic/v1/messagesFull Anthropic Messages API proxy with streaming. Set ANTHROPIC_BASE_URL=https://api.anoman.io/anthropic.

Public — no auth required (/public/v1/)

  • GET/public/v1/pricingModel pricing — 0% markup, pass-through at provider cost. No auth required.
  • GET/public/v1/statusSystem status + provider health. No auth required.
  • GET/public/v1/modelsPublic model catalog. No auth required.

Anoman-native (/anoman/v1/)

  • POST/anoman/v1/keysCreate virtual API key.
  • GET/anoman/v1/batch/{job_id}Poll batch job status or retrieve result.
  • GET/anoman/v1/usage/savingsCache and batch cost savings breakdown.
  • GET/anoman/v1/events/streamSSE stream — real-time gateway events.
  • POST/anoman/v1/burst/activateBuy temporary rate limit multiplier (2x/5x/10x).

Full API reference with request/response schemas: Docs / API Reference →

Batch routing

50% savings on non-interactive workloads

Document analysis, background agent tasks, overnight report generation — eligible requests are automatically routed to provider native batch APIs (Anthropic, OpenAI, Google) with 50% cost savings.

Python
from openai import OpenAI
client = OpenAI(base_url="https://api.anoman.io/v1", api_key="anm-sk-...")

# Returns 202 for batch-eligible requests
response = client.chat.completions.create(
    model="claude-sonnet-4-6",
    messages=[{"role": "user", "content": "Summarise this doc..."}],
    extra_headers={"prefer_batch": "true"}
)

if response.status_code == 202:
    job_id = response.json()["id"]
    # Poll until complete
    result = client.get(f"/anoman/v1/batch/{job_id}")
Batch SLAs
Starter:    30 minutes SLA
Pro:        15 minutes SLA
Enterprise: Real-time only (SLA guarantee)

Auto-escalation: if job reaches 80% of SLA,
worker automatically promotes to real-time.

Cost savings: ~50% via provider native batch API
(Anthropic /v1/messages/batches, OpenAI /v1/batches)

Response transparency

Every response carries full context

Response JSON
{
  "choices": [{ "message": { "content": "..." } }],
  "usage": { "prompt_tokens": 512, "completion_tokens": 128 },
  "_anoman": {
    "guardrails": {
      "injection": { "status": "pass", "score": 0.12 },
      "pii": { "status": "pass" },
      "content": { "status": "pass" },
      "policy": { "status": "pass" }
    },
    "routing": {
      "mode": "realtime",
      "region": "id",
      "provider_type": "cloud_direct"
    },
    "cache": { "hit": false, "type": "none" },
    "weighted_tokens": 1240,
    "cost_usd": "0.0023",
    "burst": { "active": false }
  }
}
Response Headers
X-Anoman-Guardrail-Injection: pass score=0.12
X-Anoman-Guardrail-Pii: pass
X-Anoman-Guardrail-Content: pass
X-Anoman-Cache: none
X-Anoman-Region: ID
X-Anoman-Tier: pro

Integrations

Works with every AI agent and framework.

One-line setup for OpenAI-compatible and Anthropic-compatible agents. No SDK changes required.

OpenAI-compatible agents

Set OPENAI_BASE_URL — everything else stays the same

Works with: Cursor, Codex CLI, Aider, Cline, Continue, LangChain, CrewAI, AutoGen, OpenClaw, and any OpenAI SDK client.

Cursor / Codex CLI
export OPENAI_BASE_URL=https://api.anoman.io/v1
export OPENAI_API_KEY=anm-sk-...
LangChain
from langchain_openai import ChatOpenAI

chat = ChatOpenAI(
    base_url="https://api.anoman.io/v1",
    api_key="anm-sk-..."
)
AutoGen / CrewAI
import os
os.environ["OPENAI_BASE_URL"] = "https://api.anoman.io/v1"
os.environ["OPENAI_API_KEY"] = "anm-sk-..."

Anthropic-compatible agents

Set ANTHROPIC_BASE_URL for Claude Code and Agent TARS

Full Anthropic Messages API proxy including streaming. guardrails, routing, and observability apply to every call.

Claude Code
export ANTHROPIC_BASE_URL=https://api.anoman.io/anthropic
export ANTHROPIC_API_KEY=anm-sk-...

claude  # All calls now go through Anoman guardrails
Python SDK
import anthropic

client = anthropic.Anthropic(
    base_url="https://api.anoman.io/anthropic",
    api_key="anm-sk-..."
)
Self-hosted (Ollama / vLLM / LM Studio): Connect via localhost — these bypass the Anoman gateway so guardrails, observability, and policy enforcement do not apply. For compliance workloads, use Tier 3a managed open-source models (routed through OpenRouter) instead.
Coding agent? Skip the manual env vars — our one-line installer points Claude Code, Cursor, Cline, Aider, and more at Anoman automatically, with a full revert path.

Step-by-step guides

Per-tool setup for the most popular agents and editors.

Platform · Decision Models

The guarded decision layer for AI agents

Your agents make judgment calls constantly — route this, flag that, is this risky. Decision Models turn those calls into typed, calibrated, auditable decisions, powered by JEV and run through the same guarded Anoman gateway as every chat call.

Typed questions in, calibrated answers out

A Decision Models call takes a state — the thing being judged — and a set of typed questions, and returns a typed answer for each one with a calibrated probability attached. There's no free text generated at any point, so there's nothing for the model to hallucinate outside the schema you defined.

noul

Yes / no

A binary judgment with a calibrated probability — is this urgent, is this spam, does this need review.

choice

Classify

Pick one label from a fixed set you define — route a ticket, categorize content — with a confidence score per option.

score

Rate

A graded rating along a scale you define — tone, severity, quality — with the full probability distribution across grades.

Billed on input tokens only — output is free — and a typical call resolves in well under a second. See every decision model and its live pricing on the AI Models Directory.

Where typed decisions replace a chat call

Ticket & intent routing

Classify an inbound message into a team or workflow branch — before a human, or an expensive model, ever sees it.

Content & severity classification

Flag risky, sensitive, or policy-relevant content with a calibrated score, not a guess.

Scoring an LLM's own output

Rate a chat completion's quality, tone, or groundedness before it reaches a user — a cheap judge in front of an expensive generator.

Agent tool-call risk gating

Decide whether a proposed tool call is safe to run automatically, or needs a human in the loop.

One call, a typed answer

Send a state and one or more typed questions to POST /anoman/v1/decisions. It runs through the same guarded Anoman gateway as every chat call — authenticated, rate-limited, and metered — and returns a typed answer per question with a confidence score you can gate your routing logic on.

Request
curl https://api.anoman.io/anoman/v1/decisions \
  -H "Authorization: Bearer anm-sk-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-latest",
    "state": "Customer: My invoice charged me twice this month and I need this fixed today.",
    "questions": {
      "category": {
        "type": "choice",
        "instructions": "Which team should this ticket route to?",
        "criteria": {
          "billing": "Billing",
          "technical": "Technical issue",
          "account": "Account access"
        }
      }
    }
  }'
Response
{
  "data": {
    "model": "jev-1.13.0",
    "answers": {
      "category": {
        "type": "choice",
        "choice": "billing",
        "confidence": 0.91,
        "probabilities": { "billing": 0.91, "technical": 0.06, "account": 0.03 }
      }
    },
    "usage": { "input_tokens": 118, "output_tokens": 0 }
  },
  "error": null,
  "meta": { "request_id": "req_...", "timestamp": "2026-09-22T10:00:00Z", "region": "id" }
}

Full request/response shapes, error codes, and language examples are in the Decision Models docs.

What actually happens to your data

Only state is redacted. PII detected in state is masked before it leaves Anoman. Your questions and instructions are sent to JEV exactly as authored and are not scanned or redacted — don't put secrets or sensitive identifiers in your question text.

US routing. Unlike our Jakarta-hosted gateway, Decision Models route your request to JEV's API in the United States. If your content must stay in-region, don't send it to this endpoint.

Injection scoring is monitor-only. A judge legitimately has to evaluate adversarial or attack-shaped text as part of its job — so prompt injection on state is scored and logged, but this endpoint never blocks on it.

Decisions are probabilistic. Every answer is a calibrated probability, not a certainty. Review important decisions before acting on them, especially where the returned confidence is low.

Route your first AI call through Anoman.