One API. Every provider.
OpenAI-compatible and Anthropic-compatible endpoints. Drop-in replacement — change base_url and api_key, nothing else.
API surface
Complete endpoint reference
OpenAI-compatible (/v1/)
- POST
/v1/chat/completionsMain completion. Returns 202 for batch-eligible requests. - POST
/v1/completionsLegacy text completions. - GET
/v1/modelsModel catalog with tier and batch support flag. - POST
/v1/embeddingsEmbeddings passthrough.
Anthropic-compatible (/anthropic/)
- POST
/anthropic/v1/messagesFull Anthropic Messages API proxy with streaming. Set ANTHROPIC_BASE_URL=https://api.anoman.io/anthropic.
Public — no auth required (/public/v1/)
- GET
/public/v1/pricingModel pricing — 0% markup, pass-through at provider cost. No auth required. - GET
/public/v1/statusSystem status + provider health. No auth required. - GET
/public/v1/modelsPublic model catalog. No auth required.
Anoman-native (/anoman/v1/)
- POST
/anoman/v1/keysCreate virtual API key. - GET
/anoman/v1/batch/{job_id}Poll batch job status or retrieve result. - GET
/anoman/v1/usage/savingsCache and batch cost savings breakdown. - GET
/anoman/v1/events/streamSSE stream — real-time gateway events. - POST
/anoman/v1/burst/activateBuy temporary rate limit multiplier (2x/5x/10x).
Full API reference with request/response schemas: Docs / API Reference →
Batch routing
50% savings on non-interactive workloads
Document analysis, background agent tasks, overnight report generation — eligible requests are automatically routed to provider native batch APIs (Anthropic, OpenAI, Google) with 50% cost savings.
from openai import OpenAI
client = OpenAI(base_url="https://api.anoman.io/v1", api_key="anm-sk-...")
# Returns 202 for batch-eligible requests
response = client.chat.completions.create(
model="claude-sonnet-4-6",
messages=[{"role": "user", "content": "Summarise this doc..."}],
extra_headers={"prefer_batch": "true"}
)
if response.status_code == 202:
job_id = response.json()["id"]
# Poll until complete
result = client.get(f"/anoman/v1/batch/{job_id}")Starter: 30 minutes SLA
Pro: 15 minutes SLA
Enterprise: Real-time only (SLA guarantee)
Auto-escalation: if job reaches 80% of SLA,
worker automatically promotes to real-time.
Cost savings: ~50% via provider native batch API
(Anthropic /v1/messages/batches, OpenAI /v1/batches)Response transparency
Every response carries full context
{
"choices": [{ "message": { "content": "..." } }],
"usage": { "prompt_tokens": 512, "completion_tokens": 128 },
"_anoman": {
"guardrails": {
"injection": { "status": "pass", "score": 0.12 },
"pii": { "status": "pass" },
"content": { "status": "pass" },
"policy": { "status": "pass" }
},
"routing": {
"mode": "realtime",
"region": "id",
"provider_type": "cloud_direct"
},
"cache": { "hit": false, "type": "none" },
"weighted_tokens": 1240,
"cost_usd": "0.0023",
"burst": { "active": false }
}
}X-Anoman-Guardrail-Injection: pass score=0.12
X-Anoman-Guardrail-Pii: pass
X-Anoman-Guardrail-Content: pass
X-Anoman-Cache: none
X-Anoman-Region: ID
X-Anoman-Tier: proIntegrations
Works with every AI agent and framework.
One-line setup for OpenAI-compatible and Anthropic-compatible agents. No SDK changes required.
OpenAI-compatible agents
Set OPENAI_BASE_URL — everything else stays the same
Works with: Cursor, Codex CLI, Aider, Cline, Continue, LangChain, CrewAI, AutoGen, OpenClaw, and any OpenAI SDK client.
export OPENAI_BASE_URL=https://api.anoman.io/v1
export OPENAI_API_KEY=anm-sk-...from langchain_openai import ChatOpenAI
chat = ChatOpenAI(
base_url="https://api.anoman.io/v1",
api_key="anm-sk-..."
)import os
os.environ["OPENAI_BASE_URL"] = "https://api.anoman.io/v1"
os.environ["OPENAI_API_KEY"] = "anm-sk-..."Anthropic-compatible agents
Set ANTHROPIC_BASE_URL for Claude Code and Agent TARS
Full Anthropic Messages API proxy including streaming. guardrails, routing, and observability apply to every call.
export ANTHROPIC_BASE_URL=https://api.anoman.io/anthropic
export ANTHROPIC_API_KEY=anm-sk-...
claude # All calls now go through Anoman guardrailsimport anthropic
client = anthropic.Anthropic(
base_url="https://api.anoman.io/anthropic",
api_key="anm-sk-..."
)Step-by-step guides
Per-tool setup for the most popular agents and editors.
Platform · Decision Models
The guarded decision layer for AI agents
Your agents make judgment calls constantly — route this, flag that, is this risky. Decision Models turn those calls into typed, calibrated, auditable decisions, powered by JEV and run through the same guarded Anoman gateway as every chat call.
Typed questions in, calibrated answers out
A Decision Models call takes a state — the thing being judged — and a set of typed questions, and returns a typed answer for each one with a calibrated probability attached. There's no free text generated at any point, so there's nothing for the model to hallucinate outside the schema you defined.
Yes / no
A binary judgment with a calibrated probability — is this urgent, is this spam, does this need review.
Classify
Pick one label from a fixed set you define — route a ticket, categorize content — with a confidence score per option.
Rate
A graded rating along a scale you define — tone, severity, quality — with the full probability distribution across grades.
Billed on input tokens only — output is free — and a typical call resolves in well under a second. See every decision model and its live pricing on the AI Models Directory.
Where typed decisions replace a chat call
Ticket & intent routing
Classify an inbound message into a team or workflow branch — before a human, or an expensive model, ever sees it.
Content & severity classification
Flag risky, sensitive, or policy-relevant content with a calibrated score, not a guess.
Scoring an LLM's own output
Rate a chat completion's quality, tone, or groundedness before it reaches a user — a cheap judge in front of an expensive generator.
Agent tool-call risk gating
Decide whether a proposed tool call is safe to run automatically, or needs a human in the loop.
One call, a typed answer
Send a state and one or more typed questions to POST /anoman/v1/decisions. It runs through the same guarded Anoman gateway as every chat call — authenticated, rate-limited, and metered — and returns a typed answer per question with a confidence score you can gate your routing logic on.
curl https://api.anoman.io/anoman/v1/decisions \
-H "Authorization: Bearer anm-sk-..." \
-H "Content-Type: application/json" \
-d '{
"model": "jev-latest",
"state": "Customer: My invoice charged me twice this month and I need this fixed today.",
"questions": {
"category": {
"type": "choice",
"instructions": "Which team should this ticket route to?",
"criteria": {
"billing": "Billing",
"technical": "Technical issue",
"account": "Account access"
}
}
}
}'{
"data": {
"model": "jev-1.13.0",
"answers": {
"category": {
"type": "choice",
"choice": "billing",
"confidence": 0.91,
"probabilities": { "billing": 0.91, "technical": 0.06, "account": 0.03 }
}
},
"usage": { "input_tokens": 118, "output_tokens": 0 }
},
"error": null,
"meta": { "request_id": "req_...", "timestamp": "2026-09-22T10:00:00Z", "region": "id" }
}Full request/response shapes, error codes, and language examples are in the Decision Models docs.
What actually happens to your data
Only state is redacted. PII detected in state is masked before it leaves Anoman. Your questions and instructions are sent to JEV exactly as authored and are not scanned or redacted — don't put secrets or sensitive identifiers in your question text.
US routing. Unlike our Jakarta-hosted gateway, Decision Models route your request to JEV's API in the United States. If your content must stay in-region, don't send it to this endpoint.
Injection scoring is monitor-only. A judge legitimately has to evaluate adversarial or attack-shaped text as part of its job — so prompt injection on state is scored and logged, but this endpoint never blocks on it.
Decisions are probabilistic. Every answer is a calibrated probability, not a certainty. Review important decisions before acting on them, especially where the returned confidence is low.