Simple, transparent pricing.
0% markup — you pay exactly what providers charge. Revenue comes from the platform fee, not a tax on your tokens.
Pay-As-You-Go
$18/mo + raw cost$9/mo + raw cost
Rp 159rb
- All 400+ LLMs
- 0% base markup
- Full guardrails
- PAYG billing
Starter
$24/mo$12/mo
Rp 209rb
- Budget-tier models
- Full guardrails
- 1 API key
- IDR billing
Pro
$78/mo$39/mo
Rp 699rb
≈ Rp 23rb/hari ☕
- Budget + Mid + Premium
- Batch routing (50% off)
- Semantic cache
- 5 API keys
Enterprise
$499+/mo
Rp 8,5jt+
- All tiers incl. Ultra
- SSO/SAML
- Audit log export
- SLA guarantee
- Net-30 invoicing
Weighted token formula
Not all tokens are equal
weighted_tokens = raw_tokens × provider_multiplier × tier_multiplier × quota_discount
Provider Multiplier
| Cloud Direct (OpenAI, Anthropic, Google) | 1× |
|---|---|
| Bedrock / self-hosted | 0.5× |
| Local ID (Jakarta edge) | 0.3× |
Model Tier Multiplier
| Budget (Gemini Flash, Llama, Haiku) | 1× |
|---|---|
| Mid (Claude Sonnet) | 4× |
| Premium (Claude Opus) | 17× |
| Ultra (largest models) | 42× |
Quota Discounts
| Cache hit (semantic) | 0% of quota used |
|---|---|
| Provider cache (prefix) | 10% discount |
| Batch routing | 50% discount |
| Real-time | 100% (full cost) |
Token consumption
How far do your tokens go?
Estimated at 1,000 tokens per request (typical chat turn)
| Model | Tier | Weight | Starter (500K wt) | Pro (20M wt) |
|---|---|---|---|---|
| Gemini 2.5 Flash | Budget | 1× | 500K calls | 20M calls |
| Llama 3.3 70B | Budget | 1× | 500K calls | 20M calls |
| Claude Haiku 4.5 | Budget | 1× | 500K calls | 20M calls |
| Claude Sonnet 4.6 | Mid | 4× | 125K calls | 5M calls |
| Claude Opus 4.7 | Premium | 17× | ~29K calls | ~1.2M calls |
Pay-As-You-Go markup
Volume-based progressive markup
PAYG customers pay a $9/mo platform fee + raw provider cost, with a small progressive markup based on monthly call volume.
| Monthly calls | Markup | Note |
|---|---|---|
| 0–1,000 | 0% | Pure pass-through |
| 1,001–5,000 | 1% | Near pass-through |
| 5,001–20,000 | 2% | Below OpenRouter (10–25%) |
| 20,001–50,000 | 3% | 3–5× cheaper than OpenRouter |
| 50,001+ | 5% cap | Still 2–5× cheaper than OpenRouter |
Models by region
Pick where your data is processed
Every model declares the region where the request is processed. Filter the catalog by region to align your model choice with your compliance posture. Cross-border data movement is opt-in.
Indonesia
Framework: UU PDP
Models: Frontier + curated chat
Best for: All Indonesian-resident workloads
Singapore
Framework: PDPA
Models: Frontier + curated chat
Best for: PDPA workloads (Q3 2026)
China
Framework: PIPL
Models: 9 OSS text + 3 OSS vision
Best for: Cost-optimized OSS, ~50–75% cheaper
United States
Framework: —
Models: Most frontier flagship
Best for: Latency-insensitive global
Europe
Framework: GDPR
Models: EU-resident frontier
Best for: GDPR workloads
Global
Framework: Multi-region
Models: Routed for best latency
Best for: Default residency-agnostic
Pricing across regions is pass-through — Anoman applies 0% markup. The platform fee covers the gateway, guardrails, and observability.
Compare plans
Everything you get
Full feature breakdown across all tiers.
| Feature | Starter | Pro | Pay-As-You-Go | Enterprise |
|---|---|---|---|---|
| Weighted tokens/mo | 500K | 20M | Unlimited | Custom |
| Overage | Hard cap | $1.20/1M wt | 0–5% markup | Negotiated |
| Token cap (budget, mo) | 25M | 120M | No cap | No cap |
| Token cap (mid, mo) | — | 24M | No cap | No cap |
| Token cap (premium, mo) | — | 4M | No cap | No cap |
| Ultra-class token cap | — | — | No cap | No cap |
| RPM burst guard | 30 | 120 | 60 | 30,000 |
| TPM burst guard | 60K | 500K | 200K | 100M |
| Inflight requests | 60 | 150 | 125 | 500 |
| Budget models | ||||
| Mid models | ||||
| Premium models | ||||
| Ultra models | ||||
| API keys | 1 | 5 | 5 | Unlimited |
| Prompt injection detection | ||||
| PII masking (4 modes) | ||||
| Content moderation | ||||
| Conversational guardrails | ||||
| Agent policy control | ||||
| MCP governance | ||||
| Batch routing (50% savings) | ||||
| Token packs | ||||
| Semantic cache | ||||
| Team management + RBAC | ||||
| Organization workspaces | ||||
| Member groups + entitlements | ||||
| SSO/SAML | ||||
| Audit log export | ||||
| Custom rate limits | ||||
| SLA guarantee | ||||
| Net-30 invoicing | ||||
| IDR billing | ||||
| Support | Community | Dedicated |
Ready to secure every AI call?
Start for free. No credit card required. Upgrade when you need more.
Curious how we stack up against other providers? Compare Anoman vs OpenRouter, OpenAI & more →
Frequently Asked Questions
Everything you need to know about Anoman AI and LLM gateway security.
An LLM gateway is a proxy between your application and the upstream model providers. It routes requests, enforces security policies, meters usage, and gives you observability — through one OpenAI-compatible endpoint. Anoman is the guarded LLM gateway: prompt-injection defense, PII masking, and policy control run on every call.