Batch jobs.
POST · /v1/chat/completions + GET · /anoman/v1/batch — async routing for non-interactive workloads. ~50% cheaper than realtime with 5-30 minute SLA depending on tier.
When to use batch
Pick batch when…
- Latency does not matter — overnight summarization, eval suites, data enrichment pipelines, batch RAG.
- Volume is high — thousands of independent requests where ~50% savings adds up.
- You can poll or wait — your code path is fine with 5-30 minute response times.
Pick realtime when…
- You are streaming output to a user UI.
- Latency is part of the product (chat, autocomplete, voice agents).
- You need
stream: true— batch does not support streaming.
SLA by tier
| Tier | Batch SLA | Notes |
|---|---|---|
| Starter | 30 min | Default; overflow + opt-in |
| Pro | 15 min | Default for new accounts |
| Pay-As-You-Go | 15 min | Same as Pro |
| Enterprise | 5 min | Or custom contract SLA |
When a queued job hits 80% of its SLA deadline without completing, Anoman auto-promotes it to realtime at full cost. You never see a breach — but you also stop saving.
Enqueue
Same endpoint, opt-in header
Batch is opt-in per request via x-anoman-prefer-batch: true (or set default_routing: batch on the API key in the dashboard for set-and-forget). The response is 202 with a poll_url.
Poll
GET /anoman/v1/batch/{job_id}
Hit the URL returned in the enqueue response. While processing, you get 202 with a hint poll_again_in_seconds. When complete, 200 with the full chat completion result plus savings_usd.
Python helper
Enqueue + poll in one call
Equivalent helper exists in the official Anoman SDK as client.poll_batch(job_id).
Cancel
DELETE /anoman/v1/batch/{job_id}
Queued jobs can be cancelled with a single DELETE. Once execution starts upstream, cancellation returns 409 (you are still billed for whatever has already run).
What batch does not support
Constraints
- No streaming —
stream:trueforces realtime regardless of headers. - No realtime guarantees — the Enterprise tier realtime SLA does not apply.
- Auto-promotion to realtime at 80% SLA — when SF degrades or queue is hot, we promote rather than miss the SLA. You stop saving but the job still completes on time.
- Same guardrail pipeline runs before enqueue — a batch request that fails injection detection returns 403 immediately and is never queued.
See batch economics in the dashboard.
Per-month savings broken down by model + provider.