Every call, fully traced — without exposing what's inside it.
An Overview with KPIs and a usage-over-time chart, a Requests log with per-request metadata and guardrail findings, and an Errors & Blocked view — metadata-only by default, and the underlying model provider is never shown to your team.
What you see
Three views, one request pipeline
Overview
KPI strip — request volume, spend, guardrail block rate, error rate — plus a usage-over-time chart so you can see traffic shape and cost trend at a glance.
Requests
A per-request log: model tier, token counts, cost, latency, routing mode, cache hit, and guardrail findings for that call. Prompt and completion content is never stored in this log — metadata-only by default.
Errors & Blocked
Coarse error categories and guardrail block reasons for the calls that didn't complete cleanly — surfaced without any prompt or response content attached.
Privacy posture
Metadata-only, provider-hidden by design
Observability is built to answer "what happened" without ever needing to show "what was said." Prompt and response content stays out of the log by default — full content capture is opt-in and masked when you turn it on. Your data is isolated per tenant, the underlying model provider is never named to your team, and everything is built residency-first for Indonesia under UU PDP.
Provider-hidden
Anoman routes the call — the specific model provider behind a response is never surfaced in the dashboard or the API. You see the model tier and outcome, not the vendor.
Per-tenant isolation
Every request, trace, and log line is scoped to your account. Nothing is pooled or visible across tenants — not even in aggregate.
Residency-first (UU PDP)
Metadata-only logging by default, opt-in masked content capture when you need deeper debugging, and Indonesia-residency-first infrastructure aligned with UU PDP.
Anomaly Detection
Catch unusual AI behaviour before it becomes a problem.
Two-layer detection: Z-score/Welford sliding windows for fast signals, Isolation Forest per-customer models for subtle drift.
Detection layers
Statistical + ML anomaly detection
Layer 1 — Z-score / Welford
Online algorithm, no history required. 4 rules run on every request:
- token_spike — tokens >> rolling average
- request_burst — RPM spike
- guardrail_storm — block rate spike
- cost_outlier — cost >> session average
Layer 2 — Isolation Forest
Per-customer models, trained nightly from a 14-day analytics lookback (200-sample minimum). 8-dimensional feature vector:
total_tokenscost_usdrequest_rate_60stool_call_countguardrail_rate_300ssession_lengthprompt_len_charshour_of_dayFail-soft design
Exceptions never block the completion pipeline. Anomaly scoring is fire-and-forget — a model load failure returns a degraded score, not a 500.
Inline scoring
ML score computed with a 5-minute process-local model cache. Models stored in the in-memory cache with 48h TTL, refreshed nightly by training job.
Ops alerts
Email alerts via Resend for anomalies, batch failures, and SLA breaches. Configurable threshold and recipient.
Dashboard integration
Acknowledge anomalies, link to trace for investigation. Severity cards in dashboard with open/high counts.
Platform · Response Faithfulness
Know when your AI is making things up
Faithfulness scores whether a model's answer actually follows from the source you gave it — so you catch hallucinations and ungrounded claims before they reach a customer. Calibrated on English and Bahasa Indonesia.
Retrieval-augmented and agent answers fail in a specific way: they sound confident but assert things the source never said. Faithfulness runs a multilingual natural-language-inference model that scores how well a response is entailed by its provided context, and returns a clear verdict. It measures groundedness against your source — it is not a general fact-checker — and it needs no extra LLM call, so it's cheap enough to run on every response.
Three verdicts you can act on
Every claim in the answer is supported by the source. Safe to serve.
The answer is partially supported or ambiguous — a good threshold to log, review, or ask for a citation.
The answer asserts claims the source doesn't support — a likely hallucination. Flag, block, or regenerate.
What a check actually returns
Source you provide
Anoman processes Indonesian customer data in Jakarta and bills in Rupiah. The free tier includes 500,000 weighted tokens per month.
Model answer
On the free tier you get 500,000 weighted tokens per month, and your data is processed in Jakarta.
Grounded — every claim traces to the source.
Model answer
The free tier includes 5 million tokens and your data is processed in Frankfurt.
Ungrounded — “5 million” and “Frankfurt” are not supported by the source (the unsupported spans are surfaced).
A real model, calibrated for the region
The signal is a multilingual NLI model that scores entailment between the source and the response — not a keyword match and not a mock. It's calibrated on an English + Bahasa Indonesia evaluation set (AUC 0.97 English, 0.999 Bahasa). It runs on the same in-region infrastructure as the gateway, so nothing leaves the region to score it, and it adds no provider cost.
Try it free, then run it on your traffic
Free tester — no signup
Paste a source and an answer and see the verdict live, in the browser. Nothing is stored.
Open the faithfulness tester →
On every response
As an opt-in signal on your own traffic — surfaced in your dashboard, scored in-region, with the raw score kept internal and a coarse band shown to you.
See the guardrail platform →