anoman
MasukDapatkan API Key
Recipes
Beranda Dokumentasi

Chat retrieval-augmented dengan penghematan cache 90%.

Embed dokumen Anda sekali. Tempatkan system prompt yang stabil di posisi yang tepat agar provider cache aktif pada setiap query. Bayar 10% pada prefix yang ter-cache, harga penuh hanya pada pesan user yang kecil.

Kenapa pola ini — Cache prefix-nya, bukan pertanyaannya

Rantai RAG naif menaruh semuanya di pesan user — instruksi, chunk hasil retrieval, dan pertanyaan. System prompt-nya kosong. Ini berfungsi tetapi seluruh prompt tidak ter-cache pada setiap panggilan.

Rantai yang lebih cerdas menaruh instruksi yang stabil di system prompt (jarang berubah) dan retrieval + pertanyaan yang bervariasi di pesan user. Anoman otomatis menyisipkan penanda cache saat system prompt cukup panjang; panggilan berikutnya mengenai provider prompt cache dan membayar 10% pada prefix yang ter-cache.

Lihat /docs/concepts/caching untuk mekanika cache secara detail.

1. Indeks — Embed sekali, gaya batch

Embeddings itu murah (~$0.02 per juta token untuk text-embedding-3-small). Batch 100 chunk per panggilan untuk mengamortisasi overhead HTTP.

Indeks dokumen dengan embeddings

# index.py — embed every doc once, store in any vector DB
import os
import json
from pathlib import Path
from openai import OpenAI
 
client = OpenAI(
    base_url="https://api.anoman.io/v1",
    api_key=os.environ["ANOMAN_API_KEY"],
)
 
EMBED_MODEL = "text-embedding-3-small"   # 1536 dims, $0.02/Mtok
DOCS = Path("./knowledge_base")
 
# Cheap chunk function — use a real chunker (e.g. langchain's
# RecursiveCharacterTextSplitter) for production.
def chunk(text: str, size: int = 1200, overlap: int = 200):
    chunks = []
    i = 0
    while i < len(text):
        chunks.append(text[i:i+size])
        i += size - overlap
    return chunks
 
# Batch embed for throughput
records = []
texts: list[str] = []
metadata: list[dict] = []
for doc_path in DOCS.glob("*.md"):
    text = doc_path.read_text()
    for j, ch in enumerate(chunk(text)):
        texts.append(ch)
        metadata.append({"source": doc_path.name, "chunk": j})
        if len(texts) >= 100:    # batch of 100 per embedding call
            response = client.embeddings.create(model=EMBED_MODEL, input=texts)
            for emb, meta in zip(response.data, metadata):
                records.append({**meta, "embedding": emb.embedding})
            texts, metadata = [], []
 
# Flush remainder
if texts:
    response = client.embeddings.create(model=EMBED_MODEL, input=texts)
    for emb, meta in zip(response.data, metadata):
        records.append({**meta, "embedding": emb.embedding})
 
# Persist — replace with your vector DB (Qdrant, pgvector, etc.)
Path("./index.jsonl").write_text(
    "\n".join(json.dumps(r) for r in records)
)
print(f"Indexed {len(records)} chunks")

2. Query dengan system prompt yang bisa di-cache — Polanya

Query — embed + retrieve + chat

# query.py — retrieval-augmented chat call
import json
from pathlib import Path
import numpy as np
from openai import OpenAI
 
client = OpenAI(
    base_url="https://api.anoman.io/v1",
    api_key=os.environ["ANOMAN_API_KEY"],
)
 
# Load the index. Replace with your vector DB query.
INDEX = [json.loads(line) for line in Path("./index.jsonl").read_text().splitlines()]
EMBEDDINGS = np.array([r["embedding"] for r in INDEX])
 
def top_k(query: str, k: int = 5) -> list[dict]:
    q = client.embeddings.create(
        model="text-embedding-3-small",
        input=query,
    ).data[0].embedding
    sims = EMBEDDINGS @ np.array(q)        # cosine if normalized
    idx = sims.argsort()[-k:][::-1]
    return [INDEX[i] for i in idx]

Bangun panggilan chat dengan system prompt yang bisa di-cache

# This is the load-bearing trick: put STABLE content in the system
# prompt, VOLATILE content in the user message. The provider caches
# the system prefix and only re-processes the user message.
 
STABLE_SYSTEM = """
You are a knowledgeable assistant for the Anoman platform.
You answer based on the provided context chunks. If the answer
isn't in the context, say so.
 
Style: concise, technical, no marketing language.
""".strip()
 
def ask(question: str) -> str:
    chunks = top_k(question, k=5)
    # IMPORTANT: keep the CONTEXT in the user message, not system.
    # If you put retrieved chunks in the system prompt, the cache
    # invalidates on every query.
    context = "\n\n".join(
        f"[{c['source']}#{c['chunk']}]\n{INDEX[c['chunk']].get('text', '')}"
        for c in chunks
    )
    response = client.chat.completions.create(
        model="claude-sonnet-4-6",
        messages=[
            {"role": "system", "content": STABLE_SYSTEM},
            {"role": "user", "content": f"""
Context:
{context}
 
Question: {question}
"""},
        ],
        max_tokens=500,
        temperature=0.0,
    )
    return response.choices[0].message.content
 
print(ask("How do batch jobs handle SLA breaches?"))

Anti-pola cache: jika Anda menaruh chunk hasil retrieval ke system prompt, setiap query punya system prompt berbeda dan cache tidak pernah aktif. Jaga system tetap stabil; taruh konteks di pesan user.

3. Verifikasi cache aktif — Baca cached_tokens

Verifikasi cache aktif

# After the FIRST call, the system prompt is cached.
# Subsequent calls should show cached_tokens > 0 in usage.
 
response = client.chat.completions.create(
    model="claude-sonnet-4-6",
    messages=[
        {"role": "system", "content": STABLE_SYSTEM},
        {"role": "user", "content": "Different question..."},
    ],
)
 
cached = response.usage.prompt_tokens_details.cached_tokens
print(f"Cached prefix tokens: {cached}")     # > 0 = cache fired
 
# Cost comparison via _anoman extension:
print(f"Cost this call: ${response._anoman['cost_usd']}")

Untuk model Anthropic, system prompt harus ≥ 2.048 token agar bisa di-cache. Untuk OpenAI otomatis untuk panjang berapa pun. Untuk Google 32.768+ token (via CachedContent API eksplisit). Lihat tabel per-provider di /docs/concepts/caching.

4. Stream untuk UX — Tumpukan streaming + cache

Set stream: true. Provider cache tetap berlaku — Anda menghemat 90% pada prefix dan men-stream bagian variabel yang kecil token demi token.

Stream retrieval + respons streaming

# For UX, run retrieval + streaming in parallel. The first chunk
# from the model arrives in ~1s; retrieval takes ~200ms.
 
stream = client.chat.completions.create(
    model="claude-sonnet-4-6",
    messages=[
        {"role": "system", "content": STABLE_SYSTEM},
        {"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
    ],
    stream=True,
)
 
for chunk in stream:
    delta = chunk.choices[0].delta.content or ""
    yield delta   # if running in a FastAPI streaming response

Tips produksi

  • Masa hidup cache ~5 menit. Trafik idle tidak diuntungkan. Untuk aplikasi bertrafik jarang, kirim permintaan pemanasan sintetis setiap 4 menit agar cache tetap hangat.
  • Versioning system prompt — perubahan apa pun membatalkan cache. Gunakan komentar versi berbasis content-hash di bagian atas system prompt agar Anda bisa mendeteksi pergeseran tak sengaja di kode Anda.
  • Jangan cache PII pengguna di system prompt — guardrail PII Anoman berlaku, tetapi Anda sebaiknya tidak menyimpan data pelanggan sensitif di lapisan cache multi-tenant.
  • Semantic cache di atasnya — untuk aplikasi gaya FAQ di mana pertanyaan yang sama berulang, tambahkan semantic cache Anoman (x-anoman-cache: semantic) di atasnya. Permintaan identik lalu kembali tanpa panggilan upstream sama sekali.
  • Reranking setelah retrieval — top-k berdasarkan cosine itu kasar. Untuk kualitas lebih tinggi, lewatkan kandidat melalui model rerank murah (qwen-2.5-7b) sebelum dimasukkan ke pesan user.

Periksa penghematan cache secara langsung.

Halaman Usage menampilkan cache_savings_usd sebagai kolom terpisah.