Blogs
Long-form essays on distributed systems, backend engineering, databases, and AI systems.
2026
(8)- Rate Limit an LLM Agent by Tokens, Not Requests
Counting requests can't control LLM spend: one request may be a one-line question or an agent run making a dozen model calls. Limit by tokens instead — estimate, reserve atomically in a shared Redis counter, then reconcile with actual usage. Also covers token buckets vs. window quotas and per-run agent budgets.
- When a Redis Cache Hit Is the Wrong Answer
Cache-aside breaks in front of an LLM agent: a question has no fixed key, rebuilding an answer means paying the model again, and answers depend on document version and permissions. Three Redis caches — exact match, semantic, and tool results — and why a semantic hit for timer T3410 must never answer T3411.
- Rate limit LLM tokens with Redis when the cost arrives after the call
TL;DRA token limit is a reservation: reserve an estimate before the model call, correct it after, and give every agent run its own budget.
- 01Decide what each counter protects
- 02Estimate before the call
- 03Reserve in one atomic script
- 04Call the model, then reconcile
- 05Give every agent run its own budget
- 06Tell the caller which limit it hit
- How to cache an LLM agent with Redis without serving wrong answers
TL;DRPut three caches in Redis and check them in order of cost — exact match, semantic, then tool results inside the agent loop — and know the cases where a hit must be treated as a miss.
- 01Check in order of cost
- 02Add the exact-match cache
- 03Add the semantic cache, with an identifier check
- 04Add the tool-result cache inside the agent loop
- 05Re-check authorisation on every hit
- 06Invalidate by version
- 07Fail open when Redis is down
- 08Leave prefix caching to the model provider
- Multi-Tenant RAG Leaks Through the Search, Not the Login
In a shared RAG pipeline the login is the easy part; leaks happen in the search. The rule: a client may select a workspace but never assert one. Four checkpoints — entry, retrieval, prompt, and proof — keep the workspace inside every SQL and vector query, backed by a 30+ attack suite that passes only when nothing leaks.
- Multi-Tenant RAG Leaks Through the Search, Not the Login
TL;DROne pipeline, many users: the client may select a workspace but never assert one. Four checkpoints keep every answer inside the asker’s own documents.
- 01The server decides the workspace on every request
- 02Every path to the data carries the limit
- 03Document text never makes an access decision
- 04Prove it with refusal tests and an audit trail
- My RAG API Never Signs Tokens or Sees Passwords
If an attacker stole everything inside your API, who could they become? With an identity provider issuing tokens and the API only verifying them, the answer is no one. Four questions for RAG auth — including two attackers most designs forget: the model and the documents — and why every failure should close the door.
- Secure a RAG API without minting tokens or storing passwords
TL;DRLet an identity provider (Keycloak, Okta, Cognito) issue tokens; the RAG API only verifies them, locally, against cached public keys.
- 01The config holds nothing that signs
- 02Verify locally, then check membership
- 03The route and every tool take identity from the principal
2025
(8)- Placeholder Post 09
- Placeholder Post 10
- Placeholder Post 11
- Placeholder Post 12
- Placeholder Post 13
- Placeholder Post 14
- Placeholder Post 15
- Placeholder Post 16
2024
(6)- Placeholder Post 17
- Placeholder Post 18
- Placeholder Post 19
- Placeholder Post 20
- Placeholder Post 21
- Placeholder Post 22