Hands-on guides for people who ship.
Every tutorial here comes out of work we actually did. They assume you are building something real and skip the parts that only work in a notebook.
Production-grade tool use with Claude
Strict schemas, idempotency keys, structured errors, and the observability that makes tool calls debuggable six weeks later.
- Tool use
- Claude
- Reliability
By Anand Subramanian
How we evaluate RAG systems in production
Build the labeled query set first, measure retrieval separately from generation, and gate deploys on the result.
- RAG
- Evaluation
- Retrieval
By Sana Qureshi
Building a prompt regression test suite
Turn prompt engineering from art into practice with CI-friendly evals you can run on every commit.
- Evaluation
- CI
- Prompting
By Sana Qureshi
Hybrid retrieval from scratch
Combining BM25 with dense embeddings and a cross-encoder reranker — and measuring whether it actually helped.
- RAG
- Retrieval
- Search
By David Okafor
Adding deterministic checkpoints to an agent
Confidence gates, policy checks, and blast-radius limits implemented in code rather than left to the model.
- Agents
- Safety
- Architecture
By David Okafor
Attributing LLM spend to features and teams
Instrument token usage end-to-end so the month-end invoice stops being a surprise.
- Observability
- Cost
- Platform
By Mariana Costa
We will build the eval harness with your team.
A fixed-scope engagement that leaves you with a working evaluation suite, running in your CI, against your own data.
