Evaluating retrieval quality at scale
Good answers start with good retrieval. We share the evaluation harness we use to measure recall and precision across millions of documents. The post includes the metrics that…
Author
Good answers start with good retrieval. We share the evaluation harness we use to measure recall and precision across millions of documents. The post includes the metrics that…
Bigger is not always better. We benchmarked a family of smaller models on routing tasks and found they matched larger ones at a fraction of the cost. We…
Guardrails fail when they are bolted on late. This post explains how we model policy as code so the same rules apply in testing and production. We also…