Author

Author alpware

Research

Evaluating retrieval quality at scale

Good answers start with good retrieval. We share the evaluation harness we use to measure recall and precision across millions of documents. The post includes the metrics that…

alpware 1 min

Research

Smaller models, bigger wins

Bigger is not always better. We benchmarked a family of smaller models on routing tasks and found they matched larger ones at a fraction of the cost. We…

alpware 1 min

Research

Designing guardrails that hold

Guardrails fail when they are bolted on late. This post explains how we model policy as code so the same rules apply in testing and production. We also…