01
Retrieval & RAG
Embedding, chunking strategy, evals. We build retrieval pipelines that survive contact with the actual corpus — not synthetic benchmarks.
- OpenAI
- Anthropic
- Voyage
- pgvector
- Qdrant
Now booking Q3 2026 4 seats left
A senior studio of ex-OpenAI, ex-Anthropic engineers and applied researchers. We take your AI roadmap from “interesting Slack thread” to a measurable production line in 6–12 weeks.
In production at
01 — Capabilities
We don’t outsource. Every engagement is staffed by the same team — researcher, applied engineer, designer, ops — across the full delivery.
01
Embedding, chunking strategy, evals. We build retrieval pipelines that survive contact with the actual corpus — not synthetic benchmarks.
02
LoRA / QLoRA / DPO on your data. We start with the eval, never the loss curve.
03
Stateful, observable, recoverable. We ship agents that survive their second week.
04
Custom eval harnesses, cost dashboards, drift alerts. The boring part most teams skip.
05
Surfaces, latency budgets, fallback UX. What does “loading” look like at 1.2 s of TTFT?
06
Buy-vs-build, model selection, vendor lock-in. Hourly retainer for CTOs.
02 — Process
01
Week 0
Two 90-minute calls. We map your data, your constraints, and your definition of success. Output: a one-page scope you can show your board.
02
Week 1–3
A working slice — not slides. End-to-end through your stack. We bring our own eval harness; you keep it.
03
Week 4–8
Hardening, observability, cost controls, on-call runbooks. By the end of week 8 your team is on duty and we’re paired with them.
04
Week 9–12
Two weeks of pair-shipping. We document, we offboard, we leave the studio number with you for emergencies.
03 — Selected work
Strato AI · Customer Operations · 2026
A retrieval-grounded support agent trained on 12,840 resolved tickets. Hand-off to a human happens only on the 4% of conversations the model marks low-confidence — a number we instrument and tune weekly.
47s
p50 first response
↓ from 6h 12m
96%
CSAT after launch
↑ from 81%
$2.1M
Annualised savings
Year 1
4.1×
ROI on engagement
Validated Q1 2026
Northrun
Real-time pricing agent that recomputes 320k SKUs every 4 minutes.
Pricing · Retail
Cohostly
Multimodal listing QA — image + copy — caught 11% policy drift on day one.
Marketplace · Trust & Safety
Ink Foundry
Internal RAG over 4,200 PDFs of regulatory filings. Saved a $1.4M outside-counsel line item.
Legal · RAG
04 — Studio
Founder · Applied research
ex-OpenAI alignment, ex-DeepMind
Principal · LLM engineering
ex-Anthropic, ex-Stripe ML
Principal · Product & design
ex-Figma, ex-Linear
Studio · Operations
ex-Stripe COO staff
05 — Book
You bring a problem you’re actually trying to solve. We bring a senior engineer and a clear yes-or-no by the end of the call.
Discovery calls are free and confidential. We respond within one business day.