The Eval Harness Behind Our Voice Agent
How we test a voice agent on real calls; score every call, calibrate the judges, and replay real failures before shipping.
Urvesh Patel / August 25, 2026
How we test a voice agent on real calls; score every call, calibrate the judges, and replay real failures before shipping.
Urvesh Patel / August 25, 2026
How Eve built Plaintiff Bench, a gold-standard evaluation dataset of 400+ plaintiff-law tasks, to catch citation, temporal reasoning, and retrieval failures before they reach lawyers.
How Eve uses Claude across legal workflows and internal engineering systems to accelerate settlements, improve case outcomes, and cut incident triage from hours to minutes.
Allen Deng / March 19, 2026
OpenSearch is built for a few large shards, but we needed thousands of small ones. Here’s our engineering deep dive on why we migrated to Turbopuffer.
David Zeng / August 20, 2025
Help us build the systems that turn legal work from hours into minutes. Join an ambitious team working at the frontier of AI and law.
See open roles