Harvey introduced Harvey Tenet, a post‑trained checkpoint of the open‑weight Kimi K3 model, trained with asynchronous reinforcement learning built by Fireworks. Training used sandboxed legal environments with partner‑style instructions, document sets, and expert rubrics of atomic pass/fail criteria. Data were synthetic, public legal, and human‑expert samples; no customer data was used. Optimization employed a rank‑64 LoRA over the full K3 network via GSPO, processing eight task groups of eight rollouts per step across roughly 1,750 environments and >10,000 rollouts per epoch. Reward combined rubric satisfaction, a count of resolved legal issues, and an all‑pass bonus, graded by an LLM‑as‑a‑judge.
On Harvey’s LAB, Tenet completes almost twice as many held‑out tasks as base K3, and on LAB: Contracts it improves 20 %, lifting the all‑pass rate by 9 and 2 points. Harvey reports state‑of‑the‑art on LAB: Contracts and second place overall. Gains transfer untrained to Mercor’s APEX Agents and Crosby’s Redline Bench, while knowledge‑benchmark scores (LegalBench, CUAD, MAUD, PRBench) stay unchanged, showing agentic training preserved core legal reasoning. Weights and API are not yet public; the release is a research preview from August 20 2026.
Why this matters
The results show that targeted post‑training of an open‑weight LLM can boost specialized legal‑agent performance without degrading general legal knowledge, offering a plausible path for firms to adapt models. However, because weights and API are not yet released, the ability to own and customize the model remains prospective.
Share this article
Found this insightful? Share it with your community on Reddit, X, or copy the link.
