Harvey Unveils Initial Results of Legal AI Benchmark LAB

Blockchain News
Harvey's Legal Agent Benchmark reveals frontier AI models complete less than 10% of complex legal tasks end-to-end, highlighting challenges in legal AI automation.

Summary

Harvey has released the first results from its Legal Agent Benchmark (LAB), an open-source framework designed to evaluate AI agents on complex, long-horizon legal tasks. The initial findings underscore significant limitations in current-generation AI models. Despite rapid advancements, frontier models completed less than 10% of LAB tasks end-to-end under a strict all-or-nothing evaluation standard. LAB evaluates AI agents across over 1,200 tasks spanning 24 legal practice areas. Each task mirrors real-world law firm workflows, requiring AI models to produce review-ready legal work products graded against 75,000 expert-created rubric criteria. Harvey's "all-pass" scoring system demands perfection—every rubric criterion must be satisfied for a task to pass. Among the evaluated models, Claude Opus 4.7 led with a 7.1% success rate, followed by Sonnet 4.6 at 5.4%, Opus 4.6 at 4.2%, GPT-5.5 at 2.1%, and Gemini 3.5 Flash at just 0.8%. The findings also revealed uneven competence across practice areas, with models displaying "jagged intelligence." No single model dominated across all categories, reinforcing the need for multi-model strategies in AI deployments. Another major hurdle is operational efficiency. The best-performing model, Opus 4.7, costs approximately $50.90 per task and has a latency of 22 minutes. Faster alternatives like Gemini 3.5 Flash offer lower latency but at the expense of accuracy. Harvey's study also analyzed agent behavior, identifying key patterns that improve task performance. The most effective agents demonstrated behaviors akin to those of skilled human associates: thorough research before drafting, post-draft validation, and iterative revisions. The benchmark's next phases will focus on expanding its task library, improving cost-efficiency, and fostering collaboration with AI labs to refine model performance.

(Source:Blockchain News)

Siasat.com

In search of a flat in Bengaluru, woman lands a job offer

Asianet Newsable

"This Can Only Happen in Bengaluru": Woman's Flat Hunt Turns Into Unexpected Job Offer, Post Goes Viral

Complete Ai Training

LexisNexis opens customer innovation lab to build legal AI with clients

Complete Ai Training

Anthropic legal team leans on AI for contract review and workflow automation

Aitechtrend

10 Best Contract Analytics Tools 2026 | AITechTrend

Pr Sync

Knovos to Showcase Latest Platform Innovations in Data Governance, AI-Assisted Search, and Workflow Automation at ILTACON 2026

Et Now

Bengaluru woman goes flat hunting, ends up meeting legal AI startup founders and gets a job offer after asking questions about their product

Newsbreak

📣 Several High-Paying Sales & Marketing Roles in Palo Alto

Newsbreak

⚖️ Wichita legal & compliance roles with multiple $100K+ options

Newsbreak

💻 Salt Lake City IT Hiring: Multiple Roles From $45/hr to $297K+

Blockchain News

Harvey Launches Tenet, Open-Weight Legal AI Model

Aba Journal

Why the next competitive advantage in the AI era is partnership, not just technology

South China Morning Post

OpenAI-backed legal tech firm pivots to Chinese Kimi K3 open-weight model

News 18

Individual Goes Flat Hunting In Bengaluru, Lands Job Offer After Chatting With Startup Founders

The Nassau Guardian

Before we remove the land, let us first consider what we could build upon it