“Build a diligence agent” sounds specific. It names a technology and an activity. It does not tell a product team which part of the work should change. In a fictional corporate practice, I would turn that request into a decision about the next evidence to buy. The goal is a pilot that can fail usefully, before the organization commits to a larger product.
One request can hide three different problems
Suppose attorneys spend too much time completing an agreement review. They might be struggling to locate the right documents, identify material deviations, or verify a draft report. A search tool, a triage queue and a drafting assistant address different bottlenecks. Choosing the agent first risks improving a step that was never the constraint.
I would observe the existing workflow and separate preparation, substantive review, corrections and escalation. Then I would ask the sponsor which outcome matters: earlier identification of material issues, less total attorney effort, or a faster client handoff. Those objectives may lead to different scope and different measures of success.
Test: Retrieval + eligibility
Test: Source-linked triage
Test: Evidence + review
Triage finds material exceptions earlier without increasing omissions. Compare total workflow time on held-out examples; revise or stop if either condition fails.
Make the first test small enough to reject
For this example, assume the observed bottleneck is prioritizing agreements for deeper review. The hypothesis is now bounded: a source-linked queue helps attorneys identify material exceptions earlier without increasing missed issues. Automated legal conclusions and a complete diligence report are outside the pilot. That boundary makes both the intended benefit and the retained responsibility easier to evaluate.
Before showing an assisted output, establish a baseline on an approved sample. Keep a separate held-out set for the comparison. Measure the entire workflow, including time spent checking citations and correcting the queue. Ask experienced reviewers to define a material omission and resolve disagreements about the rubric before scoring the results.
Write the stop decision before the success story
The pilot brief should name the sponsor, the reviewer, the operating owner and the person who can pause the experiment. Its decision rule should distinguish three outcomes: continue when the bounded quality and time criteria are met; revise when the evidence reveals a remediable workflow problem; stop when the proposed intervention does not address the bottleneck.
A real citation does not rescue an omitted exception. Faster preparation does not compensate for slower total review. And a passing pilot in one agreement family does not establish readiness across all matters. Those are separate questions, each deserving explicit evidence rather than a blended success score.
My recommendation to the committee would explain what the next funding buys, what remains uncertain and which alternatives we are deferring. That is also how I would compare this opportunity with a shared access foundation: the more visible feature should not automatically receive the next dollar.
An original decision framework with authored examples. These are not measured model results or descriptions of a particular firm.