AI Due Diligence Software for Law Firms: Five Axes and a Bake-Off You Can Run Yourself
Key Takeaways
- •Diligence software and a general legal AI assistant are different purchases. One produces structured, citable findings across a corpus; the other answers questions one at a time. Buying the second for the first job is the most common and most expensive mistake in this category.
- •Run the bake-off on a closed deal where you already hold the answer key. Score recall against the issues your own team found, not against a vendor's demo corpus, and score it blind so nobody grades their preferred tool generously.
- •The decisive metric is the cost of checking. A tool whose findings take as long to verify as they would have taken to produce has saved nothing, which is why an uncited summary is unusable in diligence regardless of how good it reads.
- •Vendor accuracy percentages are unreproducible without a published task definition, corpus, and scoring rubric. Kira advertises 90 percent accuracy with none of the three, and no competitor publishes a methodology either, so the only number worth trusting is the one you measure.
- •A missed change of control provision is the firm's malpractice exposure, not the vendor's. That makes verifiable output a professional obligation question rather than a preference, and it is the reason to read the limitation of liability clause before the feature list.
AI due diligence software reviews the documents in a transaction and produces structured findings that trace back to a passage in a source document: extracted provisions, issue lists, disclosure schedules, and memos. It is a different product from a general legal AI assistant, which answers questions one at a time. That distinction decides most of the purchase, and it is the thing most buyer's guides in this category blur.
Disclosure first: I run Mage, which sells diligence software. Every ranking of this category is written by somebody with a product in it, including this one. The only defense is a method you can run without trusting the author, so most of what follows is a test protocol rather than a leaderboard. If you read nothing else, read the bake-off section and go run it.
What separates diligence software from a legal AI assistant?
The category splits on one question: does the tool make a claim about the whole document set, or only answer the question you happened to ask?
An assistant is question-shaped. You ask about a provision, it answers, and the session is only as good as your questions. Nothing in it tells you what you did not think to ask. Diligence software is corpus-shaped. It asserts that it looked at every agreement in scope, found every instance of the provision types you specified, and can show you where each came from. The second claim is harder to make and far easier to falsify, which is the point.
Here is how the significant products position themselves, in their own words, read on August 1, 2026.
| Product | How the vendor positions it | What that implies for a diligence buy |
|---|---|---|
| Harvey | AI software for legal and professional services, led by agents that execute legal work end to end, with Vault for bulk document analysis and a transactional workflow for due diligence and contract review | A firm-wide platform with a diligence surface. Compelling if the purchase is one system for the whole firm |
| Kira (Litera) | The number one contract intelligence solution for law firms, built on multi-layered AI combining generative and proprietary models | Structured extraction across large contract portfolios. Litera agreed to acquire Kira in an announcement dated August 10, 2021, so roadmap questions belong to Litera |
| Luminance | Legal-Grade AI and the legal brain of your organization, on a multi-model architecture spanning draft, negotiate, analyze, comply, and investigate | Its stated center of gravity is contract activity across a business rather than deal diligence. Ask specifically about the deal workflow |
| Datasite | The most trusted data room, first with multi-LLM access, offering semantic search, in-room AI with citations, and redaction at scale | Lives where the documents already sit. Scope is the room, an advantage for search and a constraint for deliverables |
| Mage | Transactional diligence covered end to end, from data room to closing, producing disclosure schedules, tabular review, and drafted memos | Ours. Judge it with the protocol below, and note the diligence platform has no command line or API surface |
Two framing points matter more than the table. Funding is not capability: Harvey announced a $200 million round at an $11 billion valuation co-led by GIC and Sequoia, which CNBC dates to March 25, 2026, and Luminance announced a $75 million Series C led by Point72 Private Investments on February 18, 2025, taking its trailing twelve month total above $115 million. Those numbers speak to durability and roadmap velocity, not to recall on your documents.
Neither does adoption. Kira states adoption among 70 percent of the top 50 global law firms, four of the five UK Magic Circle firms, and three of the Big Four accounting firms. Vendor-reported, undated, unaudited, and even at face value a statement about procurement rather than output. For the deeper comparison of the assistant, extraction, and infrastructure paradigms, see Harvey vs. Kira vs. infrastructure.
The five axes that decide the purchase
Every serious evaluation collapses to five questions. The rest of this guide is how to answer them with evidence instead of impressions.
Axis 1: accuracy you can check. Not the accuracy you are told. The only accuracy number worth anything is the one you measured on your own documents, and the only way to measure it is against a set of findings you already trust.
Axis 2: time to value. Processing speed rarely differentiates these tools; all of them beat manual review by an order of magnitude. What differs is how long before an associate produces usable output without a support call. A tool that runs in ten minutes and takes three weeks to configure is slower than one that runs in an hour and works on day one.
Axis 3: security and confidentiality. Baseline: SOC 2 Type II rather than Type I, encryption in transit and at rest, client data isolation, no training on your documents, and a deletion commitment you can point to in the contract. For sensitive matters, add data residency and audit logging of every access. Ask the training question in writing and accept only an unambiguous answer. More on what to demand is in SOC 2 and legal AI.
Axis 4: setup cost. Does the tool understand purchase agreements, employment agreements, IP assignments, leases, and credit facilities without you building templates? Who maintains those templates when they exist? A tool requiring a dedicated internal champion carries a headcount cost nobody puts in the business case.
Axis 5: output quality. A deliverable or a paragraph? Structured findings feed a schedule, an issue list, or a memo. Narrative summaries get re-read and re-extracted by an associate, which is where tools give back the time they saved. On the tracking side of that question, M&A transaction management software covers what belongs in the management layer rather than the review layer.
Four of the five are cheap to test. Security is a document review; setup cost, time to value, and output quality reveal themselves in an afternoon. Accuracy is the expensive one, and it is the one every vendor asks you to take on faith.
How do you run a bake-off on a deal you already closed?
The trick is to stop evaluating on new work and start evaluating on finished work, because finished work comes with an answer key.
- Pick a closed matter with a final work product. Ideal shape: a completed M&A deal, 100 to 400 documents, where your team produced an issue list or a set of disclosure schedules that a partner signed. Confidentiality first: confirm you can process those documents under your engagement terms and the vendor's data processing agreement, and use a matter where that is clean rather than arguing about a hard one.
- Freeze the answer key before you run anything. Write down the finding set your team actually produced: every change of control provision, every assignment restriction, every exclusivity, every indemnity cap, with the document and section for each. This list is the ground truth, and it must exist on paper before any tool sees the corpus. Freezing it afterward is how evaluations get corrupted.
- Load the identical corpus into each tool, with the same scope and the same provision list. If one tool gets a cleaner set because someone tidied it midway, the comparison is dead. Resist the urge to prompt your favorite well.
- Score blind. Strip the tool names from the outputs before anyone grades them. The person who championed a product will grade it generously without meaning to, and defeating exactly that bias is the point of the exercise.
- Time the verification separately from the run. Stopwatch on how long an associate takes to confirm 20 findings. This is the number that decides whether the tool is worth anything, and it is the one nobody measures.
- Have a second associate use each tool cold. Not the evaluator. Time from login to first useful output is your real onboarding cost.
Budget two to three hours per tool. That is less than a single vendor demo cycle and produces evidence instead of impressions.
What do you score, and what does failure look like?
| Criterion | How to measure it | What failure looks like |
|---|---|---|
| Recall against the answer key | Frozen findings the tool surfaced, divided by the total | Misses cluster in scans, amendments, and odd drafting. One missed change of control provision is a failing grade, not a rounding error |
| Precision on details | For each hit, check cap, basket, survival, carve-outs | Provision found, numbers wrong, associate re-reads every one anyway |
| Citation verifiability | Click 20 findings at random and time each click-through | The link opens a document rather than a passage, or opens nothing |
| Cost of checking | Stopwatch on verifying those 20 findings | Verification takes as long as the original review. Disqualifying |
| Behavior outside the trained set | Feed it an unusual agreement type deliberately | Confident output on a document type it does not handle, with no signal it is guessing |
| Deliverable readiness | Populate your standard schedule template from the output | Hours of reformatting, which is where the saved time goes back |
The fifth row is the one most evaluations skip. Extraction systems built on trained provision sets behave predictably inside that set and unpredictably outside it. Ask directly whether the system flags a document type as out of scope or answers anyway. A tool that fails loudly is safer than a tool that fails smoothly.
What does a citation-backed finding actually mean?
It means the software points at the document, the page, and the passage its claim came from, and clicking through lands you on that language.
Anything short of that is not a citation. A document-level reference tells you which of 300 files to read, which is where you started. The standard is functional: can an attorney confirm or reject the finding in seconds without leaving the tool?
This is why an uncited summary is unusable in diligence regardless of how good it is. Not because it is wrong, but because you cannot tell. An unverifiable finding has to be confirmed independently, which means the software produced a to-do list.
Vendor accuracy numbers deserve the same scrutiny. Kira advertises 90 percent accuracy from a hybrid of generative and proprietary models trained on 45,000 lawyer hours, on Litera's Kira product page accessed August 1, 2026. No task definition, corpus description, or scoring rubric accompanies it. That is not a criticism of Kira specifically, because nobody in this category publishes a methodology. It is a reason not to compare any two vendor percentages, and not to let a vendor compare its number against your measured one. The only reproducible number is yours.
Treat review-site scores with the same care. Datasite labels the G2 testimonials embedded on its own diligence page as incentivized, dated September and October 2025. Some of the review volume here is paid for, and a score restated on a vendor's page is not the artifact you think it is.
Who is responsible when the software misses a change of control provision?
The firm is, and no vendor page will tell you that.
Your obligations run to the client. The vendor's obligations run to you, in a contract that will cap liability at something like the fees you paid, exclude consequential damages, and disclaim any warranty that the output is complete. Read the limitation of liability clause before the feature list. It is the most informative page in the agreement.
That asymmetry is why verifiability is a professional question rather than a preference. You cannot supervise output you cannot check. If a finding cannot be traced to a passage in seconds, the only way to supervise it is to redo the work, and the tool has produced a suggestion rather than a work product. The signature on the memo is yours either way.
One practical consequence: write into your protocol which categories of finding require human confirmation before they reach a client deliverable, and keep the record of that confirmation. The question "how did you satisfy yourself" arrives after the deal, not during it.
What can you learn about price before you talk to sales?
Very little, which is itself information.
Datasite publishes no price on its diligence product page and routes to a quote request, with a trial offer of up to 90 days, accessed August 1, 2026. Ideals places its MCP connector and API integration add-on under the Enterprise tier of its pricing page with no figure attached, accessed August 1, 2026. That is the shape of the category: a quote gated behind a sales conversation, with feature access tiered.
So use the call for what it is good for. Three questions that produce usable answers:
- What changes when our deal count doubles? Per-matter and per-seat models diverge sharply at volume, and you want to know which cliff you are walking toward.
- What is included versus an add-on? Agent access, API access, and advanced analytics are commonly tiered rather than standard.
- What can we bill through to the client? Whether the cost lands on the firm's P&L or a disbursement line changes the internal politics of the purchase more than the number does.
How to model the return rather than negotiate the price is covered in the ROI of legal AI for M&A.
What should you pilot first?
Pick the narrowest workflow with the clearest answer key. For most M&A groups that is provision extraction across one deal's material agreements.
Do not pilot an open-ended assistant on a live matter. There is nothing to score it against, so the result is a collection of anecdotes and the loudest one wins. Do not pilot on a deal that is on fire either, because everyone will grade the tool on whether the week went well.
Start upstream of review if you can. Getting the documents into one place, typed, organized, indexed, and gap-checked is a mechanical problem with visible right answers, and it determines whether review starts on day one or day nine. Our data room is a separate product from our diligence platform: it is self-serve and free for a limited time, so that half can be tested without procurement. The diligence platform is a conversation, like everyone else's here. What the handoff into review looks like is in from data room to diligence workflow.
When you are ready to test the review layer, the protocol above is the whole ask: one closed deal, one frozen answer key, blind scoring, and a stopwatch on verification. We are happy to be measured that way and will run it on your documents, which you can start from a conversation with us. More of our writing on evaluating these tools, including how attorneys should evaluate LLM-powered tools, sits in the due diligence topic hub.
Frequently Asked Questions
What is AI due diligence software?
AI due diligence software reviews the documents in a transaction and produces structured findings that trace back to a specific passage in a source document: extracted provisions, issue lists, disclosure schedules, and memos. It differs from a general legal AI assistant, which responds to questions one at a time without guaranteeing coverage across the corpus. In a diligence context, coverage and citation matter more than conversational quality, because the deliverable is a claim about every document rather than an answer about one.
How do you evaluate AI due diligence software before buying it?
Run a bake-off on a deal you already closed, where your own work product is the answer key. Load the same document set into each tool, run the same scope, and score three things: recall against the issues your team actually found, whether every finding carries a citation you can open and confirm, and how long verification takes. Score it blind if you can, because the person who championed a tool will grade it generously without meaning to.
What does a citation-backed finding mean in diligence software?
It means the software points at the specific document, page, and passage its claim came from, and that clicking through lands you on that passage. A document-level reference is not a citation, because confirming it still requires reading the document. The test is simple: pick twenty findings at random, click each one, and count how many put the supporting language in front of you within a few seconds. Anything less than close to all of them means you are re-doing the review.
Who is responsible when AI diligence software misses a provision?
The firm is. Your professional obligations run to the client, and the vendor's contract runs to you with a liability cap that will typically be tied to fees paid. Read that clause before the feature list. This asymmetry is the reason verifiability is a professional question rather than a preference: you cannot supervise output you cannot check, and you remain answerable for work you signed.
What should a law firm pilot first with AI diligence software?
Pilot the narrowest workflow with the clearest answer key, which for most M&A groups is contract extraction across a single deal's material agreements. Avoid piloting an open-ended assistant on a live deal, because there is nothing to score it against and the result is anecdote. Start with the corpus step, which is cheap to test and easy to grade, then extend to schedules and memos once the extraction quality is known.
How much does AI due diligence software cost?
Almost nobody publishes a number. Datasite routes its diligence product to a quote request and offers a trial of up to 90 days, and Ideals places its MCP connector and API integration in its Enterprise tier without a published figure. Expect a per-seat or per-matter quote gated behind a sales conversation, and ask three questions on that call: what happens to the price when the deal count doubles, what is charged per matter versus per seat, and what the client can be billed for.
Ready to transform your diligence?
See how Mage can help your legal team work faster and more accurately.
Contact UsRelated Articles
Mage vs. Harvey: A Feature-by-Feature Comparison for M&A Counsel
An honest, sourced comparison of Mage and Harvey for M&A diligence work. Where each is built to win, where each falls short, and how to evaluate them on a real deal.
Mage vs. Luminance: How They Compare for M&A Diligence
Mage and Luminance both serve transactional teams with AI contract review. An honest comparison for M&A counsel — architecture, workflow, and where each is built to win.
Mage vs. Legora: How They Compare for M&A Counsel
Mage and Legora are both modern, LLM-native legal AI platforms but built for different scopes. An honest comparison for M&A practices choosing between firm-wide and specialist.