Don’t Ask AI to Solve a Data Problem
Why Better AI-Enabled Discovery Starts with Data, Not the Model
Artificial intelligence has created an understandable temptation in legal technology: when confronted with a large, complicated data set, apply the most powerful model available and let it sort things out.
As generative AI has become more capable, much of the discussion around AI-enabled discovery has centered on the model itself. Which model reasons best? Which produces the highest accuracy? Which handles legal nuance most effectively? Which platform should a legal team choose?
Those are legitimate questions. But they may also distract from a more fundamental one: Are we giving AI the right problem to solve?
In discovery and investigations, what appears to be an AI challenge is often, at least in part, a data challenge. The population may be unnecessarily large. Relevant information may be buried beneath redundant email, duplicates, automated content, irrelevant file types, and other noise. Important context may be fragmented across multiple systems. And the process of collecting and transforming data can sometimes obscure relationships that were clearer at the source.
A better AI model can improve the analysis. But it cannot make poor data strategy disappear.
In fact, the growing power of AI may make disciplined data strategy more important, not less.
Start With the Population, Not the Prompt
Imagine beginning a matter with one million potentially reviewable documents.
Today, it is technically possible to apply sophisticated AI across an extraordinarily large population. That capability can lead to a deceptively simple question: How can we use AI to review all one million documents?
A more valuable question may be: Why are we reviewing one million documents?
Discovery teams already possess an extensive toolkit for understanding and refining data before substantive review begins. Deduplication, email threading, metadata analysis, domain filtering, near-duplicate detection, communication analysis, conceptual analytics, and early case assessment can identify large portions of a population that do not warrant the same level of scrutiny.
At Lineal, that principle is foundational to our Amplify™ approach. Before deciding how documents should be reviewed, we first seek to understand the population itself, its composition, patterns, relationships, redundancies, and likely areas of significance.
In many matters, analytics and population refinement can result in only approximately 30% to 40% of the original document population requiring eyes-on review. The significance is not simply that fewer documents are reviewed. It is that human and technological resources can be concentrated on a smaller population with a higher likelihood of containing information that matters.
That becomes especially important when generative AI is layered into the workflow. A more focused population means the model is spending less time and cost evaluating material that mature analytics could already identify as redundant, irrelevant, or low value, and more of its reasoning capacity on the information most likely to require substantive analysis.
The 30% to 40% range is not, by itself, evidence that AI becomes more accurate. It is evidence of something more fundamental: the quality of the population presented to AI is something we can actively improve before the model ever begins its work.
The objective should not be to give AI as much data as possible.
It should be to give AI the right data.
Consider an exceptional attorney asked to investigate a complex dispute. One approach would be to deliver several million unsorted records with little explanation and ask, “What matters?”
A better approach would provide the allegations, relevant individuals, terminology, chronology, known facts, examples of significant documents, and an intelligently organized body of information.
The attorney has not changed.
The environment has.
The same is true of AI. AI performance is a workflow outcome, not simply a model characteristic.
From Data Reduction to Data Enrichment
Reducing the population is only the beginning. The next opportunity is to make the information that remains richer.
Who communicates most frequently? Which individuals connect otherwise separate groups? What concepts repeatedly appear together? Which documents are related? Where does activity suddenly increase? What happened immediately before or after an important event?
Those questions begin turning a document population into an information environment.
The same principle applies when generative AI is introduced. Asking whether a document is responsive is fundamentally different from giving a model the factual background of the matter, clearly defined legal issues, relevant terminology, important people and products, time periods of interest, examples of prior decisions, and explicit criteria for reaching a conclusion.
The difference is not simply a better prompt. It is better context.
Much has been written about prompt engineering. Increasingly, however, the more meaningful discipline may be context engineering: determining what information an AI system needs in order to reason effectively about the task in front of it.
The words we use to ask the question matter.
But what the AI knows when we ask it may matter even more.
What If We Could Start Closer to the Source?
There is an even larger opportunity ahead.
For decades, discovery has largely followed a familiar sequence: identify potentially relevant sources, collect broadly, process the information, move it into a review environment, and then begin narrowing the population.
AI creates the possibility of moving intelligence much earlier in that process.
Instead of automatically extracting enormous quantities of information from complex business systems, imagine being able to interrogate that information closer to where it originates.
Natural-language interaction could allow investigative teams to explore structured business data, financial information, communications, transaction histories, operational systems, and other complex repositories before unnecessarily broad downstream populations are created.
Rather than beginning with “Which documents are responsive?” teams could ask more fundamental investigative questions:
Who was involved? What changed? Which transactions were unusual? Where do particular people, accounts, communications, or events intersect? What activity occurred during the critical period? Which information actually warrants deeper collection and review?
This represents more than another improvement in document review. It suggests a shift from collecting first and understanding later to understanding earlier.
For investigations and litigation alike, that could materially change how scope is established, how collections are conducted, and ultimately how much information ever needs to enter a traditional discovery workflow.
Sometimes the most efficient document to review is the one that never needed to be collected in the first place.
Human Expertise Becomes the Orchestrator
None of this diminishes the role of people. It changes where their judgment creates the greatest value.
In traditional review, human expertise is often deployed at the document level, with teams repeatedly making similar decisions across thousands or millions of records.
In a more sophisticated AI-enabled workflow, experienced professionals move upstream. They define the investigative questions, understand the data sources, determine how the population should be refined, select the appropriate technology, build the context presented to AI, test its performance, examine disagreement populations, refine instructions, evaluate exceptions, and validate the results.
The goal is not simply to keep a “human in the loop.” It is to put human judgment where it matters most.
Where Discovery Goes From Here
At Lineal, we believe this is where AI-enabled discovery is heading.
The future is not simply about applying increasingly powerful models to increasingly large collections of data. It is about creating a more intelligent continuum between the underlying information, the technologies used to understand it, and the people responsible for making consequential legal decisions.
Today, that philosophy is reflected in our Amplify workflows: understand the population, remove unnecessary noise, enrich what remains, supply meaningful context, and then apply the right combination of analytics, AI, and human expertise.
Tomorrow, we believe that intelligence will move even closer to the source. Legal and investigative teams will increasingly be able to interrogate complex information earlier, understand relationships before broad collections are created, and focus downstream discovery on the data most likely to matter.
That changes the role of AI from simply reviewing what has already been collected to helping determine what should be collected, investigated, escalated, or ignored in the first place.
The competitive advantage will not belong simply to whoever has access to the newest model. Models will continue to change, improve, and leapfrog one another.
The advantage will belong to those who understand how to create the environment in which intelligence, human and artificial, can perform at its best.
Before asking AI to solve the problem, first make sure you have not handed it a data problem.
The model will change. The discipline should not.
__
About Author
Marco Nasca is the Vice President of Sales at Lineal and a 2001 graduate of DePaul University College of Law. For more than two decades, he has worked at the forefront of eDiscovery and legal technology, advising corporations and law firms on the defensible application of technology in complex litigation, investigations, and regulatory matters. His work focuses on structured data, emerging technologies, and the practical implications of evolving judicial doctrine.
__
About Lineal
Lineal is an innovative eDiscovery and legal technology solutions company that empowers law firms and corporations with modern data management and review strategies. Established in 2009, Lineal specializes in comprehensive eDiscovery services, leveraging its proprietary technology suite, Amplify™ to enhance efficiency and accuracy in handling large volumes of electronic data. With a global presence and a team of experienced professionals, Lineal is dedicated to delivering custom-tailored solutions that drive optimal legal outcomes for its clients. For more information, visit lineal.com
Índice
Inscreva-se na nossa newsletter
Obrigado por se inscrever.
Você receberá insights práticos, atualizações de produtos e conteúdos que sua equipe realmente pode usar.

