Reading Data, Not Documents: Metadata in Digital Investigations

A contract arrives in production showing a January execution date. Both signatures are there, the terms are clean, and on its face the document is exactly what the other side says it is. Then the metadata comes back. The file was created in March. Substantial edits landed after the purported signing date. The final save was made by someone who was not a party to the agreement.

Nothing on the page changed. The case did.

This is what it means to treat discovery as a data problem rather than a document problem. The visible content of a file tells you what it says. The data surrounding that file tells you whether to believe it.

What metadata actually establishes

Metadata is the descriptive layer a system records about a file: who created it, when, what application produced it, when it was last modified, who saved it last, and what happened to it in between. Every enterprise application generates it, most users never see it, and almost nobody thinks to alter it.

That last point is why it carries evidential weight. Content is authored, which means content can be shaped. Metadata is generated as a byproduct of activity, which makes it far harder to curate after the fact. When testimony and metadata disagree, the metadata is usually the more reliable witness.

In practice it answers four questions that content alone cannot. Whether a document is authentic. Whether it was altered, and when. Who had access to it and when they exercised that access. And where a given file sits in a sequence of events.

The fourth question is the one that decides investigations.

Investigations are timeline problems

Most digital investigations are not really about a single document. They are about establishing an order of events precisely enough that the innocent explanation stops being plausible.

Consider a departing employee suspected of taking client data. The content of the files tells you what was taken. The metadata tells you the story.

Forty five minutes. Four artifacts, each drawn from a different system. Individually, any one of them has a harmless explanation. Assembled in sequence, they describe intent.

Notice that no single tool produced that table. It comes from correlating endpoint activity, removable device history, mail routing data, and document properties, which means it depends on all four sources having been preserved correctly before anyone started looking.

Where metadata goes missing

Metadata is fragile in a way that document content is not. You can lose it without ever touching what the file says, and you usually lose it during the steps meant to protect it.

The most common loss happens at collection, when files are gathered as PDFs rather than in native format. Conversion flattens a file into a rendering of itself. Author information, revision history, and system timestamps do not survive the trip. The document still reads correctly, which is precisely what makes the loss easy to miss until authenticity is challenged and there is nothing left to authenticate with.

The second loss happens through handling. Opening a file updates its last accessed timestamp. Copying it can rewrite file system records. Well intentioned review of source data by someone without forensic training will quietly overwrite the evidence of who touched what and when. This is why defensible collection favors imaging and minimal interaction with originals over convenience.

The third loss is jurisdictional. When data cannot leave a country, the temptation is to export a subset and work with it elsewhere. Every export is a conversion, and every conversion is an opportunity to strip the layer you may later need. Preservation has to happen where the data lives, which is an operational constraint rather than a technical one, and one that cross-border investigations surface almost immediately.

The cloud problem is the current problem

The failure modes above are well understood. The harder issue is that the most valuable metadata increasingly does not sit on a device at all.

Collaboration platforms changed what an attachment is. People no longer attach documents, they share links, and the collection consequences of that shift deserve their own treatment, which we have given them in Modern Attachments in eDiscovery. The metadata dimension is distinct and less discussed. Version history, sharing permissions, access logs, and edit attribution all live in the platform rather than in the file. Collect the message and you get a pointer. The record of who opened the document, who changed it, and which version the recipient actually saw stays behind.

Courts have been working through the linked file question since Nichols v. Noom in 2021, which declined to treat hyperlinked cloud files as traditional email attachments and pushed parties toward targeted requests rather than wholesale production of linked content. That framework has held. In James v. Cerebras Systems Inc. (N.D. Cal., July 2026), the court entered an ESI protocol built on the Northern District of California model order that addressed hyperlinked documents, short message data, and generative AI workflows together, and kept the targeted approach: a party identifying a specific produced document containing a relevant hyperlink requests that linked file separately.

The practical consequence is that cloud and hyperlink handling is now negotiated at the protocol stage rather than argued at the motion stage. Platforms are adapting as well. RelativityOne can now treat linked cloud files as associated files, which closes part of the gap between how people work and how discovery has historically modeled attachments.

None of that helps if the underlying platform metadata was never preserved. The protocol you negotiate is only as good as the collection that came before it.

When to bring in forensics

Not every matter needs a forensic examiner. Several situations reliably do.

Bring one in when the authenticity of a document is disputed, when there is reason to believe data was altered or deleted, when trade secret theft or employee misconduct is alleged, when the matter turns on the sequence of events rather than their content, or when relevant data sits in a jurisdiction with export and data sovereignty restrictions. In each case the risk is not that you will fail to find the evidence. It is that ordinary handling will degrade it before anyone realizes it mattered.

A forensic examiner does four things a standard collection does not. Preserves in a defensible manner that withstands challenge. Extracts metadata that standard processing discards. Correlates artifacts across systems into a single timeline. And explains the result in language a court will accept, which is a distinct skill from producing the analysis in the first place.

Where data cannot leave the country, that examiner has to be in the country. This is why in-region teams and local infrastructure matter more than tooling. Preservation has to happen where the data lives, not after an export has already stripped what mattered.

The last one is underrated. An unexplained timeline is not evidence. It is a spreadsheet.

What to do before it matters

The decisions that determine whether metadata survives are made early, usually before anyone has thought about authentication.

Before collection, identify where the data actually lives, including the cloud platforms that hold the version history rather than the files. Decide whether native format is required, and assume it is unless you have a specific reason otherwise. Understand which jurisdictions the data touches. Our Forensic Collections Guide 2026 sets out preservation methods source by source, including cloud platforms, mobile, and AI-generated content.

During collection, preserve originals, document chain of custody, and keep handling to a minimum. During review, read metadata alongside content rather than as a separate exercise, and treat timeline inconsistencies as leads rather than noise.

Metadata is invisible to almost everyone who generates it, and it is frequently the most revealing part of the record. As more evidence moves into collaborative and cloud environments, the gap widens between teams who plan for it and teams who discover its absence at the worst possible moment.

Content tells you what happened. Metadata tells you when, and by whom.

Talk to Lineal about how metadata should be preserved in your next collection, or see how we structure investigations where the sequence of events is the case. Our forensic teams operate across five continents and eleven data centers, with in-country collection where data cannot cross a border.

__

About Author

Laura Collins is an accomplished digital forensics examiner, currently serving as Vice President of Shared Services & Forensics at Lineal. With extensive experience overseeing global forensic operations, complex investigations, and eDiscovery delivery, she has built her career across corporate, legal, and incident‑response environments. Laura’s background spans hands‑on forensic analysis, major incident response, and leading high‑performing teams to deliver innovative, defensible solutions for clients worldwide. Recognised for her operational leadership and deep technical expertise, she is committed to advancing high‑quality forensic services while driving collaboration, efficiency, and excellence across the organisation.

__

About Lineal 

Lineal is an innovative eDiscovery and legal technology solutions company that empowers law firms and corporations with modern data management and review strategies. Established in 2009, Lineal specializes in comprehensive eDiscovery services, leveraging its proprietary technology suite, Amplify™  to enhance efficiency and accuracy in handling large volumes of electronic data. With a global presence and a team of experienced professionals, Lineal is dedicated to delivering custom-tailored solutions that drive optimal legal outcomes for its clients. For more information, visit lineal.com 

订阅我们的新闻通讯

    感谢您的订阅。

    您将获得实用洞察、产品更新以及您的团队真正可以使用的内容。