The Refinery
Trusted Data Collections
Preserving metadata provenance for AI-Ready Data. Enterprise content captured together with its metadata, ownership, security attributes, and chain of custody before it ever reaches a downstream platform.
Trusted Data Collections
Organizations are investing heavily in Microsoft Fabric, Snowflake, Databricks, and generative AI. The value of these platforms depends entirely on the quality and trustworthiness of the data entering them. Most organizations focus on moving data into modern platforms while overlooking the metadata, lineage, and business context that make that data trustworthy.
Executive Summary
Enterprise information often arrives without the ownership, security, classification, lineage, and business context needed to determine whether it is accurate, relevant, compliant, or appropriate for AI. When this context is missing, organizations are left to reconstruct it after ingestion, an expensive and often incomplete process that erodes confidence in analytics and increases governance and compliance risk.
Zantaz Trusted Data Collections solve this by capturing enterprise content together with its metadata, ownership, security attributes, and chain of custody before it ever reaches a downstream platform. Through intelligent metadata enrichment, enterprise information is discovered, classified, and organized into governed collections that preserve their original provenance.
The result is not simply another copy of the company's files. It is a curated, policy-aligned, AI-ready data asset that can be traced to its source, governed throughout its lifecycle, and used with confidence across analytics, compliance, eDiscovery, enterprise search, and artificial intelligence.
The Business Challenge
Enterprise information is distributed across Microsoft 365, SharePoint, Exchange, OneDrive, network file shares, cloud storage, collaboration platforms, and legacy archives. Within these environments, organizations frequently encounter redundant, obsolete, and trivial information, unknown or outdated ownership, inconsistent classifications, sensitive information stored in unexpected locations, files with missing business context, broken lineage, data retained beyond its business or regulatory value, and content that was never intended for analytics or AI use.
Moving this information directly into a modern data platform does not automatically make it trustworthy. It can simply transfer existing data problems into a more powerful, and more expensive, environment. Without trusted metadata and provenance, organizations struggle to answer fundamental questions:
- Where did this information originate, and has it been changed?
- Who owns it, and what permissions originally applied to it?
- Does it contain sensitive or regulated information?
- Why was it included in this dataset, and is it current and relevant?
- Can it safely be used by an AI application?
- Can an analytical result be traced back to supporting evidence?
These questions become especially important when AI-generated answers, regulatory decisions, legal matters, or executive reporting depend on the underlying data.
What a Trusted Data Collection Is
A Trusted Data Collection is a curated, policy-aligned, metadata-enriched dataset created for a defined business purpose. It may be organized around a customer, legal matter, department, project, region, investigation, regulatory requirement, AI use case, or other business need. Using the Trusted Data Refinery, enterprise information is identified, analyzed, classified, and refined before it is made available to a downstream platform. Rather than simply copying files, Zantaz captures:
- Content: The original enterprise document.
- Metadata: File system attributes, ownership, permissions, and timestamps.
- Provenance: Complete chain of custody and original source information.
- Security Context: Access permissions and security attributes.
- Business Context: Classification, ownership, and governance metadata.
In practice, each collection can preserve the original source system and location, document or record identifiers, ownership and custodians, creation and modification and collection timestamps, original permissions and security context, sensitive-data and business classifications, integrity identifiers and file hashes, collection and chain-of-custody history, the policies and classification rules applied, transformation and processing history, and Trusted Collection identifiers.
The original content remains connected to the metadata that explains what it is, where it came from, why it was selected, and how it may be used, giving organizations a comprehensive understanding of every object in their enterprise rather than relying on basic file properties alone.
How Provenance Continues Downstream
Zantaz establishes provenance at the source. Microsoft Fabric, Snowflake, and Databricks then preserve and extend that context as the data is processed. When a Trusted Data Collection is projected into a downstream environment, its provenance attributes can be represented through record-level fields, structured metadata, linked manifests, catalog entries, tags, or external-lineage relationships. The destination platform then adds operational lineage as the data is loaded, transformed, queried, shared, and consumed.
This creates two complementary layers of trust. Source provenance explains the information's original location, ownership, security context, integrity, classification, and collection history. Platform lineage shows how the information was transformed and which tables, pipelines, notebooks, models, reports, or applications used it. Together, these layers create traceability from the original enterprise source through analytics and AI consumption.
What Organizations Do with Trusted Data Collections
AI performs best when it receives relevant, accurate, and well-contextualized information. Trusted Data Collections allow organizations to provide copilots, agents, retrieval-augmented generation applications, and enterprise search tools with curated information instead of exposing them to the entire uncontrolled data estate. AI does not always need more data. It needs the right data, with the context required to interpret it responsibly.
- Reusable Collections: Build a collection once for a defined business purpose and make it available to multiple authorized workloads instead of rediscovering the same information for every project.
- Explainability: Determine which source records contributed to an output, where they originated, who owned them, what classifications applied, and which transformations occurred downstream.
- Governance and Compliance: Identify PII, PHI, PCI, and other sensitive information, apply retention and disposition requirements, and restrict what can enter an AI environment.
- eDiscovery and Investigations: Search and organize by custodian, source, date, matter, classification, and collection history while reducing the volume sent into review.
- Enterprise Search: Discover information through business meaning, including ownership, department, customer, matter, region, sensitivity, record type, and relevant dates.
- Reduced Processing: Identify ROT, duplicates, expired records, and irrelevant content before ingestion so expensive platform resources focus on information with a defined business purpose.
- Data-Quality Investigations: Trace a questionable dashboard, report, model, or AI answer back to the original source record, owner, collection date, policies applied, and downstream transformations.
Value Across the Modern Data Stack
Microsoft Fabric. Collections can be projected into OneLake with their source identifiers, classifications, ownership, security context, and collection history. Fabric can then use these governed collections across Lakehouses, analytics, Power BI, and AI workloads, while Microsoft Purview extends visibility by cataloging Fabric assets and recording supported downstream lineage. The result is more useful Lakehouse data, less low-value content entering Fabric, better Copilot grounding, and stronger traceability between source information and business outputs.
Snowflake. Collections provide curated, provenance-aware information rather than raw enterprise content. Provenance attributes can be mapped into table fields, classifications, object tags, catalog metadata, or supported external-lineage relationships, producing more reliable data products, stronger policy enforcement, more secure sharing, and better traceability from Snowflake outputs to original business records.
Databricks. Collections provide the Lakehouse with a governed starting point for data engineering, analytics, machine learning, and AI. Source provenance is maintained in Delta tables and governed through Unity Catalog metadata, tags, properties, and external-lineage relationships, while Databricks captures additional runtime lineage. The result is a governed bronze layer instead of an uncontrolled landing zone, better inputs for feature engineering, and more explainable models.
Why Zantaz
Modern platforms are designed to process, analyze, and activate data, but they cannot automatically determine whether a source document is relevant, properly owned, appropriately retained, accurately classified, or safe for a particular AI use case. Without a trusted preparation layer, organizations are forced to solve these problems after ingestion, once information has already been copied, transformed, or exposed to additional systems.
Trusted Data Collections are powered by the Trusted Data Refinery, which combines enterprise-scale discovery, intelligent classification, metadata enrichment, ROT identification, and policy-driven governance. Rather than acting only as a catalog of what an organization has, Zantaz transforms distributed enterprise information into curated collections that are ready for approved business use.
What Changes for the Customer
- Copilot becomes governed, explainable, and defensible.
- Fabric becomes productive instead of expensive, as compute waste falls and query precision rises.
- Purview becomes more powerful because the underlying data is healthier.
- OneLake becomes an intelligence surface, not a storage swamp.
- AI applications use curated, policy-aligned information instead of uncontrolled data sprawl.
- Business users find relevant information faster.
- Compliance teams gain stronger evidence of origin, ownership, and handling.
- Legal teams create more focused and defensible collections.
- Data teams spend less time reconstructing missing context.
- Organizations trace important outputs back to the records that supported them.
Key Benefits
- Preserved Provenance: Source location, ownership, permissions, integrity identifiers, and chain of custody travel with the content.
- Explainability: Every document traceable to the logic and policy that placed it in its collection.
- Policy Alignment: Organized around governance, compliance, and security requirements, not arbitrary folder structures.
- AI Precision: Copilot and AI agents reason over curated collections, reducing ungrounded answers and wasted compute.
Conclusion
Trusted Data Collections move governance earlier in the lifecycle. They help organizations determine what data should be used, what should be excluded, what protections should apply, and what evidence must remain connected to it. Microsoft builds the platform. Zantaz prepares the data the platform depends on to deliver AI-Ready Trusted Data @ Speed and Scale.

Metadata Provenance Through the Modern Data Stack
Provenance is established at the source and carried all the way to the outcome.
Information where it already lives, across the Microsoft estate and beyond.
- Microsoft 365 and SharePoint
- Exchange and OneDrive
- Windows File Shares
- Cloud storage and collaboration platforms
- Legacy archives
Work through the interactive walkthrough before your next conversation.
The guide turns this material into a step-by-step sequence, with an interactive provenance pipeline, a collection builder, and clear answers to the questions this topic always produces.
