Designing Data for Trusted Enterprise AI Decisions
As more enterprises use AI to support decisions, many find a mismatch between what AI requires and what their existing data architecture was built to deliver. These environments can move, store, and report information effectively, yet still struggle to establish which data carries authority, how it should be interpreted, and when it can be used. As AI moves deeper into enterprise workflows, those weaknesses can translate into operational risk.
Enterprise data architecture therefore needs to do more than connect systems and deliver data. It must provide the business context and controls that allow machines to interpret information reliably and act within defined boundaries. In that role, data architecture becomes the technical foundation for a broader decision infrastructure.
At its core, this model brings together four interdependent requirements: business meaning for consistent interpretation, decision fitness for reliable use, access and provenance for governed retrieval and traceability, and controls over the actions AI can take with that information.
Strong decision infrastructure is necessary for reliable enterprise AI, but it is not sufficient on its own. Model validation, prompt and orchestration controls, tool authorization, output evaluation, human oversight, and ongoing monitoring remain essential for governing how AI performs and acts.
The priority is to strengthen these decision-infrastructure capabilities in support of the AI strategy, trusted decisions, and broader reuse while giving the enterprise a consistent base as new AI use cases emerge.
AI Changes the Role of Data Architecture
Traditional enterprise data architecture has concentrated on systems of record, integration, warehousing, analytics, and reporting. Those capabilities remain essential, but they do not fully address the requirements of AI systems that must assemble and interpret information at the moment of a decision.
Generative AI applications may draw on policy documents, customer histories, product specifications, contracts, transcripts, and operational records. Agents can go further by using that information to update applications, initiate workflows, or interact with external parties. AI data architecture must therefore support controlled retrieval, trusted interpretation, and governed action alongside movement and storage.
Unstructured information exposes one of the clearest weaknesses in enterprise AI readiness. A recent global study found that only 29% of organizations fully knew where their AI-relevant unstructured data resided. Seventy percent said less than half was discoverable and usable for analytics or AI, while 35% could not trace how unstructured data was used across systems.
AI access creates risk in both directions. Agents may gain access to sensitive information that was technically available within enterprise systems but previously difficult to discover or assemble, increasing the risk of unintended exposure or inappropriate use. They may also lack access to relevant sources, leaving decisions based on incomplete context. NIST has highlighted related AI governance risks when agents interact with diverse datasets, tools, and applications, emphasizing identification, authorization, auditing, and attribution of agent actions.
The strategic issue is therefore no longer simply whether information can reach an AI system. Organizations also need to determine which sources should be trusted, who or what is authorized to retrieve them, whether information remains current, and how conflicting or incomplete context should be handled. Sensitive information must be classified, but classification alone is insufficient. Retrieval systems also need to carry user and agent identity, entitlements, and purpose restrictions so access controls remain enforceable as information moves downstream.
Documents require the same data governance rigor around ownership, classification, access, versioning, retention, and approval status as structured records. Enterprises also need reusable data capabilities rather than a new set of pipelines and retrieval stores for every AI use case. This reduces duplicate logic, inconsistent controls, and unnecessary maintenance. Domain teams should organize priority data around a common data strategy, with enterprise standards for meaning, quality, access, and provenance.
Key Takeaway
Enterprise AI readiness depends on trusted and appropriately controlled context at decision time, not just well-prepared training data.
Explore Our AI, Data & Cognitive Sciences Consulting Services
P&C Global’s AI, data, and cognitive sciences practice helps enterprises build the decision infrastructure trusted AI depends on.
Business Meaning Must Become Machine-Readable
Enterprise AI can connect information across systems, but it may interpret that information incorrectly when business meaning is unclear. Terms such as “customer,” “revenue,” “available inventory,” or “active contract” may vary across business units, channels, geographies, and applications. Individuals often reconcile those differences through judgment and context. Once AI begins joining records and making recommendations across them, inconsistent definitions can negatively affect business decisions.
The goal is not to force one definition across every context, but to establish which definition applies to each decision. “Customer,” for example, does not always refer to the same entity. In billing, it may mean the account holder or contracting entity. In healthcare, it could refer to the payer or beneficiary. In hospitality, it may be the traveler even when a corporate account pays for the stay. AI needs enough organizational and situational context to determine which meaning applies to the task or decision.
Organizations can make those distinctions explicit through master data governance that establishes agreed definitions, accountable business ownership, and ongoing data stewardship. Data dictionaries and business glossaries can codify those standards, while master and reference data, semantic models, common identifiers, ontologies, relationship graphs, and temporal metadata translate them into forms that systems and AI can apply consistently.
These disciplines predate AI, but AI raises the stakes by applying shared definitions across more systems, decisions, and workflows at greater speed and scale.
The need for governed business meaning does not stop at the enterprise boundary. As data moves across suppliers, partners, platforms, and regulatory ecosystems, organizations also need shared standards that allow information to be interpreted consistently across systems and institutions. The European Union’s Digital Product Passport Registry demonstrates how machine-readable meaning is becoming more important for cross-enterprise data exchange. Launched in July 2026, the Registry uses unique product identifiers, standardized metadata, verification mechanisms, and machine-readable data models and vocabularies across product groups. Harmonized standards extend that foundation to interoperability, APIs, data exchange, and storage.
As organizations embed AI across customer experience, supply chains, compliance, service delivery, and partner ecosystems, shared business meaning becomes essential to coordinated execution.
Business leaders must own this work. Technology teams can implement a semantic layer, but they cannot independently decide what constitutes a customer, when a product becomes sellable, which margin measure governs a pricing decision, or which policy applies to a transaction. Domain executives should define core concepts, relationships, exceptions, and decision contexts, while data and technology teams translate them into usable standards across the enterprise.
A durable semantic layer makes the broader decision infrastructure more adaptable. Models, vendors, interfaces, and retrieval technologies can change faster than many core business concepts. Keeping machine-readable business logic in an enterprise-controlled layer reduces the need to redefine those concepts as technologies change.
Key Takeaway
Consistent business definitions are essential to enterprise operations. Organizations need to define them clearly so people and AI systems interpret and apply them consistently.
Explore Our Modern Data Architecture Services
P&C Global’s modern data architecture services turn enterprise data into a foundation for trusted AI decisions.
Data Architecture Must Be Fit for the Decision
“Improve data quality” is too broad to guide enterprise AI investment. Not every decision requires the same level of data accuracy, freshness, or control. Leaders need to judge whether information is reliable enough for the decision it supports.
Data can be assessed across seven dimensions:
- Accuracy: Is the information correct enough for the intended use?
- Currency: Is it recent enough for the decision?
- Authority: Does it come from the source that should govern the decision?
- Provenance: Can the organization establish where it came from and how it changed?
- Coverage: Does it include the records, populations, conditions, and contextual signals necessary to support the decision without material gaps or distortion?
- Access: Is this person, system, or agent authorized to retrieve the information?
- Permitted use: Can the information be used for this purpose and by this system?
These dimensions can fail independently. An inventory balance may be accurate but too old to guide a restocking decision. A policy document may be authentic but superseded. A customer preference may remain valid for one use but not another. A dataset may be accurate and current while still omitting relevant customers, markets, operating conditions, or other context needed for the decision. Data quality cannot be judged in isolation from how the information will be used.
Article 10 of the EU AI Act reinforces this principle for covered high-risk AI systems. It requires training, validation, and testing datasets to be governed according to their intended purpose, including their origin, preparation, suitability, gaps, bias, and deployment context. NIST guidance likewise emphasizes data provenance and evaluation under real-world operating conditions.
The standard of assurance must reflect the potential impact of the decision. Low-impact recommendations can operate with lighter controls or require user confirmation, while systems that change prices, approve payments, release inventory, or modify customer records require stronger evidence and oversight. The greater the potential impact and the harder the action is to reverse, the higher the assurance standard.
Organizations should define an AI governance framework and assurance thresholds for priority decisions rather than applying the same data standard everywhere. Those thresholds may specify record age, approved sources, minimum completeness, required corroboration, permissible inferred attributes, and conditions for human review.
Measurement should evolve as well. Enterprise-wide defect counts say little about whether the data supporting critical AI decisions is reliable enough for use. More useful indicators include the share of high-consequence decisions supported by approved sources, the frequency of stale or conflicting information, and the time required to resolve exceptions.
Key Takeaway
Data quality should be judged against the decision it supports, with higher-consequence decisions requiring stronger evidence that the underlying information is reliable enough for use.
Data Reuse Requires Rights & Interoperability
The economic value of enterprise data grows when the same information can support multiple applications, AI systems, regulatory requirements, and partner interactions. A strong big data strategy therefore depends not only on collecting and processing more information, but on making it reusable across the enterprise and its ecosystem. Reuse, however, depends on more than technical access. It requires clear internal entitlements, enforceable legal and contractual usage rights, and mechanisms for governing information as it moves across organizational boundaries.
Within the enterprise, reuse begins with internal entitlements. Access controls should account for identity, role, sensitivity, system context, and business purpose rather than assuming that information available somewhere within the enterprise should be discoverable by every downstream system. As AI makes previously difficult-to-find information easier to locate and combine, effective entitlement management becomes increasingly important.
Authorized access does not automatically confer the right to use information for every purpose. Legal and contractual usage rights may be shaped by intellectual property, confidentiality, localization, retention requirements, customer choices, licensing terms, and other restrictions. Enterprises need mechanisms that preserve those conditions as data moves between systems and is incorporated into new applications or AI workflows. Organizations using external AI services, for example, need clear terms governing enterprise inputs, model outputs, derived information, retention, and reuse.
The challenge becomes more complex when information crosses organizational boundaries. Cross-enterprise exchange requires usage conditions, restrictions, and responsibilities to travel with the data so suppliers, customers, service providers, platforms, and other partners can interpret and apply them consistently.
The EU Data Act reinforces this shift. It gives consumers and businesses greater control over data generated by connected products, supports sharing with third parties, addresses unfair contractual restrictions, and establishes requirements around provider switching and interoperability. The European Commission has also published model contractual terms for data access and use, providing a framework for defining responsibilities alongside technical exchange.
These requirements have practical implications across business models. Manufacturers need to determine what connected-product data customers and service providers may access and use. Platform businesses must establish what partners may retain, derive, or redistribute. Data contracts can formalize permitted uses, access rights, retention, change requirements, and responsibilities across these relationships.
Technology choices matter as well. Enterprises that embed critical data rules too deeply within one provider can make future transitions slower and more costly. Portability should therefore extend beyond the data itself to the rules, permissions, and usage conditions that govern it.
Key Takeaway
Data reuse creates value only when access entitlements, usage rights, and exchange conditions remain enforceable as information moves across systems, providers, and partners.
AI Must Be Governed as a Data Producer
Most data-readiness programs focus on governing the information AI consumes. Decision infrastructure must also govern what AI creates, changes, and introduces into enterprise systems. Embedded agents can generate customer summaries, supplier classifications, forecasts, workflow updates, and transaction records that become new inputs for employees, applications, and other AI systems. Once AI begins contributing to the enterprise information base, data governance must address not only what systems consume, but what they create.
Organizations therefore need to distinguish information by how it was created: declared, observed, inferred, generated, or verified. An AI-produced attribute should not acquire the same authority as a confirmed customer instruction, approved accounting record, or validated operational event without appropriate review. Enterprises should record the sources behind machine-generated information, how it was produced, and whether it has been validated.
NIST’s Generative AI Profile emphasizes provenance tracking to establish the origin and history of AI-generated content. Research published in Nature has also shown that repeatedly training models on generated outputs can degrade the underlying data distribution over time.
The enterprise risk extends beyond model training. When generated or inferred information enters operational systems without clear data provenance, errors can circulate and appear to be independent evidence. A generated customer summary may influence a later recommendation, which creates another inferred attribute that feeds a subsequent model. Without traceability, organizations may struggle to determine whether multiple signals reflect separate facts or the same machine-generated assumption.
Data governance for write access requires a higher control standard than read access. Enterprises should separate permissions for retrieval and modification, validate high-impact outputs before they enter authoritative systems, and apply human review where the consequence or reversibility of a change warrants it. Versioning and rollback should make erroneous updates identifiable and reversible, while inferred and machine-generated records remain distinguishable throughout their lifecycle.
When AI can modify enterprise records, accountability must extend to the change itself. Organizations should be able to trace which agent made an update, the authority under which it acted, and the policy boundary that governed the action.
Key Takeaway
Enterprise decision infrastructure must govern not only what AI can read, but also what it can create, alter, and introduce into future decisions.
Decision Infrastructure in Practice
Airbnb offers a practical example of organizing data capabilities around business decisions rather than separate technical initiatives. Its Minerva metric platform standardizes definitions across analytics, reporting, and experimentation, helping teams work from consistent business logic.
That consistency is reinforced by data governance. Airbnb’s Metis platform captures ownership, certification, lineage, authorization, audit history, and approval workflows, while its Himeji system centralizes permissions at the data layer.
Airbnb also carries these controls into decision-making. Its experimentation framework uses company-level guardrail metrics to identify potentially harmful effects before product changes are launched, with threshold breaches triggering escalation and stakeholder review.
Together, these capabilities show how decision infrastructure can connect governance and technology around business decisions: shared definitions establish meaning, ownership and certification establish authority, permissions govern access, lineage preserves provenance, and guardrails constrain consequential action.
Key Takeaway
The value of decision infrastructure comes from connecting data governance to the decision itself, so trusted information and appropriate controls persist from definition through action.
Build Decision Infrastructure Around Business Decisions
Enterprise decision infrastructure should start with the decision and work backward to the information it requires. Leaders should identify the relevant data, determine which sources and definitions apply, and set the controls needed for reliable use. Data architecture provides the technical foundation, but the broader objective is to ensure that information can be interpreted, accessed, and acted on within defined boundaries.
The advantage comes from creating decision infrastructure that remains reliable as AI use cases, models, and technologies evolve. Enterprises that preserve trusted business context can adapt without having to redefine the information required for each decision. AI may shape the output, but trusted information determines whether the enterprise can reliably act on it.