Enterprise Context
The Missing Operating Layer for Reliable AI
Why 95% of enterprise AI pilots fail to move the needle — and the reference architecture, governance model, and 90-day roadmap to fix it.
Why 95 percent of enterprise generative AI pilots still aren’t paying off — and the unglamorous operating layer the other 5 percent quietly built first.
Somewhere in your organization right now, there is an AI pilot that looked fantastic in the demo. The model summarized the contract correctly. It answered the policy question in a friendly, confident sentence. Everyone in the room nodded. And somewhere else in your organization, that same pilot is quietly stalling — not because the model got worse, but because the demo’s easy conditions never look like Tuesday afternoon production traffic, permissioned data, and a compliance reviewer asking, “where did that number come from?”
Here’s the version of this story we hear most often. A team spends six weeks building something genuinely impressive: a chat assistant that can answer questions about company policy, pull up account history, or draft a first response to a customer. Leadership sees the demo and loves it. Funding gets approved for a wider rollout. Then, three weeks into that rollout, someone asks the assistant a slightly different version of the demo question — a real one, from a real account, with a real edge case — and it gets confidently, embarrassingly wrong. Not because anyone did anything foolish. Because the demo was never actually testing the thing that breaks in production.
If that sounds familiar, you are not alone, and you are not doing anything unusually wrong. You are experiencing the single most common failure pattern in enterprise AI today, and it has a name, a cause, and — this is the good news — a fix that does not require a bigger model or a bigger budget.
| 95% | 60% | 17–33% |
|---|---|---|
| of enterprise GenAI pilots show no measurable P&L impact (MIT, 2025) |
of AI projects lacking AI-ready data will be abandoned through 2026 (Gartner) |
hallucination rate in commercial RAG tools marketed as “hallucination-free” (Stanford) |
Figures drawn from the sources cited in the full white paper’s appendix.
This post is the fast version of a much longer argument we make in our new white paper, Enterprise Context: The Missing Operating Layer for Reliable AI. If you read nothing else, read this: the gap between a slick AI demo and a trustworthy production system is almost never the model. It is the absence of what we call enterprise context — a governed capability for assembling business meaning, relationships, permissions, and current state at the exact moment a decision or action is needed. Below, we’ll walk through why that gap opened up, what it actually looks like when it bites you, and the disciplined, unglamorous path the organizations in that successful 5 percent are actually taking.
The Wall Every Enterprise Eventually Hits
For the first couple of years of the generative AI cycle, the interesting question was whether large language models were good enough. That question is mostly settled. The interesting question now is why a model that is clearly good enough keeps producing unreliable results the moment it leaves a curated demo and touches a real workflow.
The data on this is not subtle. A 2025 MIT study that looked at 300 public AI deployments, surveyed 153 executives, and interviewed 52 organizations found that 95 percent of generative AI pilots were producing no measurable effect on profit and loss — this despite an estimated $30 to $40 billion in enterprise investment. Separately, Gartner’s research into enterprise data readiness found that 63 percent of data management leaders either lacked, or were not sure they had, the right data practices to support AI, and Gartner’s resulting forecast is that 60 percent of AI projects without AI-ready data will be abandoned through 2026.
Put those two data points together and a pattern falls out immediately: the problem showing up as “the AI doesn’t work” is, underneath, almost always a data and governance problem wearing an AI costume. A demo only has to be right once, in front of a friendly audience, on a cherry-picked example. Production has to be right on the messy, contradictory, permissioned version of the same question — asked by someone whose access rights actually matter, about a policy that changed last week, pulling from three systems that each define “customer” slightly differently.
“Generic tools work well for individuals precisely because they’re flexible and undemanding — and they fail in the enterprise for exactly the same reason.”
That line, paraphrased from the MIT research, is worth sitting with. A general-purpose assistant that can talk about anything is a wonderful personal tool and a fragile enterprise one, because it was never built to learn your organization’s specific data, rules, and workflows. It performs the same in the pilot as it will in production — which is exactly the problem, because production is where the fragility shows up.
It’s Not the Model. It’s the Context.
So what is “context,” exactly, and why does it deserve to be treated as its own capability rather than a feature of whichever AI platform you buy?
Here’s the definition we use in the white paper: enterprise context is the governed combination of business meaning, relationships, rules, permissions, history, and current state that a system needs to make a decision or take an action correctly, safely, and traceably.
Every word there is doing work, but two matter most. “Governed” means someone is accountable for whether a given piece of context is accurate, current, and permitted for a given use. Context without governance is just more data sitting in more places. And “combination” means no single technology supplies it — context is assembled fresh, at the moment of use, from several capabilities working together.
This is where a lot of well-intentioned AI programs quietly go wrong: they buy or build one piece of the puzzle and treat it as the whole picture. A vector database is not context. A knowledge graph is not context. Retrieval-augmented generation (RAG) is not context. Each of these is a legitimate, often necessary component — but none of them, alone, is a governed context capability, and treating any one of them as if it were is exactly how a team ends up with a system that is fast, fluent, and occasionally, expensively wrong.
There’s a simple three-question test we use to tell the difference between “we have data” and “we have governed context.” For any fact your system uses in an answer, can you answer all three of these:
- Provenance — can the system name the source, its owner, and the date it was last verified?
- Permission — was the specific requester allowed to see that source, under the rules that apply to them personally, not a generic admin account?
- Currency — is there a mechanism that flags when a source has gone stale relative to how often the underlying policy or fact actually changes?
If your system can’t answer all three for a given response, it’s operating on data, not governed context — no matter how sophisticated the retrieval or generation technology looks from the outside.
“Won’t the Next Model Just Fix This?”
This is the objection we hear most from technically sophisticated teams, and it deserves a straight answer: no, not on its own, and betting your roadmap on it is a mistake we’d gently steer you away from.
A more capable model genuinely does help. It reasons better over messy input, follows instructions more reliably, and handles longer context windows than it used to. None of that changes whether the model was given the right facts in the first place. A smarter model reading the wrong version of a policy document will give you a smarter-sounding wrong answer. A more capable model with no entitlement enforcement will confidently disclose something it should not have seen. Model quality raises the ceiling on what’s possible; it does nothing to fix a leaking data pipe underneath it.
Think of it this way: if you handed a brilliant new analyst a stack of outdated, unlabeled, sometimes-contradictory binders and told them to answer customer questions immediately, no amount of raw intelligence would save them from getting some of those answers wrong. The problem isn’t the analyst. It’s what they were handed. Every model upgrade cycle, some portion of the enterprise AI market re-learns this lesson the expensive way, upgrades the model, sees the same failure modes reappear in a new form, and only then goes looking for the actual cause.
The Shadow AI Problem You Probably Already Have
There’s a second-order cost to a context capability that doesn’t work well, and it rarely shows up on the same budget line as the failed pilot. When the sanctioned AI tool is too brittle or too limited to be genuinely useful, employees don’t stop using AI — they route around the sanctioned tool and start pasting company information into whatever consumer assistant is on their personal device. MIT’s research has a name for this: the shadow AI economy, and it is a rational response to a broken supply of context, not a discipline problem among your employees.
That’s a governance and security exposure hiding behind what looks, on the surface, like an adoption success story — people really are using AI, just not the version anyone approved, reviewed, or can audit. Fixing the underlying context capability doesn’t just make your sanctioned tool better. It removes the reason your best people quietly went looking for an alternative in the first place.
If you want a quick, low-drama way to size this risk in your own organization, don’t start with a survey — start with your network and endpoint logs. Traffic to consumer AI domains from corporate devices, especially from your highest performers, is usually a more honest signal than any self-reported usage question. Most leaders who run this check are surprised twice: once by how much usage is already happening, and again by how little of it touches anything the security team ever reviewed.
Nine Ways This Breaks in Production (and You’ve Probably Already Seen a Few)
Most leaders have already lived through these failures without anyone in the room calling them a context problem. Naming them correctly matters, because it tells you which layer to fix instead of quietly blaming the model and requesting a bigger one.
- Confidently wrong answers — the system states something false with the same fluent tone it uses for something true, because nothing in the pipeline checked provenance before generating.
- Stale policy cited as current — a customer or employee is told an outdated rule applies, because nothing tracked when the source document was superseded.
- Conflicting definitions across departments — two teams get two correct-sounding answers to the same question, because “active customer” means something different in each system of record.
- Missing entitlements enforced — retrieval indexed content without carrying forward the access controls the source system already had, creating real regulatory exposure.
- Broken handoffs between systems — the workflow needs a human to manually bridge two systems, quietly defeating the productivity case the pilot was funded on.
- Untraceable recommendations — nobody can reconstruct which passages fed which part of an answer, which is a serious problem the first time a regulator or an unhappy customer asks.
- Local optimization — every business unit builds its own retrieval and its own glossary, multiplying cost and multiplying the number of slightly different “correct” answers floating around the company.
- Runaway retrieval and token cost — broad, imprecise retrieval compensates for weak relevance ranking, and the bill grows faster than anyone budgeted for.
- Loss of learning after deployment — the system that looked great at launch drifts as source systems change and nobody re-evaluates it, so the first person to notice the drift is a customer or a regulator.
None of these is a reason to avoid production AI. They’re a reason to treat context the way your organization already treats identity and access management: as an engineered capability with real owners, not an afterthought bolted onto a chat interface.
Why This Is Genuinely Hard — and Where the Real Trade-Offs Are
It would be dishonest to pretend building a governed context capability is easy, and a program that pretends otherwise usually doesn’t survive its first budget review. The good news is that most of the hardest conversations here are trade-offs with a defensible answer, not open problems with no solution.
Centralization Versus Federation
A central platform team can enforce consistent policy and cut down on duplicated work. But the business meaning of a “vendor” or a “claim” genuinely lives with the domain teams who create that data. The workable pattern in practice is federated ownership on top of a centralized platform — not one team trying to own everything, and not every team building its own silo.
Precision Versus Coverage
Narrow, well-curated sources produce more reliable answers. Broad indexing improves coverage but drags in noise, stale content, and material nobody meant to expose. Nearly every successful program we’ve seen starts narrow on purpose and expands only after the narrow version proves itself.
Speed Versus Governance
Skipping access-policy review to hit a pilot deadline is exactly how permission gaps end up quietly running in production. The 90-day roadmap later in this post — and in full detail in the white paper — is built specifically to preserve governance discipline without sacrificing a real, fundable proof point.
If you remember one uncomfortable fact from this section, make it this one: in Gartner’s 2024 survey, 63 percent of data management leaders admitted they lacked, or weren’t sure they had, the right data practices for AI. The highest-impact, highest-likelihood hurdles to a context program are organizational and data problems. No model upgrade fixes an ownership gap.
What a Governed Context Capability Actually Looks Like
In the white paper, we lay out a full, vendor-neutral reference architecture — five layers, each with clear responsibilities, interfaces, and a trust boundary marking where data-residency and regulatory obligations most often bite. Here’s the short version.
- Systems of record and content repositories — owned by the business, not the AI platform. Your CRM, ERP, contract repository, and identity systems stay exactly where they are.
- Ingestion, metadata, and a semantic layer — federates access, builds a metadata catalog with lineage, and defines what “active customer” actually means, once, for the whole organization.
- Context assembly, retrieval, and policy enforcement — the platform’s core. Retrieval finds candidate content; a knowledge graph resolves relational questions where that’s genuinely justified; identity and policy enforcement filters everything by the real requester’s entitlements; and everything gets tagged with provenance before it goes anywhere near a model.
- Agent or application runtime — where a grounded response or action actually happens, with an explicit refusal, clarification, and escalation path. A system that always answers is a system that will eventually answer wrong with total confidence.
- Evaluation, observability, and governance — closes the loop with regression testing, human feedback capture, audit logs, and cost monitoring, so drift gets caught by your team instead of by a customer.
One question we get constantly: do we need a knowledge graph? Usually, not at first. A graph earns its cost when your questions are inherently relational — who reports to whom, which entities share a beneficial owner, which contract clauses depend on which defined terms. If your workflow is fundamentally lookup-and-summarize, flat retrieval over a well-curated, permissioned index is usually enough, and the graph should wait until a specific relational question is actually blocking you.
A Worked Example: The $85,000 Purchase Order
Architecture diagrams describe capability. It’s more useful to see what actually happens on one real request. In the white paper we trace a procurement-policy assistant handling a common, high-stakes question: does an $85,000 purchase order require a competitive bid under current policy?
The system authenticates the requester and resolves their actual spend authority — not just their login. It retrieves the current procurement policy, category thresholds, and this specific vendor’s contract terms, filtered by entitlement. It assembles that into a single context bundle, tagging every fact with its source and date. Critically, before it answers, it confirms the retrieved policy version hasn’t been superseded — and if it can’t confirm that, the correct behavior is to say so, not to answer confidently with whatever it happened to retrieve.
Then it hits a decision point. If the system is confident and the request is clearly within an unambiguous threshold, it answers directly, with citations. If it’s ambiguous, above a threshold, or involves an unfamiliar vendor, it escalates to a human reviewer before anything is implied to be approved. At no point does the system commit spend or sign anything on its own — its job is to inform, not to act, which keeps decision authority exactly where your existing process already puts it.
That’s what “governed” looks like in practice: not a system that never makes mistakes, but a system that knows the difference between a question it’s allowed to answer confidently and one it needs to hand off.
Notice, too, what this example does not require. It does not require a knowledge graph — this is a lookup-and-threshold question, not a relationship-mapping one. It does not require a multi-agent orchestration layer. It requires disciplined retrieval, entitlement enforcement, a freshness check, and an honest confidence threshold. That’s a deliberately modest technical footprint, and it’s exactly why it’s achievable inside a 90-day window rather than a multi-year platform build. The organizations that get this right resist the temptation to reach for the most sophisticated architecture available and instead reach for the smallest architecture that actually answers the question correctly and safely.
The Cost of Waiting Is Not Zero
It’s tempting to read all of this and conclude that the safe move is to wait — let the tooling mature, let a competitor work out the kinks, revisit this next budget cycle. We’d push back on that instinct, gently but firmly, for three reasons.
First, every pilot that fails on data quality, permissions, or stale content doesn’t just cost that pilot’s budget. It costs the organizational trust the next pilot needs in order to get funded at all. We’ve watched capable teams spend a year rebuilding credibility after one bad launch, time that a disciplined first attempt would never have lost. Second, the shadow AI problem described earlier doesn’t wait for your roadmap — employees are already finding workarounds today, with or without your approval, and every month that continues is a month of ungoverned company data flowing somewhere you can’t see it. Third, and increasingly important: the regulatory floor is rising underneath you regardless of when you choose to act. Texas’s AI governance law is already in effect. Colorado’s framework is already evolving toward 2027. Waiting doesn’t pause the compliance clock; it just means you’ll eventually build the same governance capability under more pressure, with less runway, after a law already applies to what you’re doing.
None of this is an argument for moving recklessly fast, either. It’s an argument for moving deliberately now, on a scope small enough to get right, rather than either freezing or overreaching. The 90-day roadmap above exists precisely because “start now, start narrow” and “don’t cut governance corners” are not actually in tension — you just have to plan for both from day one instead of choosing one now and bolting on the other later.
Proving Value Without ROI Theater
A credible business case for context doesn’t promise a single, universal ROI number, because the return depends entirely on the workflow, the baseline, and how disciplined the measurement is. What it should do is map specific outcomes — reduced search and rework, improved decision quality, faster cycle times, fewer compliance exceptions, better experience, and reuse across future use cases — to leading and lagging indicators, an honest baseline, and the confounders that could fool you into over- or under-crediting the system.
Here’s a deliberately modest, illustrative example, not a benchmark to cite: imagine category managers handle 4,000 procurement policy questions a year, each currently taking about 18 minutes to research manually. If a governed context capability resolves 60 percent of those directly and cuts handling time on the rest by a third, that’s roughly 700 hours of category-manager time returned annually. Whether that converts into savings, capacity, or faster cycles is a call for the business owner to make — the point of the exercise is showing your math, not showing an impressive number.
Worth noting: MIT’s research found that the largest share of GenAI budgets goes toward the most visible use cases, like sales and marketing tools — while the strongest measured returns actually came from back-office automation that received far less investment and attention. Argue your value case workflow by workflow, not enterprise-wide, and you’ll avoid that trap.
It also helps to be explicit, out loud, with your finance partners about what you are not claiming. You are not claiming that context will transform the P&L on day one. You are claiming that a specific workflow, measured against a specific baseline, with specific confounders named up front, either improved or it didn’t — and that this discipline is exactly what lets a second and third workflow get funded on real evidence instead of enthusiasm. Enthusiasm got the first 95 percent of pilots approved. It has not been enough to get them to production value.
Governance Can’t Be Bolted On After the Demo Works
Here’s a pattern we see constantly: governance gets treated as a phase-two problem, something to formalize once the pilot proves itself. This is backwards, and it’s backwards in a specific, costly way. Retrofitting governance onto a system your users already trust is exactly when access-policy gaps, unowned semantic drift, and unauditable decisions tend to surface — usually during an incident, not during a calm, planned review.
The fix is a short, explicit list of decision rights, established before the second pilot starts: a business owner accountable for outcomes, a domain steward for each content area, a data or knowledge owner for the underlying systems, a security or privacy owner for entitlements, an AI product owner for end-to-end behavior, a platform team, and risk or compliance for auditability. None of this is bureaucracy for its own sake — it’s the same discipline that eventually made identity and access management unavoidable rather than optional, twenty years ago, applied now to the facts an AI system is allowed to act on.
Two decision types deserve special attention because they’re the ones most often skipped under deadline pressure. Access-policy changes should never go to production without security or privacy sign-off, even for a small pilot group — pilots are exactly where entitlement mistakes are least likely to be noticed, because usage is light and informal. And loosening an evaluation threshold to make a pilot look more successful should require joint sign-off from the platform team and the business owner, so a team under pressure can’t quietly redefine what “good” means.
None of this needs to feel heavy. A workable governance model for a first workflow can be written on two pages: eight roles, six recurring decisions, and a simple rule about who signs off on what. The point isn’t ceremony. It’s making sure that when something does go wrong — and on a long enough timeline, something eventually will — there is a clear, fast, pre-agreed answer for who gets pulled in, rather than a scramble to figure out ownership while a customer or regulator is waiting.
Where Are You on the Curve?
Not every organization needs to leap straight to an adaptive, multi-workflow platform, and trying to skip stages is the most common root cause behind the failure modes above. The progression we see in practice runs through five stages: ad hoc answers, where individuals use general-purpose tools informally; connected content, where a few approved sources are indexed without policy enforcement at retrieval time; governed retrieval, where entitlements, provenance, and staleness checks are finally enforced; contextual workflows, where one or more use cases run on assembled context with real evaluation and escalation logic; and adaptive enterprise context, where a shared platform serves multiple workflows with continuous evaluation and feedback.
“Adaptive” describes how efficiently the platform can serve a new workflow — never license for the system to change policy or business definitions on its own.
That distinction matters more, not less, as a program matures. The blast radius of a bad semantic or policy change grows with every workflow that depends on the shared platform, so the governance model from the previous section doesn’t get to relax just because the technology got more capable.
A 90-Day Plan That Actually Respects Your Calendar
If you’re starting from fragmented systems — which is most organizations — the right first move is not a platform purchase. It’s a conservative, three-phase proof on a single workflow.
Weeks 1–3: Discover
Pick one bounded workflow with a real business owner and a measurable baseline. Inventory the two or three sources it actually needs. Resist the pull toward a workflow that’s easy to demo instead of one that’s easy to measure.
Weeks 4–8: Prove
Stand up governed retrieval and context assembly for that workflow only — nothing broader. Build an initial evaluation set. Run a real access-control audit using actual (not admin) accounts; this is the single most commonly skipped step under deadline pressure, and skipping it is how permission gaps get discovered in production instead of in testing.
Weeks 9–12: Operationalize
Add escalation and feedback loops. Formally define the operating-model roles above. Bring the business owner a clear go/no-go recommendation, with the rationale documented either way — a well-run pilot that surfaces a hard blocker and recommends against expansion is a success, not a failure.
What should deliberately wait: multi-workflow platform investment, a knowledge graph, and broad source expansion. Building the shared platform layers before a single workflow has proven the pattern is a reliable way to spend a large budget on infrastructure shaped wrong for your first real use case.
Ten Questions Before You Fund the Next Pilot
If you’re the one signing off on the budget, here are the questions worth asking before you approve a pilot, expand one, or greenlight scaling past it. We go deeper on each of these — with the specific red flag to listen for in the answer — in the full white paper, but the short list alone will change most pilot conversations:
- What is the one workflow, and who is its named business owner?
- What is the measured baseline this system is being compared against?
- Who owns each source, and is that ownership documented anywhere?
- Does retrieval enforce the requester’s actual entitlements — tested with real accounts, not admin ones?
- Can every material answer be traced to a specific, dated source?
- What does the system do when it isn’t confident?
- What’s the cost per resolved query, and how does that scale with usage?
- Is the architecture portable across models and vendors?
- What’s the change management plan for the people whose workflow this touches?
- What regulatory exposure does this workflow carry, and has legal or compliance actually reviewed it?
On that last point: this is no longer a hypothetical. Texas’s Responsible Artificial Intelligence Governance Act took effect January 1, 2026. Colorado’s automated decision-making framework continues to evolve through 2026 and 2027, with obligations attaching to AI systems used in consequential decisions across employment, lending, insurance, and healthcare. A context capability that can’t show which source justified which decision, and who was authorized to see it, isn’t just a quality gap anymore — under this regime, it’s a governance gap a regulator can ask about directly.
If your legal or compliance team hasn’t been in the room for at least one of your active AI pilots, that’s worth fixing this week, not at the next quarterly review. The cheapest time to involve them is before a workflow launches, when a scope change costs an afternoon of conversation. The most expensive time is after an incident, when the same conversation costs an investigation.
The Unglamorous Truth
Nothing in this post is exciting in the way a new model release is exciting. There’s no headline capability, no benchmark chart trending up and to the right. What there is, is a repeatable, testable path off the 95 percent list: pick one workflow, name its owner, build governed retrieval before you build a knowledge graph, enforce entitlements with real accounts before you scale to more users, and put governance in place before your second pilot instead of after your first one breaks.
None of that requires a large upfront platform investment. All of it is testable inside a single fiscal quarter. What it requires is the discipline to measure a pilot honestly, escalate what doesn’t work, and resist the very real pressure to declare success before an answer can be traced back to a source a compliance reviewer would actually trust.
It’s also worth saying plainly what this is not. It is not a call to slow down on AI, and it is not an argument for a year-long infrastructure project before anyone gets to touch a use case. It’s the opposite: the fastest reliable path to production is usually the narrow one, on one workflow, with governance built in from week one — not the broad one that looks faster on a slide and then stalls exactly where the MIT and Gartner data says most enterprise AI programs stall today.
Success in twelve months doesn’t look like a platform migration. It looks like one workflow where a leader can ask a governed question, get a grounded and traceable answer, and know that answer would survive an audit. That’s achievable today. It just isn’t achieved by a better model — it’s achieved by treating context as the operating layer it actually is.
Go deeper
This post is a teaser for our full white paper, “Enterprise Context: The Missing Operating Layer for Reliable AI.” It includes the complete vendor-neutral reference architecture and diagrams, the full nine-failure-mode table, an eleven-item hurdle heat map, a RACI model for context governance, the five-stage maturity model, the full 90-day roadmap, a two-page executive decision checklist, and a sourced appendix of every claim in both documents.
Download Whitepaper
Enterprise Context
The Missing Operating Layer for Reliable AI
Why 95% of enterprise AI pilots fail to move the needle — and the reference architecture, governance model, and 90-day roadmap to fix it..
