morpho360
Executive essay
Evidence-led executive essay

From Automation to Adoption

Why AI projects stall—and what it takes to turn technical capability into measurable behaviour change.

A careful reading of the MIT NANDA and RAND failure studies—and the leadership disciplines their findings actually support.
Evidence base

MIT NANDA · RAND Corporation

Published by

morpho360

Edition

July 2026 · 16-minute read

The real distance to value

The most dangerous moment in an AI project is not when the model fails.

It is when the model works.

A successful demonstration creates the appearance that the hard part is over. The system can summarize the document, classify the request, draft the response, predict the outcome, or recommend the next action. Executives see capability. The project team sees momentum. A vendor sees a path to deployment.

But a demonstration answers only one question: can the technology perform a task under selected conditions? It does not prove that the task is worth changing, that the system fits the real workflow, that the people involved will use it correctly, that exceptions can be governed, or that any resulting improvement will reach the income statement.

That distance—from technical performance to sustained business value—is where AI projects disappear.

Two widely cited studies help explain the pattern. RAND Corporation reports that, by some estimates, more than 80 percent of AI projects fail. MIT Project NANDA’s 2025 report produced the more dramatic headline: 95 percent of organizations were seeing zero return from enterprise generative-AI investment, while only 5 percent of task-specific tools reached successful implementation.

Those numbers are alarming. They are also routinely misused.

The useful lesson is not that AI fails 80 or 95 percent of the time. It is that organizations continue to treat implementation as a technology event when it is actually a redesign of decisions, workflows, responsibilities, and behaviour.

AI does not create value when it produces an answer. It creates value when the organization changes what it does with that answer—and can prove that the change mattered.

What the failure statistics really say

RAND: failure begins upstream of the model

In The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed (2024), RAND researchers interviewed 65 experienced data scientists and engineers from industry and academia. Participants had at least five years of AI or machine-learning experience. The study was exploratory: it asked practitioners to describe the failures they had observed and the root causes they considered most frequent or consequential.

RAND did not independently calculate a universal failure rate. Its report cites an outside estimate that more than 80 percent of AI projects fail, then investigates why projects flounder. That distinction matters. The strongest evidence in the report is not the percentage; it is the recurrence of the same failure mechanisms across experienced practitioners.

Five leading causes emerged:

  • business and technical stakeholders misunderstand or miscommunicate the problem;
  • the organization lacks suitable data;
  • teams pursue fashionable technology rather than a real user problem;
  • data and deployment infrastructure are inadequate; and
  • AI is applied to problems beyond its practical limits.

The most striking finding is managerial. Eighty-four percent of RAND’s industry interviewees cited one or more leadership-driven causes as a primary reason AI projects fail. Leadership teams ask technical teams to optimize the wrong metric, solve a problem with little business consequence, or deliver against expectations that the technology and data cannot support. Months of technically competent work can therefore produce an operationally irrelevant result.

RAND also draws attention to interaction and expectation failures. Projects do not fail only because an algorithm is inaccurate. They fail when people cannot fit the technology into their work, when leaders and builders understand the objective differently, or when the value expected from the system bears little relationship to what was designed.

The study has an important boundary: it focused on machine-learning projects and included LLMs, but excluded projects that only used pretrained LLMs through prompt engineering. Its findings are therefore highly relevant to custom or integrated AI systems, but should not be applied mechanically to every employee using a general-purpose chatbot.

MIT NANDA: high usage is not transformation

MIT Project NANDA’s preliminary 2025 report, The GenAI Divide: State of AI in Business 2025, examined more than 300 publicly disclosed AI initiatives, interviewed representatives from 52 organizations, and collected survey responses from 153 senior leaders.

Its central finding was a divide between widespread experimentation and measurable enterprise impact. More than 80 percent of organizations had explored or piloted tools such as ChatGPT and Copilot, and almost 40 percent reported deployment. Yet only 5 percent of task-specific enterprise GenAI tools were described as successfully implemented.

The “95 percent failure” headline needs precision. The report defines successful implementation as a task-specific tool that users or executives said produced a marked and sustained productivity or P&L impact. The finding does not mean that 95 percent of all AI models malfunction, that every pilot in the 95 percent produced no learning, or that general-purpose tools have no utility.

The authors also label the implementation figures directionally accurate rather than definitive. Sample sizes varied, success definitions differed between organizations, and outcomes were largely interview-based rather than drawn from audited financial reporting. The work is preliminary research, not a universal law of enterprise technology.

Even with those limitations, the pattern is useful. General-purpose LLMs spread rapidly because employees can try them immediately and adapt them to small tasks. Task-specific enterprise tools stall because they are brittle, require repeated context, fail at edge cases, and do not fit the daily workflow. MIT NANDA calls this the learning gap: many systems do not retain feedback, adapt to the organization’s context, or improve with use.

The report also identifies a shadow-AI economy. Employees are already using flexible consumer tools, frequently outside formal initiatives, because those tools solve immediate problems better than sanctioned systems. Adoption is not absent. It is happening around the organization rather than through it.

The studies converge more than the percentages suggest

RAND studies a broad set of AI and machine-learning projects through the experience of builders. MIT NANDA focuses on enterprise generative AI during a period of rapid LLM democratization. Their samples, methods, definitions, and periods differ. Their headline numbers should not be averaged or treated as competing measurements of one stable phenomenon.

Their causal stories nevertheless converge:

RAND MIT NANDA Shared implication
Wrong problem or wrong metric Investment follows visibility rather than workflow value Start from a consequential business problem.
Weak communication between domain and technical teams Tools do not match day-to-day operations Co-design with the people who know the work.
Insufficient or unsuitable data Systems lack context and learning Build evidence, feedback, and data discipline into the workflow.
Infrastructure and deployment gaps Pilots fail to integrate and scale Design the operating system, not just the model.
Misunderstanding AI’s limits Brittle tools fail on complex, high-stakes work Bound the use case and preserve human judgment.
Interaction and expectation failure High adoption but low transformation Measure behaviour and business consequence, not access or activity.

The studies are not telling leaders to avoid AI. They are telling leaders to stop confusing technological availability with organizational readiness.

LLM democratization changed the entry cost—not the work of transformation

Generative AI has removed several traditional barriers. A team no longer needs to train a model from scratch to test whether AI can draft, classify, extract, summarize, translate, or reason across text. An employee can move from idea to useful output in an afternoon. Prototypes that once required specialists and infrastructure now require a browser and a well-formed prompt.

This matters. It makes experimentation cheaper, increases the number of people capable of discovering use cases, and shifts useful innovation closer to the frontline.

It also creates a new failure pattern: local utility is mistaken for enterprise value.

An individual saving twenty minutes on a report is real value to that person. But that saving does not automatically become profit, capacity, faster service, or lower risk. The organization may still require the old process. A manager may recheck every output. Data may be copied manually between systems. The employee may use the freed time on other unmeasured work. The tool may create new review, privacy, or quality obligations.

The gap is not evidence that personal productivity is imaginary. It means value is being trapped between the task and the operating model.

Democratized LLMs therefore change where leaders should look for opportunity. Employees who have already developed effective practices can reveal valuable workflow edges. But informal success must be converted carefully into an organizational capability:

  • the problem must be named;
  • the workflow must be understood;
  • acceptable and unacceptable uses must be clear;
  • the right context and data must be available;
  • output quality and exceptions must be governed;
  • the behaviour change must be supported;
  • and the business consequence must be measured.

The cost of producing an answer has collapsed. The cost of integrating trustworthy behaviour has not.

The real unit of AI transformation is the workflow

Organizations buy models, licences, agents, platforms, and APIs. Value appears in workflows.

A workflow is where information enters, decisions are made, exceptions appear, responsibilities change hands, customers experience service, and financial consequences accumulate. It is also where AI meets organizational history: old controls, informal workarounds, political boundaries, missing data, overextended managers, and employees who remember the last transformation that never delivered.

This is why tool-first deployment underperforms. A tool can be excellent in isolation and still make the overall system worse.

Consider an AI assistant that drafts customer responses in seconds. If employees do not trust it, they rewrite every answer. If supervisors require two approvals, cycle time increases. If the assistant cannot see account history, it produces polished but irrelevant language. If performance is measured by drafts generated rather than resolution time or customer retention, the project can look active while the workflow deteriorates.

The correct design question is not “How accurate is the assistant?” It is:

What must change in this workflow—from information and authority to behaviour and measurement—for the organization to create value safely?

That question forces leaders to consider the full system:

  • the trigger for the work;
  • the person accountable for the outcome;
  • the decisions AI may support or perform;
  • the context the system needs;
  • the human review appropriate to consequence and uncertainty;
  • the handling of edge cases and failure;
  • the new behaviour expected from users;
  • the controls that protect customers and the organization;
  • and the signals that demonstrate value.

It also reveals when AI is not the right intervention. Sometimes the better answer is simpler rules, cleaner data, fewer approvals, better use of existing software, or clearer ownership.

Human adoption is not downstream from implementation

Many AI programmes treat adoption as the final stage. First the solution is selected and built; then communications, training, and change management are added to help people accept it.

By then, the most important human decisions have already been made without the humans who understand them.

Adoption should shape the choice from the beginning. Frontline employees know where the process description is false, where exceptions concentrate, which information arrives late, what customers will not tolerate, and which “inefficiency” is actually a safety mechanism. Their resistance may expose a weak solution rather than a weak attitude.

This does not mean every employee has veto power over change. It means operational knowledge and behavioural reality are evidence.

People adopt a changed operating contract, not a tool

When AI enters a workflow, employees are rarely being asked merely to click a new button. They may be asked to:

  • trust a recommendation they cannot fully inspect;
  • remain accountable for an output they did not create;
  • detect subtle errors while working faster;
  • surrender discretion or status built around expertise;
  • document context that previously lived in their heads;
  • teach the system through feedback;
  • or accept that performance is now measured differently.

Those are changes to the operating contract between the employee and the organization. A training video cannot resolve them.

Trust must be calibrated, not maximized

Blind trust is dangerous; total distrust destroys value. Users need a practical understanding of where the system is strong, where it is uncertain, what evidence accompanies an output, when human review is mandatory, and how to challenge or escalate.

The goal is calibrated reliance: the right level of trust for the task and consequence.

Feedback requires a response

Organizations frequently ask employees to report bad outputs, then provide no evidence that the system or workflow changed. Feedback becomes unpaid quality assurance and trust erodes.

A credible learning loop names who reviews feedback, how quickly, what qualifies as a material issue, how changes are tested, and how users learn what happened. This is the organizational counterpart to MIT NANDA’s technical learning gap. A system may be capable of learning, but the organization must also know how to learn around it.

Adoption must be observed in behaviour

Licence activation, logins, training completion, and prompts submitted are activity measures. They do not prove that a changed workflow is being used correctly or producing value.

Better signals include:

  • the share of eligible cases handled through the new workflow;
  • the rate of correct use without unnecessary rework;
  • override and escalation patterns;
  • error severity and recovery time;
  • time from decision to adopted routine;
  • user trust calibrated against actual system performance;
  • and the business outcome connected to that behaviour.

If leaders cannot describe the behaviour that must change, the project is not ready for deployment.

What leadership must do differently

There may be four useful moves in a particular organization. There may be six, eight, or three. The evidence does not justify a universal numbered formula.

What it does support is a connected set of leadership disciplines. They form a chain; weakness in one can invalidate the rest.

Choose a problem worthy of sustained attention

Name the business consequence before the technology. Identify who experiences the problem, how often, and why it matters. Compare AI with simpler options, including changing the process or doing nothing. RAND recommends choosing enduring problems rather than chasing the opportunity of the moment.

Put domain experts, users, and technical people in the same decision loop

Miscommunication is not solved by a handoff document. The people who understand the business context, the workflow, the data, and the technology need recurring contact throughout discovery, design, validation, and operation.

Design around the real workflow

Map the current process, including exceptions and unofficial work. Define how AI changes decisions, roles, controls, and customer experience. Start narrowly where value and learning can be observed.

Make uncertainty governable

List assumptions, risks, limits, and evidence gaps. Use a proportional, reversible next step with explicit success, escalation, and stop conditions. A pilot should answer a decision—not merely demonstrate activity.

Build technical and organizational learning loops

The system needs context, feedback, monitoring, and a path to improvement. The organization needs the same: people who review signals, authority to change the workflow, and a cadence that turns experience into better decisions.

Engineer adoption and value together

Define the behaviour change, involve affected people, assess change capacity, provide the right enablement, and connect usage to a business baseline. Do not declare victory at launch. Measure whether the new behaviour sticks and whether value survives after the project team leaves.

These disciplines are not sequential departments. They must inform one another. A human adoption barrier may change the technical design. A data constraint may narrow the business case. A workflow observation may prove that the original problem was wrong. Measurement may show that a successful model creates no economic value.

That is not project failure. That is disciplined learning before scale.

Why the morpho360 Method™ is built for this problem

The morpho360 Method™ begins from the premise that AI transformation is a decision-and-adoption problem before it is a technology problem.

Its six steps address the failure chain identified across the RAND and MIT NANDA research.

1. Clarify the real problem

The method refuses to begin with a preferred tool. It identifies the business consequence, stakeholders, workflow reality, evidence, constraints, and decision owner. This directly counters RAND’s most common leadership-driven failure: asking technical teams to solve the wrong problem.

2. Translate complexity

Executives, operators, and technical specialists rarely use the same language. morpho360 converts technical, operational, data, and organizational complexity into a shared mental model so disagreements become visible before they become expensive.

3. Isolate decisions and map trade-offs

The work names the actual decision, the realistic options, and the risks across strategy, finance, data, technology, workflow, governance, vendors, and adoption. Assumptions become evidence tasks. “Pilot AI” is replaced by a defensible choice: investigate, buy, build, partner, validate, defer, stop, or do nothing.

4. Build a proportional roadmap

The roadmap is sequenced by value, risk, feasibility, dependencies, and the organization’s absorption capacity. It protects capital by funding the next evidence-producing step rather than treating a successful demo as permission to scale.

5. Enable ground-level adoption

Human adoption runs in parallel with the decision track. Stakeholder dynamics, resistance, capability, trust, behaviour change, and operational ownership are allowed to reshape the solution itself. Adoption is not a communications layer applied after design.

6. Establish continuous measurement and feedback loops

Success measures connect system performance to use, behaviour, and business consequence. Feedback loops capture what happens after deployment, improve the workflow, and reopen the problem when evidence changes. This responds directly to the learning gap identified by MIT NANDA.

Two tracks, one outcome

The method operates through a decision track and a human track at the same time.

The decision track asks: Is this the right problem, option, risk posture, and investment sequence?

The human track asks: Will the people who must live with the change understand it, trust it appropriately, use it correctly, and sustain it under real operating pressure?

Neither track can approve the transformation alone.

This is why morpho360 does not define success as a model in production, a licence activated, or a pilot completed. Success is a defensible decision converted into trusted behaviour, with evidence that the value is real and a mechanism for learning when it is not.

The objective is not to become one of the lucky five or twenty percent. It is to stop relying on luck.

Conclusion: deployment is the midpoint

AI project failure is often presented as proof that the technology is immature or that organizations need better models. The two studies tell a more demanding story.

RAND’s experienced builders point upstream to leadership, problem selection, data, infrastructure, communication, and technical limits. MIT NANDA observes a world in which employees adopt LLMs rapidly while enterprise transformation stalls at workflow fit, context, learning, ownership, and measurable impact.

The common failure is not an inability to produce AI output. It is an inability to reorganize work around that output responsibly.

Leaders should therefore stop asking only whether an AI system can perform. They should ask:

  • Is this a consequential problem?
  • Does the solution fit the real workflow?
  • What evidence would justify the next commitment?
  • Who remains accountable when the system is wrong?
  • What behaviour must change for value to appear?
  • Can the organization absorb and sustain that change?
  • How will the system and the organization learn?
  • What would cause us to stop?

The answers determine whether automation becomes adoption—and whether adoption becomes value.

Research sources

  1. James Ryseff, Brandon De Bruhl, and Sydne J. Newberry, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed: Avoiding the Anti-Patterns of AI, RAND Corporation, 2024, DOI: 10.7249/RRA2680-1.
  2. Aditya Challapally, Chris Pease, Ramesh Raskar, and Pradyumna Chari, The GenAI Divide: State of AI in Business 2025, MIT Project NANDA, preliminary findings, July 2025.

A note on interpretation

The studies use different populations, methods, scopes, and definitions of failure. RAND’s report cites—but does not independently calculate—the “more than 80 percent” estimate. MIT NANDA’s 95 percent concerns task-specific enterprise GenAI tools that had not achieved marked, sustained productivity or P&L impact under a preliminary, largely interview-based methodology. Neither statistic should be presented as a timeless failure probability for every AI initiative.

About morpho360

morpho360 helps leadership teams make and execute complex technology-enabled transformation decisions with greater clarity, commercial discipline, and adoption confidence.

Suggested call to action: Before scaling another AI pilot, examine whether the decision, workflow, evidence, governance, adoption, and measurement system are ready to carry it.

morpho360
Decision clarity for technology-enabled transformation.