April 29, 20245 min read

    From Idea to Impact: Why Enterprise Product Development Is Still Slow When Building Is Cheap

    By MASSIVUE Team

    From Idea to Impact: Why Enterprise Product Development Is Still Slow When Building Is Cheap
    Product DevelopmentProduct ManagementEnterprise TransformationAI-Assisted DevelopmentOperating ModelDecision RightsStage Gates
    Contents
    1. The short answer
    2. Key takeaways
    3. Has AI actually made building faster?
    4. Where idea-to-impact time actually goes
    5. The evidence queue
    6. The authority queue
    7. The integration queue
    8. The adoption queue
    9. How to redesign the gates
    10. Where MASSIVUE fits
    11. Frequently asked questions
    12. Related MASSIVUE resources
    13. Limitations of this article
    14. Sources

    Enterprises bought the tools that made building cheap and did not get faster. This article sets out where idea-to-impact time actually goes, and which gates to change.

    Published by MASSIVUE, an enterprise transformation and AI capability firm. Last reviewed: August 2026.


    The short answer

    Enterprise product development is still slow because building was never the long pole. The path from an idea to a measurable business result runs through four queues that AI has not touched: assembling evidence that anyone wants the thing, obtaining the authority to fund it and to stop it, integrating with the systems the business actually runs on, and getting people to work differently once it ships. AI compressed the build leg and added a new cost, the effort of reviewing output that arrives faster than the organisation can validate it. Speed comes from redesigning the gates around those four queues, not from adding build capacity.


    Key takeaways

    • The most rigorous controlled trial of AI-assisted development, run by METR, has not yet demonstrated a speedup for experienced engineers on mature codebases, and its authors now say their own signal is unreliable.
    • Google's DORA research finds AI adoption relates positively to delivery throughput and negatively to delivery stability. More output is not the same as more impact.
    • The widely repeated claim that 80 to 90 percent of new products fail is not supported by the research literature. Peer-reviewed work puts the figure at 40 percent or less.
    • Peer-reviewed field research inside a large distributed company names annual budgeting cycles and unclear decision-making authority as direct barriers to getting new ideas built.
    • The practical intervention is to change what a gate asks for. Gates that ask for documents reward writing. Gates that ask for evidence reward learning.

    Has AI actually made building faster?

    Less clearly than most boards have been told.

    The most rigorous attempt to measure it is a randomised controlled trial by METR, published in July 2025. Sixteen experienced open-source developers worked through 246 real issues in their own large repositories, with each task randomly assigned to allow or forbid AI assistance. The developers expected to be sped up by 24 percent. Afterwards they believed they had been sped up by 20 percent. They were measurably slowed down by 19 percent.

    METR has since labelled that result historical and run a continuation with 57 developers, 143 repositories and more than 800 tasks, reported in February 2026. The point estimates were still negative: a speedup of minus 18 percent for the original cohort, with a confidence interval running from minus 38 to plus 9 percent, and minus 4 percent for newly recruited developers. METR's own conclusion is the one worth carrying forward. They state that the data gives an unreliable signal of the current productivity effect of AI tools, because between 30 and 50 percent of developers declined to submit tasks they did not want to attempt without AI. The experiment has become hard to run precisely because the tools have become hard to give up.

    So the correct reading is not that AI makes developers slower. It is that a clear, measured, end-to-end speedup for experienced engineers on mature systems has not been established, while the belief that it exists is very well established. That gap between perceived and measured gain is the single most expensive assumption in enterprise product planning right now, because roadmaps, headcount and business cases have all been rebuilt on top of it.

    Aggregate industry data points the same way. Google's DORA programme surveyed close to 5,000 technology professionals for its 2025 report. Ninety percent used AI at work and more than 80 percent believed it had increased their productivity, yet 30 percent reported little or no trust in the code it produced. DORA found AI adoption positively related to software delivery throughput and negatively related to delivery stability. Its summary of the mechanism is that AI acts as an amplifier of whatever the organisation already is.

    A 2026 CloudBees survey of more than 200 enterprise technology leaders, a vendor study rather than peer-reviewed research, describes the downstream shape of that. AI now generates or assists roughly 61 percent of the average enterprise codebase, 81 percent of those leaders saw an increase in production issues tied to AI-generated code, and organisations could attribute only about a third of their AI-related spend to specific business outcomes.

    More code, arriving faster, that fewer people trust, costing money nobody can trace to a result. That is not a build problem. It is a validation and decision problem.


    Where idea-to-impact time actually goes

    An enterprise idea has to clear four queues before it becomes impact. Each has its own owner, its own clock and its own reason for stalling. None of them gets shorter because code is generated faster.

    Two horizontal bars comparing an enterprise product cycle before and after AI-assisted build. In the second bar the build segment is much shorter, the evidence, authority, integration and adoption segments are unchanged, and a new rework and review segment appears at the end.
    The build leg was compressed. The other four were not, and a fifth appeared. MASSIVUE analysis, informed by DORA 2025 and METR 2025 to 2026.
    QueueThe question it answersWho actually controls itTypical failure
    EvidenceDoes anyone outside the building want this?Product and researchInternal opinion is presented as validation
    AuthorityWho can fund it, and who can stop it?Finance and the sponsorDecision rights are undefined, so nobody can say stop
    IntegrationCan it reach the systems the business runs on?Platform and architectureThe demo never had to touch real data
    AdoptionWill people work differently because it exists?The receiving business unitLaunch is treated as the finish line

    The four-queue framing is MASSIVUE editorial synthesis. The evidence for each queue is external and is cited at the point of use.


    The evidence queue

    Making a prototype cheap does not make proof cheap. A working demo answers whether something can be built. It says nothing about whether anyone will change their behaviour to use it, and gathering that second answer still costs the same calendar time it always did, because it depends on real users being available.

    This queue is also where the industry's favourite statistic does damage. The claim that 80 or 90 percent of new products fail circulates constantly in innovation decks. It is not supported by the literature. In a 2013 perspective piece in the Journal of Product Innovation Management, Castellion and Markham traced the figure and concluded that studies since 1977 put the new product failure rate at 40 percent or less. They attributed the survival of the higher number to argument from popularity and to the self-interest of parties who benefit from quoting it.

    The distinction matters operationally. A 90 percent failure rate justifies a spray-and-pray portfolio in which nothing is examined closely. A 40 percent rate justifies the opposite: fewer bets, each with real evidence behind it, and a genuine willingness to stop the ones that are not working. Confusion between idea failure rates and product failure rates is part of why the myth persists, and the two are not interchangeable.

    The same care applies to the newer figure now being quoted in AI business cases, that 95 percent of enterprise generative AI pilots produce no measurable return. It comes from The GenAI Divide: State of AI in Business 2025, a July 2025 report from MIT Media Lab's Project NANDA. That report is explicitly preliminary and not peer reviewed, and rests on 52 structured interviews, 153 survey responses and a review of publicly disclosed initiatives, measured against a roughly six-month definition of success. Treated as a directional observation that most pilots never captured a pre-deployment baseline, it is useful. Treated as a measured failure rate, it is not, and it is currently being used to justify decisions it cannot support.

    What shortens this queue: deciding in advance what result would count as evidence, and what result would trigger a stop. An evidence threshold written before the test is worth more than any amount of analysis afterwards.


    The authority queue

    This is the queue enterprises least like to look at, because the delay is structural rather than technical.

    Field research published at the 2021 IEEE and ACM International Conference on Global Software Engineering studied five internal startups inside a single large, globally distributed company. Sporsem and colleagues identified six barriers to getting new ideas built. Three of them are pure authority problems: yearly budgeting and planning cycles, unclear decision-making authority, and missing or unclear executive sponsorship. The others were late involvement of developers, no digital infrastructure for experimentation, and limited access to external data.

    Yearly budgeting is the one most enterprises accept as immovable. Its effect is precise: an idea that becomes credible in month three of the financial year cannot be funded until month twelve, so it either waits, or it gets smuggled into an existing budget line where nobody will review it honestly. Both outcomes are slow, and the second is worse, because unreviewed work is exactly the work that never gets stopped.

    Unclear decision authority produces the mirror problem. If nobody holds an explicit right to stop an initiative, initiatives do not stop. They shrink, get renamed, and continue consuming capacity that the next idea needs. The absence of a stop decision is not neutrality; it is a funding decision made by default.

    What shortens this queue: naming one person who can both fund and stop each initiative, and moving at least part of the portfolio to a quarterly funding cadence so that evidence produced in month three can be acted on in month four.


    The integration queue

    A prototype that runs on sample data has not met the constraint. The constraint is the customer record, the ledger, the policy engine, the entitlements model, and whatever access controls sit around them.

    DORA's 2026 report on the return on AI-assisted software development is direct about where value actually comes from. Its conclusion is that the greatest returns on AI investment come not from the tools themselves but from the underlying organisational system: the quality of the internal platform, the clarity of workflows and the alignment of teams. The same research introduces a J-curve of value realisation, in which organisations reliably experience a productivity dip before any gain, driven by workflow adaptation, the verification cost of AI-generated code, and downstream process adjustment. DORA calls that dip the tuition cost of transformation.

    The report does model an illustrative case, a first-year return of 39 percent with roughly an eight month payback for a 500-person engineering organisation, but its authors explicitly describe these as high-uncertainty estimates meant to start a conversation rather than a formula. That caveat should travel with the number wherever it is quoted.

    For a product manager, the practical consequence is that integration readiness is a property of the platform, not of the project, and it cannot be created inside the initiative that needs it. If the internal platform cannot expose data safely to a new product, every product will pay the same tax, and no individual business case will ever be allowed to fund the fix.

    What shortens this queue: requiring the first meaningful test to run against a real system of record, however small the slice, and treating platform capability as a funded product with its own owner rather than as overhead.


    The adoption queue

    The last queue is the one where most measured value is lost, and it is measured in how much change a business unit can absorb, not in how much software has been shipped.

    McKinsey's State of AI survey, published in November 2025 with close to 2,000 respondents, illustrates the gap. Eighty-eight percent said their organisation regularly uses AI in at least one business function. Nearly two-thirds said they had not yet begun scaling AI across the enterprise. Thirty-nine percent attributed any EBIT impact at all to AI, and most of those put it below 5 percent of EBIT.

    That is a picture of very high deployment and very low absorption. Shipping is not the constraint. The constraint is the receiving organisation's capacity to change process, retrain people, rewrite the job description and retire the old way of working. That capacity is finite, it belongs to the business unit rather than to the product team, and it is almost never scheduled.

    What shortens this queue: naming the operating metric that must move before work starts, and confirming that the receiving business unit has capacity in the quarter the product lands. If it does not, the launch date is fiction regardless of engineering progress.


    How to redesign the gates

    Most enterprise stage gates were designed when building was the expensive step, so they concentrate scrutiny before the build and go quiet afterwards. When building becomes cheap, that design inverts: the gates now guard the wrong door. The fix is not to remove gates. It is to change what each one asks for.

    GateReplace this askWith this askStop rule
    1. ProblemA business case built on forecast benefitNamed users, the current cost of the problem to them, and the metric that would moveNo named user will describe the problem in their own words
    2. EvidenceA completed research documentA pre-declared threshold and the result of testing against itThe threshold was missed, or was rewritten after the result came in
    3. IntegrationAn architecture review of a proposed designA thin slice running against a real system of recordThe slice cannot obtain data access within the quarter
    4. AbsorptionA go-live dateWritten confirmation of change capacity from the receiving unit, plus the owner of the operating metricNo unit will accept the metric
    5. ImpactA launch announcementThe metric measured against the pre-deployment baselineThe baseline was never captured

    Three design rules make this hold in practice, and they are MASSIVUE practitioner guidance rather than research findings.

    First, every gate needs a stop rule written before the gate is reached, because a stop rule invented afterwards is a negotiation. Second, the person who can approve funding at a gate must also be the person who can stop it, or the gate only functions in one direction. Third, capture the baseline at gate one. Most of the pilots later described as having produced no measurable return were never measured at the start, which makes the failure a measurement failure rather than a product failure.

    None of this requires a new methodology. It requires the existing gates to ask for evidence rather than documents.


    Where MASSIVUE fits

    MASSIVUE is an enterprise transformation and AI capability firm. Its relevance to this problem is specific rather than general.

    The four-queue model and the gate design above are MASSIVUE editorial synthesis, built on the external research cited throughout. Everything attributed to a third party is sourced below and can be checked against the publisher's own material.

    Where an organisation concludes that the constraint is authority rather than capability, that sits with Enterprise Transformation, which covers operating model, decision rights and funding cadence. Where the constraint is the platform and the AI delivery system underneath it, that sits with AI Transformation, and the underlying operating model framework is Protum. Where the constraint is the receiving organisation's ability to absorb change, that sits with AI Workforce Transformation.

    On the Academy side, the queue this article treats as most neglected, authority, is the explicit subject of the Pragmatic Product Ownership micro-credential, which is written for the managers and sponsors who set the authority boundary that Product Owners have to work inside. Teams whose problem is proving impact after launch rather than deciding before it will find the relevant discipline in Total Economic Impact of AI, which covers benefits, costs, risk adjustment and building a model a finance function will accept.

    Where does your idea-to-impact time actually go?

    If your build capacity has increased and your delivered impact has not, the constraint is in one of the other four queues. Working out which one is a short diagnostic exercise, not a programme.

    Talk to MASSIVUE about Enterprise Transformation, or start with Pragmatic Product Ownership if the authority boundary is the part you recognise.


    Frequently asked questions

    Why is enterprise product development still slow when AI has made building faster?

    Because building was one leg of a five-leg journey. Evidence gathering, funding and stop authority, integration with systems of record, and organisational adoption are unchanged by faster code generation, and a new cost has been added: reviewing and correcting output that arrives faster than the organisation can validate it. DORA's 2025 research found AI adoption relates positively to delivery throughput and negatively to delivery stability.

    Does AI actually make developers faster?

    The evidence is weaker than the belief. METR's randomised controlled trial found experienced open-source developers were 19 percent slower with AI access in 2025 while believing they were 20 percent faster. Its February 2026 continuation still produced negative point estimates, but METR states the signal is unreliable because 30 to 50 percent of developers refused to attempt tasks without AI. No clear measured end-to-end speedup for experienced engineers on mature codebases has been established.

    What percentage of new products actually fail?

    Around 40 percent or less, not the 80 to 90 percent commonly quoted. Castellion and Markham, writing in the Journal of Product Innovation Management in 2013, traced the higher figure to argument from popularity and self-interest, and reported that empirical studies since 1977 put the rate at 40 percent or below. Idea failure rates and product failure rates are different measures and are often conflated.

    Is it true that 95 percent of enterprise AI pilots deliver no return?

    That figure comes from The GenAI Divide: State of AI in Business 2025, a July 2025 report from MIT Media Lab's Project NANDA. It is explicitly preliminary and not peer reviewed, and rests on 52 structured interviews, 153 survey responses and a review of publicly disclosed initiatives against a roughly six-month definition of success. It is a reasonable directional signal that most pilots never captured a pre-deployment baseline. It is not a measured failure rate, and it should not be used to justify funding decisions on its own.

    What are the real bottlenecks in enterprise product development?

    Peer-reviewed field research on internal startups in a large distributed company identified six: late involvement of software developers, missing or unclear executive sponsorship, yearly budgeting and planning cycles, unclear decision-making authority, lack of digital infrastructure for experimentation, and limited access to external data. Three of those six are authority and funding problems rather than delivery problems.

    How should product development stage gates change now that building is cheap?

    Move the scrutiny from before the build to around evidence and absorption. Ask each gate for evidence rather than documents: named users at the problem gate, a pre-declared evidence threshold at the evidence gate, a thin slice running against a real system of record at the integration gate, written change capacity from the receiving business unit at the absorption gate, and the metric measured against a baseline at the impact gate. Write the stop rule before the gate is reached.

    Why do organisations report high AI adoption but low financial impact?

    Deployment and absorption are different things. McKinsey's November 2025 State of AI survey found 88 percent of organisations use AI in at least one function, nearly two-thirds have not begun scaling it enterprise-wide, and only 39 percent attribute any EBIT impact to it, with most of those reporting under 5 percent. The limiting factor is the receiving organisation's capacity to change how work is done.



    Limitations of this article

    The four-queue model is a framing device, not a measured decomposition. The proportions in the figure are directional and illustrative, not data. METR's trials measure experienced open-source developers working on large mature repositories and do not generalise to all software work. The CloudBees figures come from a vendor survey of just over 200 leaders. The internal startup research studied five initiatives inside one company. The DORA return-on-investment model is described by its own authors as high-uncertainty. Where a source is preliminary, small-sample or narrowly scoped, that is stated at the point of use above.


    Sources

    Each figure cited above was checked against the publisher's own material at the last review of this article.

    • METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
    • METR, We are Changing our Developer Productivity Experiment Design, February 2026. https://metr.org/blog/2026-02-24-uplift-update/
    • DORA and Google Cloud, State of AI-assisted Software Development, 2025. https://dora.dev/dora-report-2025/
    • DORA and Google Cloud, The ROI of AI-assisted Software Development, 2026. https://dora.dev/ai/roi/report/
    • CloudBees, 2026 State of Code Abundance Report, May 2026. https://www.cloudbees.com/blog/2026-state-of-code-abundance-report
    • Castellion, G. and Markham, S. K., Perspective: New Product Failure Rates: Influence of Argumentum ad Populum and Self-Interest, Journal of Product Innovation Management, 30(5), 2013, pages 976 to 979. https://doi.org/10.1111/j.1540-5885.2012.01009.x
    • Sporsem, T., Tkalich, A., Moe, N. B. and Mikalsen, M., Understanding Barriers to Internal Startups in Large Organizations: Evidence from a Globally Distributed Company, 16th ACM and IEEE International Conference on Global Software Engineering, 2021. https://arxiv.org/abs/2103.09707
    • McKinsey, The State of AI, November 2025. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
    • Challapally, A., Pease, C., Raskar, R. and Chari, P., The GenAI Divide: State of AI in Business 2025, MIT Media Lab Project NANDA, July 2025, circulated as a preliminary report. https://nanda.media.mit.edu/

    Share this article

    Help others discover this insight