Enterprise teams are adopting AI faster than they are redesigning around it. This guide sets out which established team frameworks survive AI conditions, what the 2026 evidence actually shows about accountability, and a working operating model for assigning ownership of human-agent output.
Executive summary
Atlassian's State of Teams 2026, a double-blind study of 12,035 knowledge workers and 173 Fortune 1000 executives fielded in January and February 2026, found that 85% of knowledge workers use AI at work, but only 29% have embedded it in their flows of work. In the same study, 89% of executives said AI increases speed, while only 6% were sure they had clear examples of organisation-wide AI ROI.
That gap is not a tooling problem. It is an ownership problem. Deloitte's April 2026 survey of 3,235 IT and business leaders across 24 countries found that only 21% of organisations have a mature governance model for agentic AI. Roughly 80% lack mature governance capabilities, which Deloitte describes as including clear boundaries defining which decisions agents may make independently, real-time monitoring of agent behaviour, and audit trails capturing the full chain of agent actions.
The classical team frameworks do not resolve this, because almost all of them assume every team member is human, carries intent, and can be held responsible. The first two assumptions are now false. The third is the one that matters.
Short answer
How do enterprise leaders build high-performing teams in the age of AI?
By reassigning accountability before adding capability. In an AI-era team, every output requires a named human owner, explicit decision rights defining what an agent may do without approval, and a defined escalation path when those limits are reached. Organisations that add AI without redesigning ownership become faster and less reliable at the same time.
Key takeaways
- Speed is no longer a differentiator. When every team has access to the same models, output velocity converges. What separates teams is whether anyone owns the output.
- Accountability does not transfer to an agent. BCG Henderson Institute research with more than 1,200 managers found that framing AI as an "employee" rather than a tool led managers to identify 18% fewer errors.
- Most enterprises have an AI strategy and no AI operating model. Only 21% report mature agentic AI governance.
- Of seven established team frameworks assessed here, two hold, two partly hold and three break. Frameworks describing human relationships survive. Frameworks assuming bounded membership or mutual accountability do not.
- The metrics most organisations track become misleading first. Velocity and utilisation rise whether or not the work is correct.
Contents
- What makes a team high-performing in the age of AI?
- Which classical frameworks still hold, and which break
- The accountability gap: what the 2026 evidence shows
- The Protum™ 3-2-1 model: who owns agent output
- Six capabilities of an AI-era high-performing team
- What to measure when agents are in the workflow
- A 90-day sequence for enterprise leaders
- Executive readiness scorecard
- Frequently asked questions
What makes a team high-performing in the age of AI?
A high-performing team in the age of AI is one where every output has a named human owner, regardless of whether a person or an agent produced it. The differentiator is no longer speed or collaboration quality, since AI raises both across the board, but clarity of accountability when the work is wrong.
For most of the past three decades, team performance was substantially a function of throughput: what a group could produce relative to its size and cost. That framing survived because output capacity was genuinely scarce and genuinely variable between teams.
Generative and agentic AI compresses that variance. When comparable models are available to every team in an organisation, and increasingly to every competitor, raw production capability stops being a meaningful differentiator. Atlassian's 2026 research found that 55% of executives believe AI is widening performance and opportunity gaps between teams. The tool is evenly distributed. The operating model is not.
Microsoft's Work Trend Index 2026, covering 20,000 knowledge workers across ten markets, found that organisational factors accounted for more than twice the reported AI impact of individual factors. This is among the most consequential structural findings available on the subject: the constraint on AI-era team performance sits above the individual, in how work is organised, owned and governed.
There is a corollary that, in our experience, is consistently underestimated. If AI raises output for everyone, then the cost of wrong output rises too, because there is more of it, produced faster, with less human contact per unit. A team that has doubled its throughput without changing who reviews and owns that throughput has not become high-performing. It has become high-volume with unchanged, and therefore proportionally weaker, quality control.
Which classical frameworks still hold, and which break
Frameworks describing how humans relate, such as psychological safety, shared purpose and trust, still hold and matter more. Frameworks describing how work is allocated and owned, including bounded membership, mutual accountability and team lifecycle stages, break, because they assume every team member can be held responsible for a decision.
The team-effectiveness literature is unusually mature. It is also unusually untested against non-human participants.
| Framework | Origin | Status | Why |
| Psychological safety | Edmondson, 1999, Administrative Science Quarterly 44(2), 350–383 | Holds, and matters more | Disclosing that an agent produced a flawed output carries a specific risk: appearing to have failed at supervision |
| Salas "Big Five" of teamwork | Salas, Sims & Burke, 2005, Small Group Research 36(5), 555–599 | Holds. Most transferable | Mutual performance monitoring and backup behaviour describe human-over-agent supervision almost exactly |
| Project Aristotle | Google re:Work, archived guide originally published 2015 | Partly holds | "Dependability" was defined for humans. Agents are consistent but not accountable, which is a different property |
| Five Dysfunctions of a Team | Lencioni, 2002 | Partly holds | The chain terminates in peer accountability. An agent cannot be a peer |
| Tuckman's stages | Tuckman, 1965, Psychological Bulletin 63(6), 384–399 | Breaks | Agents do not socialise into norms. Membership changes continuously, so re-forming never ends |
| Katzenbach & Smith | 1993, The Wisdom of Teams | Breaks | Mutual accountability is definitionally impossible where one party cannot be held to account |
| Hackman's conditions | Hackman, 2002, Leading Teams | Breaks | Bounded, stable membership is precisely what agentic workflows dissolve |
Attribution notes
These errors appear routinely in published commentary on this topic, including on pages that currently rank well. Getting them right is inexpensive and it is verifiable:
- "Adjourning" is not in Tuckman's 1965 paper. It was added by Tuckman and Jensen in "Stages of Small-Group Development Revisited," Group & Organization Studies, 1977, 2(4), 419–427, as a fifth stage.
- Hackman specified five conditions, not six: a stable team, a clear and engaging direction, an enabling team structure, a supportive organisational context, and available competent coaching. The widely marketed "six team conditions" framework is a later commercial derivative, not what the 2002 book says.
- Project Aristotle's five dynamics have a rank order: psychological safety, dependability, structure and clarity, meaning, impact. Secondary sources frequently scramble it. Google now preserves the guide as an archived record.
- The T7 model is consistently attributed to Lombardo and Eichinger across secondary sources, but no primary publication is locatable. It should be described as "widely attributed to" rather than cited.
The practical conclusion is not that the classical literature is obsolete. It is that the literature splits cleanly along one fault line: frameworks about how humans relate to each other survive; frameworks about how work is bounded and owned do not. That split is precisely why human-AI collaboration is an operating-model problem before it is a behavioural one.
The accountability gap: what the 2026 evidence shows
Enterprises are deploying AI agents substantially faster than they are defining who is answerable for agent output. Deloitte's April 2026 study of 3,235 leaders across 24 countries found that only 21% have a mature governance model for agentic AI, and roughly 80% lack mature governance capabilities altogether.
Agentic AI governance is the layer of control specific to systems that act rather than only advise. In Deloitte's framing it has three components: boundaries defining which decisions an agent may take independently, real-time monitoring of agent behaviour, and audit trails covering the full chain of agent actions. An organisation can have a mature AI policy and still have none of the three.
| Source | Date | Sample | Finding |
| BCG Henderson Institute | 6 May 2026 | >1,200 managers | When AI was framed as an "employee" rather than a tool, managers identified 18% fewer errors; individual accountability for errors fell 9 percentage points; accountability attributed to the AI rose 8 percentage points |
| Deloitte, AI agents are scaling faster than their guardrails | 24 Apr 2026 | 3,235 IT and business leaders, 24 countries | 21% have a mature agentic AI governance model; ~80% lack mature governance capabilities including decision boundaries, real-time monitoring and audit trails |
| Atlassian, State of Teams 2026 | 27 Apr 2026 | 12,035 knowledge workers + 173 Fortune 1000 executives, double-blind | 85% use AI at work, 29% have embedded it in flows of work; 89% of executives cite speed gains, 6% are sure they have clear examples of organisation-wide AI ROI; 55% say AI is widening gaps between teams |
| Microsoft, Work Trend Index 2026 | 5 May 2026 | 20,000 knowledge workers, 10 markets | Organisational factors account for more than 2× the reported AI impact of individual factors |
| Deloitte, Building high-performing teams | 14 Jan 2026 | 1,394 US professionals | High-performing teams significantly more likely to use AI tools, 78% versus 54%; 2.5× as likely to say their team can quickly change direction |
The finding that matters most
The BCG result deserves particular weight, because it runs directly against the prevailing narrative.
A great deal of current commentary encourages organisations to treat AI agents as team members: onboarding them, naming them, giving them roles. BCG's research with more than 1,200 managers found that this framing carries a measurable cost. When AI was framed as an employee rather than a tool, managers identified 18% fewer errors, individual accountability for those errors dropped nine percentage points, and accountability attributed to the AI itself rose eight points.
In other words, the language used to describe an agent changes whether a human still feels answerable for its output. Anthropomorphic framing does not just fail to help. It appears to actively erode the ownership that AI-era performance depends on.
This has a direct operating implication: the boundary between "the agent did it" and "I am answerable for it" has to be maintained deliberately, in language and in structure, because it does not hold on its own.
A note on the evidence base
Three caveats are worth stating plainly, because they are routinely omitted:
- Deloitte's high-performing-teams study is US-only and relies on self-identified membership of a high-performing team, which introduces self-report bias.
- Atlassian's widely-quoted "fragmentation tax" figure is a vendor-modelled estimate, not a measured value.
- McKinsey's The Agentic Organization (26 September 2025) is the most-cited framework document in this space and contains no primary research, no survey and no disclosed sample size. Its assertion that a human team of two to five people can supervise 50 to 100 specialised agents is drawn from consulting experience, not data. It is a useful hypothesis. It should not be repeated as a finding.
The Protum™ 3-2-1 model: who owns agent output
Massivue's Protum™ 3-2-1 model assigns ownership of human-agent work across three accountable human roles: a Value Lead, an AI Orchestration Lead, and an AI Quality Engineer. Its purpose is to prevent the accountability diffusion that occurs when agents enter a team whose ownership structure was designed for humans only.
- Value Lead: owns the outcome. Answerable for whether the work achieved what it was meant to achieve.
- AI Orchestration Lead: owns deployment. Answerable for how agents are configured and applied against that outcome.
- AI Quality Engineer: owns fitness to ship. Answerable for whether the output is correct enough to release.
Agents sit beneath all three as a tool layer. They produce work and hold no accountability, because accountability cannot rest with a party that cannot be held to account.
| Dimension | Traditional agile squad | Protum™ 3-2-1 |
| Membership | Stable, bounded, human | Fluid; humans and agents |
| Accountability model | Mutual, peer-to-peer | Named and asymmetric. Only humans are answerable |
| Quality assurance | Shared norm ("definition of done") | Explicitly owned by a named role |
| Decision speed | Governed by ceremony cadence | Governed by protocol and time-box |
| Primary failure mode | Diffusion of responsibility across peers | Diffusion of responsibility to the agent |
| Scaling constraint | Communication overhead | Human supervision capacity |
The structural argument for three roles rather than two or four is that human-agent work generates three distinct and separable accountabilities: what we are trying to achieve, how agents are deployed against it, and whether the result is fit to ship. Collapsing the third into the second is a common design error in our engagements, because it makes the party that deployed the agent also the party that judges its output.
Decision rights: what an agent may do without a human in the loop
The practical work of an operating model is deciding, in advance and in writing, which decisions an agent may take alone. The grid below is an illustrative starting point, not a standard. The specific placements matter far less than the fact that the grid exists and has been agreed.
| Decision type | Agent alone | Human approval required | Joint, named owner |
| Draft internal content | Yes | ||
| Analyse internal data | Yes | ||
| Publish external content | Yes | ||
| Commit code to production | Yes | ||
| Contact a customer | Yes | ||
| Alter a system of record | Yes | ||
| Approve spend | Yes |
Adapt these placements to your own risk appetite and regulatory position. Two rules travel across every organisation we have worked with: the person who deployed an agent should not be the person who judges its output, and every escalation trigger should be an observable condition rather than a matter of individual judgement.
The full framework, including its six capabilities and deployment model, is set out on the Protum™ operating model page.
Six capabilities of an AI-era high-performing team
Six organisational capabilities distinguish teams that convert AI adoption into performance from teams that only accelerate. Each modifies an established team-effectiveness condition rather than replacing it.
Protum defines six capabilities for AI-era operating models. The table below is our reading of how each relates to the established team-effectiveness literature, an interpretation offered to make the capabilities assessable against evidence, rather than a claim that the original researchers framed them this way.
| Protum™ capability | Established anchor | What changes under AI conditions |
| Data Culture | Structure and clarity (Google re:Work) | Clarity now includes data lineage and provenance. A team cannot own an output it cannot trace |
| Adaptive Structures | Hackman's stable team condition (2002) | Membership stops being bounded. Structure must be re-derivable rather than fixed |
| Augmented Craft | Backup behaviour (Salas et al., 2005) | Backup behaviour becomes human-over-agent supervision: the same construct applied to a new object |
| Responsible Intelligence | Psychological safety (Edmondson, 1999) | Safety must extend to disclosing agent error without implying supervisory failure |
| Flow-Based Interactions | Closed-loop communication (Salas et al., 2005) | Loops must close across human-agent handoffs, where acknowledgement is not guaranteed |
| Impact Prioritisation | Clear and engaging direction (Hackman, 2002) | Prioritisation must survive agent-driven volume. More options is not more direction |
The pattern across all six is consistent: AI does not invalidate the established conditions for team effectiveness. It changes what satisfying them requires.
What to measure when agents are in the workflow
Velocity and utilisation lose diagnostic value once agents enter a workflow, because both rise regardless of whether the work is correct. The measures that retain meaning are those tracking ownership and correction: rework rate on agent-originated output, escalation latency, and the proportion of outputs carrying a named human owner.
Most enterprise team metrics were designed under an assumption that has become false: that output volume is a reasonable proxy for output value, because producing more required more human effort. Once agents supply marginal capacity, that link breaks.
| Metric | Pre-AI | With agents in the workflow | Replace or supplement with |
| Velocity / throughput | Reliable | Inflates without any quality signal | Rework rate on agent-originated output |
| Utilisation | Reliable | Meaningless, because agent capacity is effectively unbounded | Human supervision capacity against review load |
| Cycle time | Reliable | Compresses while concealing accumulated review debt | Time to verified output |
| Engagement | Reliable | Still valid for the human layer | Unchanged |
| Defect / escape rate | Reliable | More important than before | Segment by human- vs agent-originated |
| n/a | n/a | New | Share of outputs with a named accountable owner |
| n/a | n/a | New | Escalation latency: trigger to human decision |
On engagement measurement
Engagement instruments remain valid for the human layer, and the evidence base behind them is stronger than almost anything else in this field. Gallup's Q12 meta-analysis, 11th edition (2024), draws on 736 studies covering 183,806 business units and 3,354,784 employees across 90 countries. It reports median differences between top- and bottom-quartile units of 23% in profitability, 18% in productivity (sales) and 78% in absenteeism.
A note on a widely-repeated error. The figure commonly quoted for absenteeism is 81%. That number is from the 10th edition (2020) and has been superseded. The current figure is 78%. The profitability (23%) and sales productivity (18%) figures are unchanged between editions. Several highly-ranked pages were still quoting the 2020 figure in 2026, in at least one case additionally misattributing it to a different Gallup instrument.
A 90-day sequence for enterprise leaders
Most enterprises should sequence AI team redesign in three stages: map where agents already operate and which outputs lack an owner, assign accountability and define decision rights, then instrument the new measures and run escalation live. Tool selection should follow, not precede, the ownership map.
Days 1–30: Map
- Inventory every workflow where an agent, copilot or model already produces or materially shapes output. Include unsanctioned use; in our experience it is often the larger set.
- For each, identify the human currently answerable for the result. Where no name emerges within thirty seconds, record it as unowned.
- Produce the unowned-output list. Treat that list as the operating-model backlog.
- Do not procure anything during this stage.
Days 31–60: Assign
- Assign the three accountable roles for each material workflow, using existing people.
- Define decision rights explicitly: what the agent may do without approval, what requires review, what requires escalation.
- Define escalation triggers as observable conditions, not judgement calls.
- Write down each role's accountability boundary, including what it is not answerable for.
Days 61–75: Audit the language
Where agents are described in employee terms, such as onboarded, hired or given a job title, change it. This is not cosmetic. It is the BCG finding applied: employee framing measurably reduced error detection and shifted accountability away from the humans supervising the work. Language is the cheapest lever in the entire sequence and the one most often skipped.
Days 76–90: Instrument and run
- Stand up the new measures set out above.
- Run the escalation protocol live on real decisions. Record every invocation.
- Review the unowned-output list. It should be shrinking; if it is not, decision rights were defined too loosely.
- Re-baseline: report quality-adjusted output, never raw throughput.
Executive readiness scorecard
Score each statement 0 (not true), 1 (partly true), or 2 (consistently true).
This is a structured conversation prompt, not a validated psychometric instrument. Its value lies in the disagreements the ten questions surface among a leadership team, rather than in the total.
| # | Statement |
| 1 | Every material AI-assisted output in our organisation has a named human owner |
| 2 | We have written decision rights specifying what agents may do without approval |
| 3 | Escalation triggers are defined as observable conditions, not individual judgement |
| 4 | Someone other than the person deploying an agent judges whether its output ships |
| 5 | We measure rework rate on agent-originated output separately from human-originated |
| 6 | We know our human supervision capacity and compare it to actual review load |
| 7 | Team members can disclose an agent error without it reading as personal failure |
| 8 | We can trace the data lineage of any AI-assisted output we publish or act on |
| 9 | Our internal language describes agents as tools, not as colleagues or employees |
| 10 | We report quality-adjusted output rather than raw throughput |
| Score | Reading |
| 0–6 | Pre-operating-model. AI is being adopted faster than it is being governed. Start with the unowned-output map. |
| 7–13 | Partial. Ownership usually exists but is inconsistent, and it typically fails first under time pressure. |
| 14–17 | Functioning operating model. The remaining gaps are usually measurement, not accountability. |
| 18–20 | Mature. Focus shifts to supervision capacity as the binding constraint on scale. |
A structured version of this assessment is available through Massivue's AI Maturity Assessment.
Frequently asked questions
How do enterprise leaders build high-performing teams in the age of AI?
In three stages, in this order. First, map every workflow where an agent already produces output and identify which of those outputs has no named human owner. Second, assign accountability and write down decision rights before selecting any tool. Third, instrument for quality and ownership rather than speed. Organisations that reverse this order buy capability they cannot govern.
Who is accountable when an AI agent produces a bad output?
A named human, always. Accountability cannot transfer to a party that cannot be held to account. BCG Henderson Institute research with more than 1,200 managers found that when AI was framed as an "employee" rather than a tool, managers identified 18% fewer errors and individual accountability for those errors fell nine percentage points.
Should AI agents be treated as employees or as tools?
As tools, on current evidence. BCG's May 2026 research found that employee-style framing measurably reduced error detection and diffused accountability toward the AI. This contradicts the prevailing "agents as teammates" narrative and is the strongest available argument for maintaining an explicit human ownership line in both structure and language.
What is agentic AI governance?
The layer of control specific to AI systems that act rather than only advise. In Deloitte's April 2026 framing it has three components: boundaries defining which decisions an agent may take independently, real-time monitoring of agent behaviour, and audit trails covering the full chain of agent actions. Only 21% of the 3,235 organisations Deloitte surveyed had a mature version of it.
Does psychological safety still matter when part of the team is a machine?
More than before. Edmondson's 1999 construct describes willingness to take interpersonal risk. Disclosing that an agent produced a flawed output carries a particular risk: appearing to have failed at supervision. Without safety, agent errors surface later, at higher cost, and further from the person who could have caught them.
What is the optimal size of an AI-augmented team?
No verified answer exists yet. McKinsey's The Agentic Organization asserts that two to five people can supervise 50 to 100 specialised agents, but publishes no primary research supporting it and describes the basis as consulting experience. Treat it as a hypothesis. The real constraint is human supervision capacity, which remains unmeasured in published research.
Which team metrics break when AI enters the workflow?
Velocity and utilisation break first, because both rise whether or not the work is correct. Cycle time compresses while hiding accumulated review debt. Metrics that retain meaning include rework rate on agent-originated output, escalation latency, defect rate segmented by origin, and the share of outputs with a named accountable owner.
Do classical team frameworks still apply?
Selectively, along a clear fault line. Of seven established frameworks assessed here, two hold, two partly hold and three break. Frameworks describing human relationships, such as psychological safety and shared purpose, still hold. Frameworks assuming bounded membership or mutual accountability break: Tuckman's stages (1965), Katzenbach and Smith (1993), and Hackman's stable-team condition (2002).
What is an AI operating model?
The set of decisions defining how an organisation structures work, roles, decision rights and governance once AI systems perform part of that work. It differs from an AI strategy, which defines what to pursue. Most enterprises have a strategy and no AI operating model, which is where pilots typically stall.
Why do most enterprise AI pilots fail to scale?
Because the operating model does not change alongside the technology. Deloitte's April 2026 survey of 3,235 leaders across 24 countries found only 21% have a mature agentic AI governance model. Pilots succeed under controlled conditions with informal oversight and fail when they meet organisational boundaries that were never redefined. We have written separately on why enterprise AI pilots stall before production.
What is the Protum™ 3-2-1 model?
Massivue's operating structure for human-agent teams. It assigns ownership across three accountable human roles: a Value Lead, an AI Orchestration Lead, and an AI Quality Engineer. Its purpose is to prevent accountability diffusion when agents join a team whose ownership structure assumed all members were human.
How is the 3-2-1 model different from a traditional agile squad?
An agile squad assumes stable, bounded, human membership with mutual peer accountability. The 3-2-1 model assumes fluid membership including non-human participants and replaces mutual accountability with named, asymmetric ownership, because only humans can be answerable. Quality assurance becomes an owned role rather than a shared norm.
Do we need to hire new roles, or can existing people take these on?
Existing people, in most enterprises. The roles describe accountabilities, not headcount. Common mappings are a product or business owner to Value Lead, a technical lead or architect to AI Orchestration Lead, and a quality or risk function to AI Quality Engineer. In our experience hiring is rarely the constraint; undefined decision rights are.
How long does restructuring a team around AI agents take?
In our engagements, role assignment and decision rights can generally be defined within 90 days. Behavioural change takes longer. The limiting factor is usually reaching agreement on escalation authority, meaning who can overrule whom and under what conditions, rather than technical implementation.
What is the difference between AI governance and an AI operating model?
Governance defines constraints: what must not happen, and who is answerable if it does. An operating model defines execution: how work flows, who decides, and how roles interact daily. Governance without an operating model produces policies nobody applies. An operating model without governance produces speed without control.
How do we measure whether AI is actually improving team performance?
Compare quality-adjusted output rather than raw output. Atlassian's State of Teams 2026 found 89% of executives report AI-driven speed gains while only 6% are sure they have clear examples of organisation-wide AI ROI. Speed measured without a quality and ownership denominator is not a performance measure.
Does AI widen or narrow the performance gap between teams?
Current evidence points to widening. Atlassian's 2026 research found 55% of executives believe AI is widening performance and opportunity gaps between teams. Microsoft's 2026 Work Trend Index found organisational factors accounted for more than twice the AI impact of individual factors, which is to say the differences are organisational, not individual.
What capabilities do teams need to work well with AI?
Three layers. Supervisory judgment: knowing when to accept, reject or escalate agent output. Workflow design: restructuring processes rather than inserting AI into existing steps. Accountability literacy: understanding where ownership sits and when it transfers. Tool training alone produces none of the three, which is why AI workforce transformation has to be designed around workflows rather than software.
Where should an enterprise start?
Map where agents already operate in existing workflows and identify which outputs have no named human owner. That gap list is the operating-model backlog. Selecting tools or launching a capability programme before the ownership map exists is, in our experience, the most expensive sequencing error in enterprise AI.
Conclusion
The question enterprise leaders bring to this topic is usually framed as capability: what do our teams need to be able to do. The evidence points somewhere less comfortable.
Adoption is not the constraint. Governance is. And the organisations that have closed that gap did not do it by adding capability. They did it by deciding who is answerable for what, and then holding that line, including in the language they use, which BCG's research suggests matters more than anyone expected.
The classical team literature is not obsolete here. The part that holds is the part about how humans relate to each other under uncertainty. What breaks is everything built on the assumption that a team is a bounded group of people who can all be held responsible. That assumption is now false, and no amount of tooling restores it.
High performance in the age of AI is not a throughput property. It is an ownership property. The teams that will pull ahead are not the ones producing the most, since everyone will produce more, but the ones who can still answer, quickly and without ambiguity, a question that used to be trivial: who owns this?
Where to start. Take the ten-question scorecard above to your next leadership meeting and score it as a group. The disagreements will tell you more than the total. If you want a structured version with benchmarking, the AI Maturity Assessment covers the same ground in more depth.
Related reading
- Protum™: the AI operating model. The full framework, including its six capabilities.
- AI Operating Model consulting. The structural layer this article assumes.
- AI-First Operating Model: A Framework for Enterprises Beyond Pilots.
- Why enterprise AI pilots stall before production.
- AI Workforce Transformation. Capability building at organisational scale.
- Enterprise Transformation.
- AI Maturity Assessment. A structured version of the scorecard above.
Sources
All figures verified against primary sources on 4 August 2026.
- Atlassian, State of Teams 2026, 27 April 2026. 12,035 knowledge workers, 173 Fortune 1000 executives, double-blind, fielded January to February 2026.
- Deloitte, AI agents are scaling faster than their guardrails, 24 April 2026. 3,235 IT and business leaders, 24 countries.
- BCG Henderson Institute, Why You Shouldn't Treat AI Agents Like Employees, 6 May 2026. More than 1,200 managers.
- Microsoft, Work Trend Index 2026, 5 May 2026. 20,000 knowledge workers, 10 markets.
- Deloitte, Human capabilities are at the heart of high-performing teams, 14 January 2026. 1,394 US professionals.
- McKinsey, The Agentic Organization, 26 September 2025. No primary research disclosed.
- Gallup, Q12 Meta-Analysis, 11th Edition, May 2024 (updated July 2024). 736 studies, 183,806 business units, 3,354,784 employees, 90 countries. Confirmed as the current edition on 4 August 2026.
- Google re:Work, Understand team effectiveness (archived record, originally published 2015). 180 teams.
- Edmondson, A. (1999). Psychological Safety and Learning Behavior in Work Teams. Administrative Science Quarterly, 44(2), 350–383.
- Salas, E., Sims, D. E., & Burke, C. S. (2005). Is there a "Big Five" in Teamwork? Small Group Research, 36(5), 555–599.
- Tuckman, B. W. (1965). Developmental sequence in small groups. Psychological Bulletin, 63(6), 384–399.
- Tuckman, B. W., & Jensen, M. A. C. (1977). Stages of Small-Group Development Revisited. Group & Organization Studies, 2(4), 419–427.
- Hackman, J. R. (2002). Leading Teams: Setting the Stage for Great Performances. Harvard Business School Press.
- Lencioni, P. (2002). The Five Dysfunctions of a Team. Jossey-Bass.
- Katzenbach, J. R., & Smith, D. K. (1993). The Wisdom of Teams. Harvard Business School Press.