Human–AI teaming is a measurable design, not a feeling
AI as tool, teammate or gate; who gains, who deskills, and how to pre-register the metric
The question: When you put people and a model on the same task, what actually changes in output, in quality, and in what the people learn — and how would you know?
The 2025 field experiments turned "humans plus AI" from a slogan into measured designs: gains are real, unevenly distributed, and depend on how the collaboration is shaped. That makes team design something to test, not assume.
What the lesson covers
Three designs for the same task. **AI as tool**: the person does the work and calls the model for pieces. **AI as teammate**: the model participates through the task — proposing, drafting, critiquing — and the person steers and decides. **AI as gate**: the model checks or scores the person's work. Each changes output, quality and learning differently, and the differences are measurable.
The evidence. Brynjolfsson, Li and Raymond (QJE 2025): a support assistant raised productivity ~14% on average, most for the least experienced — the tool transferred the top performers' know-how down the ladder. Dell'Acqua and colleagues' P&G field experiment ("The Cybernetic Teammate", 2025): individuals with AI matched the performance of two-person teams without it, and AI-enabled teams produced more balanced, cross-functional work — AI acted like a teammate, not just a tool. The gains were real; so was the redistribution of who contributes what.
The cost side is learning. Cognitive offloading research (Lesson 34) and the MIT "Your Brain on ChatGPT" study (2025) found weaker engagement and recall when the model did the thinking step — which matters most for the people who still have to become experts (the apprenticeship problem, Lesson 38). A design that maximises today's output can quietly starve tomorrow's capability, and only measurement shows it.
So treat team design as an experiment with a **pre-registered metric**: before the trial, write the outcome you will judge it on (quality, time, cost per completed task, and a learning measure), the comparison (tool vs teammate vs gate, or vs no AI), and the duration. Run it small, on real work, for a week or two. Decide on the number, not the mood — Lesson 32's discipline at team scale.
The organisational takeaway is that "how we work with AI" is a design decision per task, not a policy per company: some tasks want the model as a gate (verification, Lesson 38), some as a teammate (exploration), some as a tool (production) — and some want a human only, because the learning the task produces is worth more than the output.
Key points
- Three designs — tool, teammate, gate — change output, quality and learning differently.
- Gains are real and uneven: ~14% for support agents, most for the least experienced (QJE 2025); AI-enabled individuals matched two-person teams (P&G, 2025).
- The cost is learning: offloading the thinking step can starve tomorrow's experts.
- Pre-register the metric — including a learning measure — then trial small, on real work, and decide on the number.
- Team design is per task, not per company: gate for verification, teammate for exploration, tool for production, human-only where the learning is the point.
Framework — Tool / Teammate / Gate → Pre-register → Trial → Decide on the number
Three designs per task, a metric written before the trial (including learning), a small real trial, a decision on evidence.
The lab
Design and pre-register a human–AI teaming trial for one real task.
Open this lesson, its lab and its quiz
Sources and further reading
- Generative AI at Work — Brynjolfsson, Li & Raymond, QJE 2025
- The Cybernetic Teammate — Dell'Acqua et al., 2025
- Lesson 34 — cognitive offloading — Applied AI Academy