All Blogs
How AI Transfers Knowledge and Expertise Back to Humans: The Feedback Loop That Works
Capturing expertise in an AI avatar is half the job. Harvard, Stanford, and Wharton research shows how to get that knowledge back out — and make it stick.

Every GTM leader already knows the most effective way to train a rep: one-to-one coaching, with feedback that lands in the moment. The research proving it is four decades old. The problem has never been knowing what works. It’s that almost nobody can afford to deliver it.
Capturing a top performer’s expertise into an AI avatar solves half of that equation — the half we covered in our companion piece on human-to-AI knowledge transfer. This piece is about the return trip: getting that expertise back into reps’ hands. Here’s the short version. There’s a coaching problem the industry has lived with since 1984, a growing body of university research showing AI can now solve it, one well-documented way the solution backfires, and three things GTM leaders should do about it.
TL;DR
- The problem: Benjamin Bloom’s landmark 1984 study found that one-to-one tutoring beats classroom instruction by two standard deviations. It’s the gold standard almost nobody gets: only 26% of sales professionals receive 1:1 coaching at least weekly, per Salesforce research.
- The evidence: A 2024 Harvard randomized controlled trial found that students using a purpose-built AI tutor learned more than twice as much, in less time, than peers in an active-learning classroom.
- In a Stanford randomized controlled trial covering roughly 900 tutors and 1,800 students, AI assistance helped lower-rated tutors most, closing up to 9 percentage points of the performance gap with their higher-rated colleagues.
- The caveat: Wharton School research found that unrestricted, on-demand AI help can undermine learning by removing “productive struggle,” the mechanism through which real skill gets built.
- What to do: pair AI role-play with in-the-moment coaching, calibrate the coach to challenge reps rather than hand them answers (and certify mastery through simulations without the coach present), and keep line managers in the loop for the human element.
The problem: the best coaching is the one almost nobody gets
In 1984, educational psychologist Benjamin Bloom compared three ways of teaching the same material: a conventional classroom, a mastery-learning classroom, and one-to-one tutoring. The tutored students didn’t just do better. On average, they performed two standard deviations above their conventionally taught peers, ahead of roughly 98% of the classroom group. Bloom treated the finding as a problem rather than a victory, and spent much of the paper asking whether group instruction could be redesigned to match tutoring, precisely because paying for the real thing was out of reach.
Sales organizations live inside the same problem, with two complications layered on top. The first is scarcity. Only 26% of sales professionals receive 1:1 coaching at least weekly, according to Salesforce research, because frontline managers carry the coaching load alongside forecasting and deal support, and mock calls are usually the first thing to fall off the calendar. Coaching quality also varies from manager to manager — a gap the Stanford research below measures directly. The second complication is timing. When coaching does happen, it breaks a basic rule of learning: a rep runs a mock call on Tuesday, feedback arrives the following week from a reviewed recording, and by then the moment it refers to is half-forgotten and the mistake has been repeated on live calls.
The cost of leaving this unsolved is familiar to anyone who has run a sales team: long ramps, wide performance gaps between reps, and expertise that stays locked in the heads of a few top performers. Everyone agrees on what good looks like. Almost nobody can deliver it to every rep, every week.
The evidence: AI can close the coaching gap — if it’s built right
Over the past two years, university researchers have run controlled experiments on exactly this question, and the results sort cleanly into three findings. AI coaching works, and it’s always available. Too much help backfires. And the human layer still matters. Each comes from a different study, and together they define what “built right” means.
AI coaching works, and it arrives in the moment
At Harvard, physics lecturers Gregory Kestin and Kelly Miller ran a randomized controlled trial with 194 undergraduates, pitting a custom AI tutor against a live, instructor-led active-learning class covering the same material. Students working with the AI tutor learned more than twice as much in less time, as reported by the Harvard Gazette, and rated the experience more engaging and motivating.
At Stanford, researchers built an AI copilot for human tutors and tested it in a randomized controlled trial spanning roughly 900 tutors and 1,800 students. The study found the largest gains among students working with lower-rated tutors, whose outcomes improved by up to 9 percentage points, enough to close most of the gap with stronger tutors. If coaching quality varies across your management team, this is the finding to sit with: AI assistance lifted the floor.
Just as important is when the help arrives. A human role-play partner will almost always let you finish, out of politeness or time pressure, and a general-purpose LLM tends to flatter you by design. A system built to coach carries neither instinct. Mahalia VanDeBerghe, who works in sales enablement at Cavallo, described the difference: “I love that she stopped me. I have not had a role play do that where it stops me mid-conversation. I’ve only ever had it take me all the way through. That’s really cool. I’m very impressed.”
Too much help backfires
Not every version of AI-assisted practice helps. A Wharton School study led by professor Hamsa Bastani tested nearly 1,000 high school students on two versions of an AI tutor: one that offered direct, on-demand answers, and one built with guardrails that nudged students toward reasoning. Students with unrestricted access achieved less than half the performance gains of students using the guarded version, even though they felt more helped in the moment. And when AI access was later removed, they performed measurably worse than students who had never used AI at all.
“As educators, we worry about that,” Bastani said of the risk that learners treat AI as a crutch instead of building the underlying skill. The finding doesn’t argue against AI coaching. It argues that the design of the coach — whether it makes a rep think or simply hands over the answer — determines which outcome you get.
The human layer still matters
The researchers behind the strongest pro-AI results were also the clearest about where the technology’s advantage ends. Susanna Loeb, an education researcher on the Stanford study, put it plainly: AI is well suited to figuring out what a learner knows and needs next, but “people are better at caring, motivating and engaging, and celebrating successes,” she said. That human element, in her view, is a core part of why tutoring works at all. Harvard’s Kestin drew the same boundary from his own results: “This allows for the human interaction to be much richer,” he said — the AI handles first-pass instruction so that time with a human can go to deeper discussion.
What this means: three keys to making AI coaching work
Taken together, the research points to three operating principles for GTM teams.
1. Pair role-play with in-the-moment coaching; simulated calls alone aren’t enough
Simulation is the delivery vehicle, but coaching inside the simulation is the active ingredient. A rep who runs ten mock calls and gets a score at the end learns less than a rep who gets stopped, corrected, and made to try again at the moment of the mistake. Misha McPhearson, VP of Revenue Enablement at Betterworks, described why the combination is what makes knowledge usable: “I would describe Avarra as my favorite tool for taking information and making it real. Unless reps have a chance to practice in a way that feels relevant to their day-to-day, it’s not going to stick. Avarra makes information sticky.”
The proof that the pairing works shows up in live deals, fast. An Account Executive at PagerDuty: “This was so helpful! What was actually most helpful was the coaching sessions. It was very cool to walk through a real deal and she even helped me prep a call I was going to have shortly after the session.” “Honestly, it helped me a lot,” said Anna Cucci, an AE at EvenUp Law. “I just got off a demo and felt that my disco was way better.”
2. Calibrate the coach to challenge, not to answer — and certify mastery without it
The Wharton results translate into two design rules. First, the coach should be tuned to push reps through the reasoning — asking the follow-up question, flagging the miss — rather than supplying the response a rep should have given. Second, certification should happen in simulation without the coach present. Wharton’s most sobering finding was what happened when the AI was taken away; testing reps without the helper is how you learn whether the skill transferred to the rep or stayed in the tool.
Customers describe both halves. Tiffany King, an AE at EvenUp Law, pointed to calibration: “Once you start to realize he interacts fairly realistically… it’s nice to have a ‘real’ role play scenario with feedback.” Her colleague Shelby Jones pointed to certification: “I thought both the training and the cert bot were super valuable… Seriously a tool I enjoyed using and will continue to use.” And Kyle Norton at Owner.com described what a well-calibrated certification loop does to rep behavior: “We have SDRs in there at 9pm on their fourth day cranking out SIMs to win a $100 prize for scoring 100%. They rack up hours of sim time.” Nobody volunteers for 9pm practice on day four of a new job because a system feels fake.
3. Keep the line manager in the loop
AI should absorb the repetitive, low-judgment portion of coaching: mock calls, objection drills, certification runs. It should not absorb the one-on-one. Marchelle Mooney at Mangomint described the division of labor in practice: “Frontline managers aren’t spending any time running mock calls — all of that is done in Avarra. The hours we’re giving back to managers is almost not even quantifiable.” Adoption followed the same logic: “There was zero level of resistance from frontline managers. If anything, they clearly saw: ‘I will not have to do one-on-ones that are just mock scenarios possibly ever again.’”
What managers do with those hours is the layer the Stanford research protects: personal context, career guidance, motivation, celebrating the win a rep just closed. That part doesn’t automate, and the same research that validates AI coaching identifies it as central to why coaching works at all.
Frequently asked questions
Is there real academic research behind AI-based coaching, or is this mostly vendor claims?
There’s a growing body of peer-reviewed, university-conducted research. Harvard’s randomized controlled trial and Stanford’s Tutor CoPilot study are two of the more rigorous examples; both used controlled experimental designs rather than self-reported satisfaction surveys.
Doesn’t research also show AI can hurt learning?
Yes, and the nuance matters. Wharton School research found that unrestricted, on-demand AI assistance can undermine skill-building by removing the productive struggle learning depends on. Whether a tool pushes the learner to reason or simply hands over an answer determines which outcome you get.
What’s the fastest sign that AI-to-human knowledge transfer is working on a team?
Voluntary, after-hours practice. When reps run simulations on their own time, not because they were assigned but because the feedback feels worth it, the transfer mechanism is working rather than just present.
How is AI-delivered feedback different from a manager’s feedback?
The real differences are timing and consistency, not quality. AI feedback lands inside the moment of the mistake, every time, for every rep. A manager’s feedback carries context and judgment an AI doesn’t have, but it arrives less often and less predictably, mostly because of time constraints. The two are complements, not substitutes.
Should every practice session include AI feedback, or are some better left unscored?
Most enablement leaders land on a mix. Early, low-stakes practice can stay forgiving and exploratory, while certification-track simulations benefit from consistent, specific scoring. The Wharton research suggests the risk isn’t feedback itself but the removal of all friction; a system that scores accurately while still requiring the rep to work through the problem tends to outperform one that hands over the right answer.
Can this kind of AI-driven coaching replace manager one-on-ones entirely?
No, and the research doesn’t support that reading. What it replaces is the repetitive, low-judgment portion of coaching: mock calls, objection drills, certification runs. One-on-ones remain the right place for personal context, career guidance, and motivation, which the Stanford research specifically identifies as human strengths.
Why do some reps practice voluntarily while others only do the minimum required?
Customers report the difference usually comes down to whether the feedback feels real and specific. Reps who get generic, surface-level scoring treat practice as a compliance task. Reps who get feedback tied to something they can recognize and fix come back on their own, often without being asked.
In conclusion
The story here is old, and only the ending is new. Bloom proved in 1984 that one-to-one coaching with immediate feedback is the most effective way to build skill, and for forty years the conclusion stayed the same: too expensive, can’t scale. The Harvard and Stanford trials are the first rigorous evidence that the economics have changed. A well-built AI coach can outperform strong group instruction, and it can lift the weakest coaches toward the level of the best. Wharton supplies the discipline the approach needs: preserve productive struggle, or you build reps who feel helped and perform worse.
The playbook that falls out of the research is short. Give reps realistic role-play with coaching that arrives in the moment, not the following week. Tune the coach to challenge rather than answer, and certify mastery in simulations without it. Keep managers doing the one thing no model does well — the human part. And remember that none of it holds up if the expertise going into the system was captured badly in the first place, which is the half of the loop covered in the companion piece.
The signal that it’s all working was never going to be a completion dashboard. It’s Owner.com’s SDRs, running sims at 9pm on their fourth day, because the feedback is making them better and they know it.
See how Avarra turns captured expertise into practice reps actually use →
What our customers are saying
Results from teams using Avarra




















































































































