AI’s Homework Problem
15 September 2026
Committing the Company Before Knowing Who Earned the Marks

AI-generated summary
A teacher marking homework cannot tell how much a parent did, and a chief executive reading finished AI-assisted work faces the same mixing of contributions. Foster-Fletcher sets a BetterUp Labs and Stanford Social Media Lab survey, in which 1,150 US desk workers reported nearly two hours per incident dealing with polished AI content that lacked substance, against Jamie Dimon's commitments on AI and JPMorgan's workforce. Dimon excludes employees' reported time savings from the bank's valuations, and Foster-Fletcher asks what he is told about the correction that makes the output usable.
A teacher marking a piece of homework sees what is handed in and if a parent has helped, then the child's ability and the parent's help are mixed on the same page. The teacher would have to ask the child, or set the same exercise in class, to separate the two.
The same situation now presents itself in professional work. If an AI draft needs substantial human rewriting before it is usable, the finished work is a poor guide to what the AI could have done on its own. The familiar idea that the AI drafts and the human finesses can understate the work involved. In an August–September 2025 survey of 1,150 full-time US desk workers by BetterUp Labs and the Stanford Social Media Lab, respondents reported spending an average of one hour and 56 minutes per incident dealing with AI content that looked polished but lacked the substance to be usable. About 40 per cent reported receiving such content from colleagues in the previous month; using that prevalence, the researchers estimated the cost in lost productivity at about $9 million a year for a company of ten thousand people.
A chief executive might not be concerned with how the work gets done, so long as the output looks good and the method fits their AI narrative.
I believe they should concern themselves with exactly how the work gets done. If they are making public commitments about the future of their organisation on the strength of what they believe AI can do, they need to know how much of that apparent capability depends on employees correcting the AI or deciding which parts of its work can be trusted.
Jamie Dimon, chief executive of JPMorgan Chase, has made commitments about AI’s effect on the bank’s workforce. His April 2024 letter to shareholders said that AI “may reduce certain job categories or roles, but it may create others as well” and promised retraining and redeployment for employees affected by the change. At the investor update on 23 February 2026, he said: “We have displaced people from AI. And we offered them other jobs.”
At the same update, Dimon also said that employees’ estimates of time saved through the bank’s internal AI platform were excluded from its net present value calculations. He was not seeing those reported savings reflected in corresponding reductions in headcount.
I would ask a chief executive making those commitments how much the AI is delivering and how much of its apparent success depends on human judgement and correction. If the workforce is to shrink, I would also want to know what evidence shows that the same standard of work can be maintained with less of the expertise that helped produce it.
The human reviewer may have supplied twenty per cent of the finished document or seventy, and ten per cent of the system's draft may have needed correcting or ninety. Those are two separate unknowns. A reviewer may know roughly how much of a draft they rewrote. It is harder to put a figure on the judgement needed to check the rest and decide it was usable. If that work is overlooked, the AI gets credit for expertise supplied by the reviewer.
If a chief executive judges the AI's impact from the collective final output, they may overestimate what it contributed. Dimon’s caution about reported time savings leaves me wanting to know what he is told about the work involved in making AI’s output usable. Recognising the need for human judgement does not, by itself, tell him how much of an apparently successful result depended on someone repeatedly correcting a poor answer or repeatedly deciding which parts could be trusted.