The Empty Gym

Share

6 October 2026

Desirable difficulties: the effort we should not spare

El judici no es delega · 5 de 12

It is easy to believe that a good lesson is one that can be followed without effort, that good material is material that is understood first time, and that a student who moves along fluently is learning well. It is what we feel when we explain and what they feel when they study, and that is why it is so hard to distrust that feeling. In 1992, two psychologists at the University of California, Los Angeles, Robert and Elizabeth Bjork, formulated the opposite within a framework they called the New Theory of Disuse, and they have been gathering evidence in its favour ever since: the conditions that make learning easier and more fluent in the short term tend to produce poorer retention and poorer transfer in the long term (1).

It is not that effort is good for its own sake. It is something rather more uncomfortable, namely that certain difficulties do not accompany learning, they constitute it. Whoever removes them does not save work. They cancel the effect of learning.

Desirable Difficulties

What fluency hides

The Bjorks’ contribution consists in distinguishing two strengths within a single memory. Storage strength indicates how well learned something is, that is, how interconnected it is with everything else we know; retrieval strength indicates how easy it is to access right now, with the cues we have now. What is genuinely new in their theory is not the distinction, which they themselves attribute to earlier authors, but how the two strengths interact: the higher the current ease of access, the smaller the gain in learning produced by restudying or retrieving again. Forgetting a little, therefore, can enhance learning. That is why spacing and variation work, worsening visible performance while improving what is actually learned (1).

From this follows the finding that speaks most directly to a teaching staff, and which the Bjorks place in the territory of metamemory. If we interpret today’s performance as a measure of learning, we are not only wrong in judging whether learning has taken place: we end up, in addition, ‘preferring poorer conditions of learning over better conditions of learning’. The preference is the students’ and it is also ours. Douglas Rohrer, who investigated spacing in real mathematics classes, found strong support for interleaving problems rather than grouping them by type, and observed at the same time that most workbooks and classroom exercises do just the opposite. Together with Marissa Hartwig he points to the probable reason, which is the same one the series named in the second post with another protagonist: blocked practice produces an ‘illusion of mastery’ that is difficult to overcome (1).

All of this predates generative artificial intelligence. And the problem was not created by the machine, it was made impossible to ignore by it. Alejandro Espeso-García, of the Catholic University of Murcia, describes the mechanism precisely in his review on cognitive enhancement and offloading: by delivering the outcome without the prior cost of effort, generative AI creates a state of illusory efficiency, and by flattening the learning curve it can deprive the student of the ‘cognitive gymnasium’ where critical thinking develops. He calls it the efficiency paradox (2).

Relying on automated solutions creates an illusion of efficiency—one that quietly bypasses our cognitive gym.
Desirable Difficulties

Not every difficulty teaches

At this point it is worth slowing down, because this argument degrades very easily. The Bjorks themselves insist that the important word in the expression ‘desirable difficulty’ is desirable, and that there is a wide repertoire of difficulties that are undesirable during instruction and forever after. The criterion that separates one from the other is concrete and has nothing moral about it: a difficulty is desirable when it triggers encoding and retrieval processes that support learning, and it ceases to be so as soon as the student does not have the background knowledge or the skills needed to respond to it successfully. The optimal level of difficulty, therefore, is not a property of the task, but a relation between the task and whoever does it (1).

Nicola Hodges and Keith Lohse, in the same special issue, propose three criteria for recognising such difficulties in advance, and they are the three a department can apply in a meeting: that the difficulty be task relevant; that it be novel, that is, not something the learner is already doing; and that it be potentially solvable by the learner (1).

The warning counts double when AI comes in, and there is a small piece of data that illustrates it well. At a German secondary school, as José Antonio Bowen and C. Edward Watson report in Teaching with AI, a teacher taught his students the flaws of AI tools and allowed them to use them in exams. None relied on the machine alone, and several went further with additional searches of their own. Those who did worst were the ones who arrived with less confidence in the content or in their reasoning, because they adopted the generated text without criticising it (3). It is the same idea that Amaia Arroyo Sagasta, doctor in Communication and Education and lecturer at Mondragon Unibertsitatea, and Carlos Magro Mazo, coordinator of the educational programme of the Institución Libre de Enseñanza, formulate from the standpoint of assessment in the ODITE Report 2026: learning depends on a solid base of prior knowledge and ‘always requires overcoming a friction’ (4). Removing the friction for someone who already has command of the content saves them time. Removing it for someone who does not prevents them from acquiring it.

Melissa Knabe and Haley Vlach point out, in that same special issue, that spacing is one of the most extensively studied phenomena in learning and, even so, it has barely been researched in children’s school learning. They predict, moreover, that individual differences in attention, memory, prior knowledge and metamemory, all of them developing rapidly and at different rates, will determine which child benefits and which does not (1). The honest conclusion is that we have a very robust principle and an incomplete map of its application by age. What was already noted in the previous article about the maturation of executive functions in adolescence reinforces that caution; it does not replace it.

Desirable Difficulties
Removing cognitive friction for those who haven’t mastered the basics prevents real knowledge consolidation.

Which difficulty each task protects

Bowen and Watson report a distinction from the mathematical biologist David Krakauer that puts all of the above in order. There are complementary cognitive artifacts and competitive cognitive artifacts. Arabic numerals or the abacus are complementary because they amplify our abilities and leave them amplified even when the artifact is no longer in front of us; GPS and the calculator are competitive, because when they disappear we are no better at the original task and are often worse. And the authors add the only honest thing that can be said today, which is that ‘it is hardly clear which type of cognitive artifact AI will turn out to be’ (3).

Nothing guarantees that it will be one or the other, because it is not decided by the tool. It is decided by the design of the task. In the previous article the criterion of sequence was settled, that AI should be used after the first effort and never in its place. The question that remains is the hard one, namely what exactly is the effort that needs protecting in each particular task.

Luis Miguel Iglesias Albarrán, Mathematics teacher and head of the IES San Antonio in Bollullos Par del Condado, proposes in the ODITE Report 2026 an answer that fits on one page and works in any subject. He calls it a thinking trail, and it consists of four moments: an initial attempt of one’s own, which seeks authenticity and not correctness; a comparison with resources or with AI itself, to ask for alternatives or detect inconsistencies, which is comparing and not copying; a verification of what has been asserted, which he places at the core of learning; and a final improvement with a brief note on what was changed and what was learned through verifying. His formulation of the reason why is the best in the chapter: a shortcut leaves no trail (5).

The design of the interaction also counts. The templates Bowen and Watson offer for turning a model into a tutor include an explicit instruction: that the AI should not do the work, that it should ask questions instead of rewriting, and that it should give explanations and examples only after the student has tried. David Malan, who runs Harvard’s introductory computer science course, sums up in a phrase why that instruction is needed: current tools are ‘currently too helpful’ (3). And in the same book is the warning that closes the argument: learning to recognise good work normally starts with doing poor work and then gradually improving, so the outstanding question is how students will learn the basics when the machine makes all the entry-level tasks easier (3).

We must decide whether artificial intelligence will serve as a complementary tool or a competitive replacement.
Desirable Difficulties

What to do with this

In the classroom. Redesign one habitual task, just one, using the four moments of the thinking trail. Iglesias sets it out subject by subject: in Mathematics, initial approach, alternative strategy, checking and brief explanation; in Science, hypothesis, comparison, verification and final write-up; in Social Sciences, comparison of sources and verification of claims; in Language, own draft, assisted revision justifying the changes, coherence check and a final, argued version (5).

In the classroom. The AI Literacy Framework of the OECD and the European Union, published in 2026, describes a whole domain devoted to this and states it thus: deliberately dividing work between humans and AI. Among its activities there is one that works from primary school onwards: students position themselves in four corners of the classroom labelled ‘AI only’, ‘AI-assisted’, ‘Humans only’ and ‘I don’t know’, according to how they think an everyday task should be tackled, and then defend their choice before those who chose differently, with permission to change corners if they are convinced. The advanced version is exactly the exercise of this article: break down the steps of a written assignment, choosing the topic, looking for evidence, ordering the arguments, drafting, revising, and decide in which of them AI can help and in which one’s own reasoning is needed (6).

For the leadership team. A departmental agreement with two questions per task: what is the cognitive process that this task teaches, and does a trail of that process remain. If the first cannot be answered in one sentence, the problem is not AI. Iglesias proposes starting with a school code of practice and a single pilot unit, without grandiose plans (5). And it is worth bearing in mind Arroyo and Magro’s warning: this does not depend on the goodwill of each individual teacher, but on cultural, curricular and systemic decisions, and as long as marking goes on occupying centre stage, the assessment of process will go on living at the margins (4).

Bowen and Watson sum it up in a sentence that no longer admits of argument: ‘All assignments are now AI assignments’ (3). The only decision left is which of them protects something.

Desirable Difficulties
Leadership teams must establish evaluation criteria that prioritize the learning process over the final grade.

Food for thought

Of next term’s tasks, could you say in one sentence what difficulty each one teaches? And if for some it cannot be said, why is that task still being set?

Identifying the cognitive challenge behind each task is the prerequisite for assessing it.
Desirable Difficulties

Glossary

  • Desirable difficulty. An obstacle deliberately introduced into learning that worsens immediate performance and improves long-term retention and transfer. Spacing study, interleaving types of problem or retrieving from memory are examples.
  • Undesirable difficulty. An obstacle that only gets in the way, because the student does not have the background knowledge or the skill needed to overcome it. The same task can be one thing or the other depending on who does it.
  • Storage strength and retrieval strength. The two properties of a memory according to the Bjorks’ theory. The first is how well learned something is; the second, how easy it is to access at this moment. Only the second is visible.
  • Transfer. The ability to apply what has been learned to a situation different from the one in which it was learned. Together with long-term retention, it is what desirable difficulties improve.
  • Complementary and competitive cognitive artifact. A distinction drawn by David Krakauer. The complementary artifact increases our abilities and leaves them increased when it is no longer there, like Arabic numerals; the competitive one does the task for us and leaves us worse at it, like GPS.
  • Thinking trail. A requirement added to a task so that evidence remains of the student’s decisions and revisions: own attempt, comparison, verification and explained improvement. The focus shifts from whether AI was used to how the work was done.

References

  1. Bjork, R. A., and Bjork, E. L. (2020). Desirable difficulties in theory and practice. Journal of Applied Research in Memory and Cognition, 9(4), 475–479.
  2. Espeso-García, A. (2025). Generative artificial intelligence: Between enhancement and cognitive offloading. Cultura, Ciencia y Deporte, 20(66).
  3. Bowen, J. A., and Watson, C. E. (2024). Teaching with AI: A Practical Guide to a New Era of Human Learning. Johns Hopkins University Press.
  4. Arroyo Sagasta, A., and Magro Mazo, C. (2026). Cuando evaluar es aprender. Evaluación auténtica para una educación centrada en el alumnado [When assessing is learning. Authentic assessment for student-centred education]. In J. M. Muñoz, N. Lorenzo, X. Suñé, M. À. Prats and C. López (Coords.), Informe ODITE 2026: Claves para una nueva educación [ODITE Report 2026: Keys for a new education]. ODITE, Asociación Espiral, Educación y Tecnología.
  5. Iglesias Albarrán, L. M. (2026). IA y aprendizaje. Del producto final al proceso [AI and learning. From the final product to the process]. In Informe ODITE 2026: Claves para una nueva educación [ODITE Report 2026: Keys for a new education]. ODITE, Asociación Espiral, Educación y Tecnología.
  6. OECD and European Union (2026). Cómo preparar a los alumnos para la era de la IA: Marco de competencias sobre IA para la educación primaria y secundaria [Empowering learners for the age of AI: AI literacy framework for primary and secondary education]. OECD Publishing. ‘Managing AI’ domain.
Rate the post
5/5 - (5 votes)

Share

2026-10-06T16:31:41+02:00
Go to Top