No Warning Light

Share

21 September 2026

Hallucination is not a system fault; it is how the system works

Judgement Cannot Be Delegated · 2 of 12

A well-presented assignment, with its bibliography in order, its years, its volumes and its DOIs. The teacher checks the first reference and cannot find it. Checks the second, and nothing either. By the fourth, they already know what they will find in the rest. What is unsettling about that scene is not the error, which students have been making for as long as schools have existed, but the finish, because nobody invents a citation and then takes the trouble to give it a credible digital identifier. The machine does. And it does not do so to deceive anyone.

It is often said that these tools make mistakes and that engineers will polish away the fault version after version. It is a reasonable idea, since that is what we have seen happen with almost every technology that has entered schools. With this one it does not work like that. What we call hallucination is not a fault that can be separated from everything else: it is the very mechanism that allows the tool to do its job.

al·lucinacions de la IA

What is not written anywhere

In the previous article we saw that a language model calculates which word is most likely to come next and repeats the calculation until the sentence is finished. From this follows a consequence worth examining slowly: a system like this will produce a continuation even when there is nothing to continue. José Antonio Bowen and C. Edward Watson explain it precisely in Teaching with AI: because these models generate by sorting the probabilities of a next word or pixel, ‘they are prone to “generate” false data or “fabricate” fictional references’, and ‘the potential for “hallucination” is built into the “Generative” part of GPT’, just as bias is built into its training.

Both problems can be reduced, the same authors warn, but they will be difficult to eliminate, because they are a direct consequence of the way these systems learn. We are not looking at a programming oversight that someone has not yet had time to correct.

And why does it invent more about some things than others? Marc Alier, lecturer and researcher at the Universitat Politècnica de Catalunya, puts it in the ODITE Report 2026 in the way most useful for a teaching staff: the quality of the answers is inconsistent, and especially so in areas where the training data are limited or incomplete. On what humanity has written about a thousand times, the ground is well prepared. On what almost nobody has written about, the model still has to say something, and it says it.

And this is precisely the problem. A car warns you when something fails, with a light on the dashboard. In a language model there is no warning light: no internal indicator comes on when the system moves from retrieving a well-established pattern to filling a gap. The two operations are the same operation. Ray Gallon, writer and educator, president and co-founder of The Transformation Society, sums it up in his interview in the same report, recalling that the machine predicts the next group of letters and that this is why it makes mistakes, ‘hallucinations, which are not hallucinations, in a psychological sense’ . The word chosen to name this phenomenon suggests an altered perception, something that happens to a healthy mind on a bad day. But in reality it has nothing to do with that.

No internal indicator warns when the machine transitions from retrieving solid data to inventing content.
al·lucinacions de la IA

Tone is not data

If the warning does not come from within, it will have to come from outside. And this is where the second difficulty appears, which is ours and not the machine’s.
When we read, we use form as a clue to substance. An orderly text, without hesitation, with precise vocabulary and a clear structure, seems more reliable to us than a hesitant one, and in dealings between people that clue works reasonably well, because those who doubt usually doubt out loud. Alejandro Espeso-García, author of a review on cognitive enhancement and offloading published in Cultura, Ciencia y Deporte, notes that the fluency and confident tone of these answers lead people to accept them on the premise that, if it reads well, it must be true. The clue is still there, but it has stopped telling us anything. Bowen and Watson say it in a single line: ‘authoritative-sounding text can still be gibberish’ .

It sounds like a warning for the careless, doesn’t it? But it turns out to work the other way round. The same authors report the experience of a German school documented by Haverkamp in 2022, where students were taught the flaws of the tool and then allowed to use it in exams. No student relied on it alone, and several looked for additional information on their own. However, the students who were less confident about the course content or about their own reasoning did worse, because they adopted the generated text without subjecting it to criticism.

That result deserves a pause. Those who most need the help are those least prepared to judge it, because to detect that a statement is false you have to know something about the subject. Verifying requires knowing.

Students discover this too when they are given the chance. Miquel Àngel Fuentes Arjona, a secondary teacher at a state school in Catalonia and a trainer in teachers’ digital competence, documents in the ODITE Report 2026 a twelve-session experience in the first year of Bachillerato (Year 12) in which his students reconstructed in writing what they thought before and what they think now. On hallucinations, they started from the idea that it was ‘a problem with the program’ and that the tool ‘always told the truth’. By the end, they speak of ‘credible but wrong answers’ and note that the system ‘is not capable of verifying information’. These data should be taken for what they are: thirteen students, a single term, and a retrospective collection which the author himself points out is not a pre-test and post-test design. It is not a measure of effectiveness. It is something more modest and more interesting: the voice of people who have changed their minds explaining why.

al·lucinacions de la IA
Students discover in the classroom that plausible answers often hide completely incorrect statements.

The same move that invents is the one that creates

At this point, the teaching staff usually ask when it will be solved. And the uncomfortable answer is that solving it completely would come at a price we may not want to pay.

Bowen and Watson devote a whole chapter of their book to creativity, and the argument is this: the ability to combine ideas and words as no human would have combined them is the same ability that produces invented references. Hallucinations, they write, ‘may be dangerous, but they are also a feature of creativity’. When inhibitions are added to a model so that it asserts fewer unchecked things, they note with regard to one particular tool, the inventions decrease and so does its creative interest. That is the real trade-off, and it is not schools that decide it.

What schools do decide is what they do in the meantime. If invention is not going to be eliminated, verification stops being a precaution added at the end and becomes half of the work. This is how the systematic review by Iliana Guaranda-Arias and her colleagues at the State University of Milagro, in Ecuador, frames it: it analyses fifty-two studies and places source verification, authorship and academic integrity within what they call ethical literacy in artificial intelligence, alongside teacher training in case analysis. The authors themselves warn that documentary and cross-sectional designs predominate and that causality cannot be established. It is a map of what to teach, not proof that teaching it works.

And there is one last twist, the hardest to see from the outside. Espeso-García observes that the greater risk is not that the tool invents facts, but that it produces in its users a hallucination of competence: the person confuses having obtained a coherent answer with having understood the matter. The machine has no warning light. Neither do we.

The primary educational risk is confusing the coherence of the generated text with actual comprehension.
al·lucinacions de la IA

What to do with this

In the classroom. Cristian Ruiz Reinales, head of Technology and computing teacher at Colegio Juan de Lanuza in Zaragoza, documents in the ODITE Report 2026 an activity for Year 6 of primary school called ‘Citation Detectives’: students check whether a quotation attributed to someone appears in reliable sources or only on pages with no stated author, and end up formulating their own rules, such as ‘if there is no source, then I do not accept it as valid’. The same project runs a cross-curricular routine that can be adopted tomorrow in any subject, ‘AI says / Source says / I conclude’, which requires students to separate in writing what the tool proposes, what can be checked and what they conclude using their own judgement.

For secondary and Bachillerato, Bowen and Watson propose a more demanding task: set a piece of work to be done with the help of the tool and then ask for all the citations to be verified, which must exist and say what the tool says they say, handing in the original references and the corrected ones together. The verification work thus becomes the object of assessment, not a prerequisite that nobody looks at.

For the leadership team. Bowen and Watson’s template for a school policy includes two points that are usually missing from the rules of use now being approved: an explicit warning about the limits of the tool and a clear rule about the student’s ultimate responsibility for what they hand in. The same authors warn that a policy built only on prohibitions is demotivating, misses opportunities and may increase inequity, because those who already know how to use the tool will use it anyway. The difference between prohibiting unverified citations and teaching how to verify seems small on paper. In practice it is the difference between a rule and a competence. Pau García-Milà, entrepreneur and science communicator, puts it in his interview in the report with a formula that works as a compass: ‘let them learn to distrust well’.

al·lucinacions de la IA
Chaining logical steps emulates reasoning, but it does not dispel the opacity of its internal mechanism.

To conclude

Schools have been teaching students to evaluate sources for decades, and that competence has not expired. What has changed is that unreliable text used to look unreliable, and now it never does. Verification is not a formality added to the use of these tools. It is half of the use.

Food for thought. Of the tasks we have set this school year, how many can be passed without checking a single statement? And a second one, for the team meeting: when a student hands in an invented citation, what exactly are we assessing, their honesty or their ability to check?

Fact-checking sources stops being an added chore and becomes the central pillar of evaluation.
al·lucinacions de la IA
Rate the post
5/5 - (5 votes)

Share

2026-09-22T13:54:34+00:00
Go to Top