Why Hallucinated Citations Are Not an AI Failure
Why Hallucinated Citations Are Not an AI Failure
In the last few days there has been extensive discussion about reports of hallucinated citations in accepted NeurIPS papers. Many reactions frame these cases as a catastrophic failure of the publication system, blaming irresponsible authors or presenting them as proof that AI is undermining scientific fairness. I think this framing avoids to talk about something more deep.
First, the data being circulated are not statistically meaningful. We are often talking about a single paper with a small number of hallucinated citations, frequently referring to marginal or unused references. This is far too weak a signal to justify claims about the collapse of peer review or scientific rigor.
In that sense, these cases are not evidence of a systemic failure. However, the debate around them does provide an opportunity to discuss a more fundamental issue. Hallucinated citations are not evidence that AI is bad for science. They are better understood as a symptom of long standing structural problems in the research publication system that predate LLMs.
It is well known that blindly using LLMs to generate bibliographies is risky and should be avoided. But focusing only on this mistake misses the deeper issue. The problem is not that researchers have suddenly become careless because AI exists. The problem is what the system expects from them and the constraints under which they operate.
Especially at top tier conferences, researchers work under extreme time pressure. The system rewards speed, volume, and priority, leaving limited time for careful revision. In this environment, cutting corners is not an exception but a predictable outcome of the incentives.
Seen from this perspective, the use of AI is not evidence of individual irresponsibility. It reflects the kind of output the system has normalized. Many AI papers are not expected to be slow, deeply revised works developed over long periods. They are treated as vehicles to publish the next idea as quickly as possible.
Against this background, a few hallucinated citations are among the least dangerous failure modes we should expect. More concerning are rushed evaluations, weak experimental validation, overclaimed results, or irreproducible findings. These issues existed long before LLMs and were not created by AI.
The lesson, then, is not that AI is bad for science. Hallucinated citations are simply another behavior that current incentives already reward.
This leaves two options. If we are genuinely concerned about the effects of expanding LLM use in research, we must rethink the incentives of AI research itself, including what we reward, what we carefully review, and what kinds of contributions we value.
If we are unwilling to question an over productive system optimized for speed and output, then there is little reason to overstate the risks of AI. In that case, hallucinated citations are not an AI specific failure but a predictable side effect of cost cutting and productivity driven research.
Under these constraints, using AI may even be one of the safer available options. It standardizes shortcuts researchers were already forced to take rather than introducing entirely new risks. The real danger lies not in the tools themselves, but in pretending their failures are exceptional instead of acknowledging that they reflect the incentives shaping modern research.
Enjoy Reading This Article?
Here are some more articles you might like to read next: