Why an LLM Alone Is Not Enough to Produce New Knowledge

Discovery requires more than an answer

The fluency of a large language model can suggest that it is capable of conducting scientific discovery independently. Yet an autoregressive model operating without tools, observations or experiments has no mechanism for generating and testing new empirical evidence. It can formulate a hypothesis. Text generation alone cannot establish whether nature confirms it. Producing a description of an experiment is different from obtaining its results.

This distinction anchors the critique advanced by Felin and Holweg. They challenge the claim that probabilistic AI can independently generate genuine novelty and new knowledge, and consequently replace human decision making under uncertainty. This article develops their emphasis on the insufficiency of prediction as a standalone method of discovery. The broader rejection of every possibility of machine innovation remains a theoretical position, rather than an established mathematical impossibility theorem.

What determines the model’s answers

An autoregressive model produces successive tokens by assigning probabilities conditional on preceding text. Its output distribution is shaped by training data, architecture, optimization, the current prompt and the procedure used to select tokens. Training data are therefore a fundamental influence, but not the sole determinant. Information supplied in a prompt can also affect the response without changing the model’s parameters.

These operations do not themselves constitute an experiment, a measurement or an independent guarantee of truth. The probability of a continuation reflects its fit with learned patterns and context. It is not equivalent to the probability that a scientific assertion is correct. Asking the model to check its answer may improve it, but a second generated response does not automatically supply independent evidence.

None of this means that every output is copied. New combinations, useful hypotheses and correct deductions are possible. The distinction is between an original expression, a candidate solution and an established discovery. Mathematics also requires a qualification: new knowledge can emerge through proof without additional empirical observations. A generated proof must nevertheless be valid; its persuasive presentation cannot substitute for establishing that validity.

Beyond the input-output analogy

The familiar analogy between computers and minds as devices receiving inputs and producing outputs can obscure how questions themselves are formed. In the critique considered here, human cognition is organized around theories and directed towards future possibilities. Probabilistic AI is characterized as backward-looking and imitative because its predictions draw on recorded examples. These are theoretical characterizations, not a complete account of every human or machine capability.

The distinction does not make people infallible. It highlights a requirement of inquiry: deciding which observations are missing and which interventions could produce them. Describing cognition only as processing information already supplied leaves that active construction of an investigation insufficiently explained. Scientific work includes changing the conditions under which evidence becomes available, rather than merely interpreting a fixed record.

When hypotheses precede evidence

“Data-belief asymmetries” describe situations in which a conjecture precedes the evidence needed to test it. A proposed explanation motivates an investigation, the investigation leads to an intervention, and the intervention produces observations. The productive role of belief lies in making inquiry possible. It does not justify protecting a favored claim against results that contradict it.

Language acquisition supplies one illustration. The “poverty of the stimulus” argument asks how children acquire linguistic knowledge that appears to exceed the limited examples they have encountered. This remains a contested scientific argument, rather than a definitive refutation of statistical learning. It nevertheless challenges the assumption that producing grammatical sentences is sufficient evidence that human and machine learning operate through equivalent mechanisms.

Heavier-than-air flight supplies another illustration. The Wright brothers investigated problems of lift, propulsion and control, built experimental apparatus and collected measurements of their own. Their work did more than summarize earlier failures. A conjecture about a feasible machine helped organize a program that produced previously unavailable evidence. The relevant lesson is the connection between explanatory ideas, intervention and measurement, not a romantic celebration of ignoring evidence.

What model evaluations reveal

Empirical research identifies related limitations. Experiments found that familiar tasks and probable answers can improve language-model performance even when strict rules determine the solution. Reasoning optimization can mitigate this effect, as a subsequent preprint examining o1 demonstrated. Such improvements do not, by themselves, provide a channel for acquiring fresh measurements from the world.

Similarly, models trained on orbital trajectories could predict successfully without reliably applying Newtonian mechanics to new tasks. Another research program found that LLM proposals initially received higher novelty ratings than human proposals, but that this advantage diminished after implementation. These are findings from particular experiments. They support testing explanations and outcomes instead of treating an impressive response as sufficient evidence of discovery.

Discovery belongs to the wider system

FunSearch produced new mathematical constructions through a combination of a language model, program generation, external evaluation and iterative search. The model proposed candidates; checking and feedback enabled selection and improvement. The discovery resulted from these interacting functions, not from next-token prediction by itself. Recognizing the wider system also recognizes that candidate generation can be a substantive contribution.

This is the article’s central interpretation. It neither dismisses the model’s creative role nor proves that every future system will require human intervention at every step. It does reject the inference that success by an integrated research system establishes the scientific self-sufficiency of an isolated LLM. Those are different claims requiring different evidence.

Human judgment under uncertainty

Public policy and research choices frequently involve unknown consequences, competing objectives and decisions about which evidence to seek. Producing a probable answer does not adequately justify either the choice or the transfer of responsibility to a model. Decisions must remain open to scrutiny, including scrutiny of the assumptions used to frame the problem.

The practical direction we propose is AI embedded in open research infrastructure: documented datasets, inspectable software, reproducible tests and clear human accountability. These arrangements connect computational assistance to hypothesis testing and make failures visible. New knowledge is established through valid proof, empirical testing and the ability of others to challenge the result, rather than through the confidence with which a model states it.

Source of this article: glossapi.gr

Got Something To Say?