QuarterlyFounded 2026ISSN pending

Analysis

Artificial Intelligence Cannot Correct a Misdefined Biological Question

Pattern recognition can improve an answer, but it cannot repair a biological premise that was wrong at the start.

Analysis13 September 20265 min readBomi Joseph, MD, PhD

Artificial intelligence begins after someone has decided what the question is, which data represent it, and what outcome counts as correct.

Those decisions are biological and conceptual before they are computational. If they are wrong, the model does not stand outside the error and repair it. It learns within the error, optimizes it, and can reproduce it at enormous speed.

This is the central limitation of AI in human science: computational power cannot compensate for a misdefined biological construct.

Better computation × wrong premise = more efficient errorA model can optimize the relationship it is given; it cannot validate the biological meaning that was never established.

The model inherits the question

Every model contains a chain of prior choices. Investigators choose the population, signals, labels, reference standards, time windows, exclusions, and target outcome. They decide whether a temperature change represents hormonal transition, whether movement represents wakefulness, whether a proxy score represents health, or whether an administrative label represents disease.

The algorithm may discover patterns among these variables. But the meaning of those patterns remains bounded by the validity of the original choices. If the label is an imperfect proxy, the model learns the proxy. If the sensor detects several overlapping processes, the model learns their mixture. If the outcome is defined incorrectly, improved prediction only increases confidence in the wrong target.

AI cannot change biology. It can only process the representation of biology that we give it.

Pattern recognition is not physiological understanding

A model can associate input A with outcome B without measuring the mechanism that connects them. This can be useful for prediction, but prediction is not the same as explanation, direct measurement, or causation.

The distinction matters whenever a technology makes a biological claim. If a person’s self-reported event supplies the label, the person may be the true sensor while the device learns nearby patterns. If a score is trained against another unvalidated score, the second system inherits the first system’s ambiguity. If multiple physiological states produce similar signals, classification accuracy in a narrow dataset may collapse when the context changes.

Adding more variables does not automatically resolve this problem. A larger collection of proxies can produce a stronger correlation while remaining detached from the construct the system claims to measure.

The inference stack hides the weakness

Human technologies often contain several layers between the body and the claim presented to the user. A receptor captures a physical signal. Processing removes artifact and transforms the waveform. Features are extracted. A model generates a probability. A product layer converts that probability into a label, score, warning, or recommendation.

Each layer can be technically elegant. Yet every layer also creates an opportunity for uncertainty to be disguised as certainty.

The chain must be tested at every level:

  1. Biological construct: Is the phenomenon defined precisely?
  2. Receptor: Is the signal captured from the correct physiological location and compartment?
  3. Reference: Was it compared with an independent, valid measure of the phenomenon?
  4. Model: Does performance persist across people, conditions, and time?
  5. Interpretation: Is the displayed claim limited to what the evidence directly supports?
  6. Decision: Does acting on the output improve the intended human outcome?

Testing only the final prediction conceals where the system succeeds and where it fails. A global performance number cannot show whether the receptor was biologically appropriate, whether the reference was valid, or whether the output remains reliable in an individual outside the development population.

A wrong label can look impressively accurate

AI performance is commonly reported against the labels used for training or evaluation. This answers an important computational question: how well does the model reproduce those labels? It does not necessarily answer the biological question: were the labels true?

Suppose a system is trained to infer a physiological state from a convenient signal, while its reference labels come from user entries, billing codes, a weak questionnaire, or another indirect device. The model may reproduce those labels with high accuracy. What has been validated is agreement with the label source—not direct measurement of the claimed physiology.

This is not a minor technicality. The apparent authority of AI can cause users to treat inference as observation. The score arrives with decimals, trends, colors, and personalized language. Presentation precision is then mistaken for biological precision.

Establish truth before scaling inference

The proper sequence is physiology first. Define the biological phenomenon. Observe it with an appropriate reference standard. Determine which signals change faithfully with it. Validate those signals under controlled perturbation and in representative humans. Only then should AI be used to identify patterns, improve prediction, or extend the measurement into practical settings.

When that foundation is sound, AI can be extraordinarily useful. It can integrate signals across time, reveal interactions, detect subtle departures from an individual baseline, and help humans interpret complex systems. But its usefulness is conditional on the quality of the construct, measurements, labels, and validation beneath it.

The answer to weak biological foundations is not a larger model. It is better physiology.

Demand the uncompressed explanation

Before accepting an AI-derived human claim, ask what was directly measured, what was inferred, how the target was defined, who supplied the labels, which independent reference was used, and what happened when the physiology changed. Ask which populations were missing, which conditions caused failure, and whether the displayed conclusion exceeds the evidence.

If those questions cannot be answered without retreating into model complexity, the complexity is functioning as camouflage.

AI can sharpen a well-defined question. It can find structure that conventional analysis misses. It can make validated knowledge more useful. What it cannot do is reach backward through the data pipeline and correct a biological premise that was wrong at the beginning.

Author note. This analysis was written by Bomi Joseph, MD, PhD, and editorially reviewed by JHST. It was not sent for external peer review. No external funding was received for this article, and no conflicts of interest were declared.