Intelligence begins with a concession the universe makes to finite creatures: not everything that could happen does happen, and not everything that happens is independent of what happened before.
A mind cannot contain the world. It is made from a small part of the world, operates for a limited time, and encounters only fragments of what exists. Nevertheless, it can anticipate an approaching object, recognize a face in unfamiliar light, infer a mechanism from an experiment, or discover an intervention that changes the course of a disease. It succeeds not by reproducing reality in full, but by finding descriptions that preserve some of reality’s consequential relationships.
The central question is therefore not simply how matter becomes intelligent. It is how a finite physical system acquires an internal organization that allows it to exploit the organization of something larger than itself.
Three ideas provide a route through this question: compression, prediction, and control. Compression discovers what can be represented economically. Prediction uses that representation to infer what has not yet been observed. Control uses those inferences to influence what happens next. In biology, however, the relationship is more intricate than this sequence suggests. Living systems do not merely discover regularities. Through feedback, they also help create and maintain them.
That additional step changes both how we should understand intelligence and how we should attempt to treat disease.
A world small enough to learn
Consider the space of all possible images. Almost every arbitrary arrangement of pixels would be meaningless to us. Faces, forests, and rooms occupy highly constrained regions within this enormous space. A change in illumination does not independently rearrange every pixel of a face; it produces coordinated changes determined by shape, surface properties, and the direction of the light. The manifold hypothesis expresses one version of this idea: observations in a high-dimensional space may lie near a much lower-dimensional structure. It is a useful hypothesis with formal formulations, not a theorem that all meaningful data occupy one smooth surface.[1]
For biology, this qualification matters. There need not be a single, fixed manifold of life. Different cell types, developmental stages, environments, and interventions may require different local descriptions. Discrete changes, continuous variation, noise, and history can coexist. Nor should every constraint be interpreted as geometric. A logical dependency, a conservation law, and a causal mechanism can each make data learnable without reducing neatly to the same mathematical object.
The deeper proposition is that the world contains dependencies that make some observations informative about others. For discrete variables, mutual information expresses this as
Knowing something about X reduces uncertainty about Y. When the variables are the past and future of a process, this becomes predictive information: the information in history that can help constrain what comes next.[2]
A biased coin clarifies the distinction. Learning that it lands heads 99 percent of the time improves our predictions relative to assuming fairness. But, if successive tosses are independent, observing the last toss tells us nothing additional about the next. Nonuniformity is useful, but it is not the same as rich temporal structure. Likewise, an independent noise process can have a learnable distribution even when its next realization cannot be predicted from its history.
Intelligence requires neither a perfectly deterministic universe nor universal predictability. It requires regularities that are accessible, sufficiently persistent, and relevant to some task. A perfectly repetitive world would be easy to compress but offer little need for flexible intelligence. An entirely unstructured one would offer no foothold for learning. The interesting territory lies between them: enough order to support inference, enough variation to make inference valuable.
Learning what to forget
Learning and compression have a precise connection. A predictive model that assigns high probability to a sequence can be used to encode that sequence economically. In a minimum-description-length view, however, the accounting must include both the model and the data left unexplained by it. A table containing every observation is not a profound explanation simply because it reproduces its entries exactly. The achievement is to find a reusable description that makes additional observations less surprising.[3][4]
This is the attraction of the shortest-program intuition associated with Kolmogorov complexity. A compact generative description can explain many apparently separate facts. But short description length is not, by itself, intelligence. The description must be discoverable, and useful consequences must be computable within the resources available. A tiny program that takes longer than an organism’s lifetime to answer a question offers little assistance to that organism.
Nor is indiscriminate compression desirable. A representation can become smaller by forgetting precisely the distinction on which a future decision depends. The information bottleneck makes the task-dependence explicit: compress the input while retaining information about a specified relevant variable. It is not an instruction to forget as much as possible. It is an instruction to forget selectively.[5]
Imagine two patients whose measurements look almost identical at baseline, but whose responses to the same treatment differ sharply. An embedding that merges them may be excellent for reconstructing common laboratory values and disastrous for selecting therapy. The distinguishing feature may explain very little variation in the dataset while explaining a great deal of variation in the consequence of an intervention.
This is why the familiar demand to “retain the signal and remove the noise” is incomplete. Signal for what? Noise relative to which decision? A rare event may be statistically negligible and medically decisive. A feature irrelevant to tomorrow’s prediction may be indispensable to a prediction over ten years.
A useful representation is therefore a commitment about future questions. To choose what a model preserves is, implicitly, to choose the kinds of ignorance it will be allowed to have.
How regularity becomes intelligence
The existence of learnable structure does not explain how a system comes to exploit it. There must also be a process that retains successful ways of responding to the world and modifies unsuccessful ones.
Evolution supplies one such process. When environmental cues can guide beneficial responses, the machinery that detects and uses those cues can acquire a fitness advantage. Formal models of the fitness value of information connect this advantage to both the informativeness of the cue and the consequences of acting on it. Information is not valuable merely because it is information; its value depends on the available responses and the environment in which they are used.[6]
This suggests a progression, rather than a magical boundary. An inherited response can exploit a regularity without learning it during an individual lifetime. A plastic system can adjust its responses through experience. A system with memory can distinguish situations that look identical now but have different histories. A system that evaluates alternative action sequences can respond to circumstances it has never encountered in exactly that form.
These are not necessarily steps on a single evolutionary ladder. They are increasingly demanding ways of coupling information to action.
In artificial learning, optimization can similarly produce reusable internal structure. Yet a randomly initialized network is not wholly without structure: its architecture, inputs, and training procedure already constrain what it can readily learn. Training does not guarantee that it discovers the true causes of its observations. It may acquire useful abstractions, shortcuts, memorized exceptions, or some mixture. In a particularly revealing small-scale case, mechanistic analysis of transformers learning modular arithmetic identified a generalizing algorithm built from Fourier-related computations, rather than merely a collection of remembered answers.[7]
Such results make emergence less mysterious without making it trivial. Some apparent capability jumps arise from discontinuous scoring rules: gradual improvements can abruptly cross an exact-answer threshold. Other changes involve the development and increasing use of structured computational mechanisms. The scientific question is not whether a model has crossed a rhetorically satisfying boundary into “understanding,” but what internal organization now supports behavior that was previously unavailable.[8][7]
Scale enlarges the space of structures a model might acquire, but capacity is only a possibility. The data, learning objective, optimization process, and available computation determine which possibilities become usable capabilities.
Memory, analogy, generalization, and reasoning can then be understood as related uses of representations, without pretending they are identical operations. Memory preserves useful distinctions. Generalization carries a relation into a new case. Analogy maps relational structure between settings. Reasoning performs transformations whose intermediate results need not have appeared in the original observations.
A world model need not be a literal simulation sitting inside the agent. It can be implicit in a policy or distributed across computational mechanisms. Under specified controlled-Markov-process assumptions, Richens and colleagues show that sufficiently general competence across multistep goals entails predictive information about the environment that can be extracted from the policy. This is a formal result about a defined class of agents and tasks, not a proof that every intelligent behavior requires a complete internal replica of reality.[9]
The resulting picture is an account of adaptive competence, not a completed theory of consciousness. A thermostat closes a feedback loop; that alone does not make it a mind. Flexible intelligence concerns the breadth, adaptability, and compositional power with which such loops can be organized.
Before brains, there was regulation
The connection between information and control is already visible in systems without nervous tissue.
Bacterial chemotaxis provides a particularly clear example. A bacterium does not need a pictorial map of its surroundings to adjust movement in response to chemical conditions. Its signaling machinery can retain a biochemical trace of recent stimulation. Work on robust adaptation in bacterial chemotaxis identified integral feedback as a mechanism through which a signaling output can return to its previous level despite a sustained change in input.[10]
Integral feedback is a particularly clear case, not a claim that every homeostatic circuit implements exact integration or achieves perfect adaptation.
The important point is not that the bacterium secretly performs conscious arithmetic. It is that chemical dynamics can implement a mathematical operation. A variable that accumulates a discrepancy over time becomes a physical memory of that discrepancy. The computation exists in the organization of the reactions, not in a separate symbolic description of them.
This is more than an analogy imposed after the fact. Synthetic-biology experiments have constructed biomolecular integral-feedback controllers in living cells and demonstrated their regulation and adaptation properties. Control-theoretic organization can be engineered into chemistry.[11]
The embodiment matters. Information is never processed by an abstract diagram alone. It must be stored and transformed in physical machinery. Landauer’s principle and experimental tests of it establish a precise connection between logically irreversible erasure and thermodynamic cost under specified conditions. They do not establish that every inference has the same minimum energy cost, or that Shannon entropy, thermodynamic entropy, and the “free energy” of an inference objective are interchangeable quantities.[12]
We do not need to blur those distinctions to see the deeper continuity. A biochemical controller, a nervous system, and an artificial agent can all preserve selected consequences of the past and use them to shape subsequent behavior. Their implementations and capabilities differ enormously. Their shared problem is to make limited internal state useful in a changing external world.
The manifold is partly something life maintains
Here the argument turns back on its starting point.
We began by saying that intelligence is possible because observations occupy a constrained portion of the space of possibilities. In living systems, some of those constraints are actively maintained. If a feedback controller keeps an output near a reference value despite disturbances, then observations of that output will cluster near the reference. The regularity is not merely an external fact waiting to be discovered. It is partly the result of regulation.[11]
This has a troublesome consequence for biological learning. A variable may vary little because it matters greatly.
Suppose a controller continually corrects deviations in an essential physiological quantity. A dataset collected during successful regulation will contain little variation in that quantity, even while the hidden effort required to maintain it changes substantially. A representation optimized only to capture the largest observed variations might neglect the regulated output, the compensatory machinery, or both.
The resulting observational simplicity can conceal causal importance. To see what the system depends on, we may need to examine its response to disturbance rather than its appearance at rest. The observational manifold is conditional on the existing regulatory regime. An intervention can alter that regime, exposing directions of variation that were nearly absent during training.
This also separates three objects that are too easily conflated. The geometry of observed states describes where a system has been found. Its dynamics describe how it moves. Its control structure describes how available interventions can alter that movement. None can simply be substituted for the others.
A low-dimensional embedding is not automatically a causal model. A causal model is not automatically sufficient for safe control. Even a two-dimensional system cannot be fully steered by an input that moves only one coordinate if the dynamics never transmit that influence to the other. Conversely, a system with vast microscopic complexity may permit useful regulation through a small number of aggregate variables.
The number of variables needed to describe a dataset, the number of hidden states needed to predict a response, and the number of intervention points needed to change an outcome are different quantities.
In control-theoretic language, we must ask about both observability—whether relevant hidden states can be inferred from measurements—and reachability—whether the available inputs can move the system toward a desired outcome under its constraints.
An atlas of biology tells us which states have been observed. Medicine requires an additional answer: which desirable states are reachable, by what paths, and at what cost?
Why pushing biology often produces a push back
A deliberately simple model makes the problem concrete. Let y be a regulated biological output, c an unobserved compensatory drive, and u a treatment that suppresses the output:
where a and k are positive and y⋆ is the controller’s reference level. This is an illustrative local model, not a claim that a particular disease obeys these equations.
Increase u, and initially y falls. But that fall creates a discrepancy relative to y⋆. The compensatory state c accumulates until it cancels the sustained suppression. At equilibrium,
The observed output has returned to its previous value. The internal state has not.
A terminal measurement of y alone might suggest that the treatment did nothing. A longitudinal record would reveal an initial response followed by recovery. Measuring c would show the compensation. Withdrawing treatment from the adapted state would initially drive the output upward because the compensatory drive would still be elevated.
Within this model, increasing the dose does not produce a different sustained output while integral compensation remains operative and unsaturated. It produces a different amount of hidden compensation. Merely lowering the feedback gain changes the speed of adaptation, not the final reference level. Durable displacement requires changing something more fundamental: the reference, the controller architecture, the available compensatory capacity, or the relationship between that controller and the output we care about.
This gives a precise form to the driver-and-constraint view of therapy. One intervention moves the system in a desirable direction; another prevents a specified compensatory or self-reinforcing process from undoing the change.
But it is a conditional design principle, not a universal two-drug theorem. A single intervention may alter the reference itself, remove an indispensable driver, or disable several compensatory routes at once. Two interventions may be redundant or insufficient. Nor is all disease a stable attractor: irreversible tissue loss, persistent external injury, and progressive failure require different descriptions. The therapeutic task is to establish which dynamics actually apply, rather than assume that every pathology is an equilibrium waiting to be shifted.
Lipid lowering supplies a concrete example of the value of understanding compensation. Atorvastatin can increase circulating PCSK9, which promotes degradation of hepatic LDL receptors. PCSK9 inhibition provides a complementary intervention; in FOURIER, adding evolocumab to background statin therapy reduced LDL cholesterol and cardiovascular events. This supports the usefulness of complementary control points, without demonstrating that all therapeutic combinations work by the same feedback mechanism.[13][14]
Most importantly, the purpose of a second intervention should not be to abolish homeostasis indiscriminately. The same regulatory machinery may protect other functions. A successful treatment must distinguish pathological compensation from the robustness that keeps the patient alive.
A model should preserve the consequences of intervention
These considerations change what “understanding biology” should mean for a model built to support medicine.
We do not necessarily need a miniature biochemical replica of the patient. We need a representation that preserves the distinctions required to choose between relevant actions. Work on control-oriented representation learning makes an analogous distinction: states can be grouped according to their consequences for rewards and future transitions, rather than according to whether they permit detailed reconstruction of the original observations.[15]
For medicine, the requirement must be stronger than observational resemblance. Let hₜ denote the history of measurements and treatments, and let zₜ = φ(hₜ) be a compressed representation. An appropriate aspiration is
for the intervention sequences u, outcomes, and time horizons relevant to the decision. Here do(u) means imposing those interventions, rather than merely observing that they occurred.
This is a proposed criterion for a useful medical representation, not a guarantee that a dataset allows us to learn one. Two mechanisms can agree on everything observed so far and disagree on the response to an untested treatment. No amount of compression can manufacture the missing causal evidence.
The criterion also explains why history belongs in the state. In the illustrative feedback system, the same current output can conceal different levels of compensation. A snapshot that merges those conditions loses information needed to predict withdrawal or rechallenge. Where the internal state remains uncertain, the representation may need to preserve a distribution over possibilities rather than one confident estimate.
A model can be mechanistically incomplete and still be useful for a bounded control problem. If it reliably predicts the consequences of modest interventions in a validated operating region, repeated measurement and replanning may compensate for some of its omissions. But feedback is not a license for arbitrary model error. Delays, unobserved hazards, and irreversible damage may prevent correction before harm occurs.
The appropriate ambition is therefore neither exhaustive realism nor convenient simplification. It is decision-relevant fidelity, with an explicit boundary of validity.
A model that knows when its representation is no longer adequate may be more valuable than one that reconstructs familiar data beautifully and extrapolates without hesitation.
An intervention can also be a question
Once a model is understood as incomplete, action acquires another purpose. An intervention changes the system, but it can also reveal which system we are dealing with.
This is dual control in its technical sense: actions have both consequences for the controlled process and consequences for what will be learned about its uncertain dynamics. It is distinct from dual intervention. Two drugs concern the number and roles of intervention points. Dual control concerns the interaction between acting and learning, even when only one intervention is available.[16]
Consider a proposed research program in fibrosis. Experimental work has shown that matrix stiffening can promote fibroblast activation and matrix production, creating a reinforcing relationship between cells and their mechanical environment. Liu and colleagues demonstrated this feedback using lung-injury and cultured-fibroblast experiments. It provides a mechanistic foothold, not proof that any particular two-intervention regimen will reverse human fibrosis.[17]
Now imagine two competing models of a patient-relevant experimental system. In one, continuing injury supplies the dominant drive. In the other, the altered matrix has become sufficiently self-reinforcing that the original injury is no longer necessary to sustain activity. Both models might explain the same baseline expression profile. They could make different predictions about the speed of response after reducing injury, the persistence of fibroblast activation, and the effects of changing the mechanical environment.
An informative experimental perturbation would be selected partly because these predicted trajectories diverge. Measurements would follow the early response, adaptation, recovery, and, where appropriate, withdrawal. The experiment would ask not only whether a marker decreases, but which explanation survives the pattern of change.
More time points do not automatically identify the mechanism. Sampling can miss the relevant timescale; hidden variables can leave several explanations indistinguishable; an intervention can have unrecognized effects. The remedy is not “longitudinal data” as a slogan. It is the joint design of perturbations, observations, and competing models.
Nor is information gain automatically a sufficient reason to perturb a patient. Clinical use requires ethical justification, safety constraints, and an appropriate evidentiary setting. Much exploratory learning belongs first in carefully chosen experimental systems. Even there, the most useful information is information that changes a subsequent decision, not merely information that makes a model more elaborate.
The central principle remains: the best next action need not be the one that looks best under our current favorite explanation. It may be the action that improves the system while making future decisions less dependent on an explanation that could be wrong.
A world model of medicine needs a curriculum of consequences
This leads to a different vision of a medical world model.
It should not be judged primarily by how much biomedical information it can store, or by whether every modality can be placed in a common embedding. Its purpose is to preserve useful relationships between observations, latent states, interventions, and outcomes across the settings in which we expect to use it.
That expectation must be made concrete through tasks. What must the model predict within a cell? Which perturbation effects must it carry from cells into tissues? Which tissue changes must it connect to organ function? Which links can support a prediction about clinical benefit, and which remain conjectural?
A sensible curriculum would include within-scale dynamics, adjacent-scale links, and selected longer-range predictions. It would also test whether independently useful pieces remain reliable when composed. A cell-state model and an organ-function model do not become a validated cell-to-organ model merely because their embeddings can be aligned.
Population generalization is a separate challenge from physical scale. Predicting from cellular measurements to an individual’s organ response is one task. Establishing that the relationship transports across patients with different histories, exposures, and comorbidities is another. Moving from a cell to a tissue and moving from one person to a population should not be collapsed into the same notion of “going up a level.”
The tasks must also expose the distinctions that observational compression would otherwise erase. They should challenge the model with timing, treatment order, adaptation, rebound, combinations, and recovery. A representation that predicts untreated trajectories but fails after intervention has learned something real, but not yet the thing required for therapeutic control.
This is a proposal for how to build and assess such models, not a claim that one universal representation must exist. Different horizons and intervention classes may require different state descriptions. The useful architecture may be a collection of overlapping, partially compatible models, with explicit statements about when their conclusions can be combined.
The aspiration is not omniscience. It is a growing domain of reliable counterfactual competence: knowing what is likely to happen if we act, recognizing when the evidence is insufficient, and choosing observations that can reduce the uncertainty that matters.
Intelligence closes the loop
Read together, the traditions of Shannon and Wiener suggest a larger picture. Information theory asks what can be distinguished, encoded, and inferred. Cybernetics asks what those distinctions can accomplish when they are placed inside a feedback loop.
A representation earns its significance through what it enables. It can allow an organism to preserve a viable condition, a learner to recognize a new instance of a familiar relation, or a scientist to choose an experiment that separates explanations. Information becomes consequential when differences in the model produce appropriate differences in action.
Biology adds a final complication: the object of our intervention is already full of regulatory activity. We are not usually acting on passive matter that will remain wherever it is pushed. We are entering processes that preserve, compensate, repair, remember, and sometimes sustain pathological conditions. Effective medicine must learn the organization of those responses, not merely catalog the components through which they occur.
This is why compression, prediction, and control should be understood as a loop rather than a ladder. What we want to control determines what we must predict. What we must predict determines what we should preserve. What we preserve determines which actions appear possible. And the actions we take change both the world and the evidence from which the next representation will be learned.
The deepest connection is therefore not simply that intelligence uses models and medicine needs better models. It is that both confront the same problem: how to act effectively in a system whose full complexity exceeds what can be represented.
Intelligence does not require possessing the whole world. It requires learning which differences matter, preserving them in a usable form, and discovering when that form has become inadequate.
For medicine, the corresponding ambition is not to know every molecular detail before intervening. It is to know enough of the right relationships to change a trajectory—and enough about the limits of that knowledge to learn from what happens next.
References
[1] Fefferman, C., Mitter, S., and Narayanan, H. (2016). “Testing the Manifold Hypothesis.” Journal of the American Mathematical Society. Author manuscript: arXiv:1310.0425.
[2] Bialek, W., Nemenman, I., and Tishby, N. (2001). “Predictability, Complexity, and Learning.” Neural Computation, 13, 2409–2463. Author manuscript: arXiv:physics/0007070.
[3] Rissanen, J. (1978). “Modeling by shortest data description.” Automatica, 14(5), 465–471. Publication record: IBM Research, “Modeling by shortest data description.”
[4] Delétang, G., et al. (2023). “Language Modeling Is Compression.” arXiv:2309.10668.
[5] Tishby, N., Pereira, F. C., and Bialek, W. (1999; preprint posted 2000). “The information bottleneck method.” arXiv:physics/0004057.
[6] Donaldson-Matasci, M. C., Bergstrom, C. T., and Lachmann, M. (2010). “The fitness value of information.” Oikos, 119(2), 219–230. DOI: 10.1111/j.1600-0706.2009.17781.x.
[7] Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J. (2023). “Progress measures for grokking via mechanistic interpretability.” arXiv:2301.05217.
[8] Schaeffer, R., Miranda, B., and Koyejo, S. (2023). “Are Emergent Abilities of Large Language Models a Mirage?” arXiv:2304.15004.
[9] Richens, J., Abel, D., Bellot, A., and Everitt, T. (2025). “General agents contain world models.” arXiv:2506.01622, version 5. Earlier versions were titled “General agents need world models.”
[10] Yi, T.-M., Huang, Y., Simon, M. I., and Doyle, J. (2000). “Robust perfect adaptation in bacterial chemotaxis through integral feedback control.” Proceedings of the National Academy of Sciences, 97(9), 4649–4653. DOI: 10.1073/pnas.97.9.4649.
[11] Aoki, S. K., et al. (2019). “A universal biomolecular integral feedback controller for robust perfect adaptation.” Nature, 570, 533–537. DOI: 10.1038/s41586-019-1321-1.
[12] Bérut, A., et al. (2012). “Experimental verification of Landauer’s principle linking information and thermodynamics.” Nature, 483, 187–189. DOI: 10.1038/nature10872.
[13] Careskey, H. E., et al. (2008). “Atorvastatin increases human serum levels of proprotein convertase subtilisin/kexin type 9.” Journal of Lipid Research, 49(2), 394–398. PMID: 18033751.
[14] Sabatine, M. S., et al. (2017). “Evolocumab and Clinical Outcomes in Patients with Cardiovascular Disease.” New England Journal of Medicine, 376, 1713–1722. DOI: 10.1056/NEJMoa1615664.
[15] Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S. (2021). “Learning Invariant Representations for Reinforcement Learning without Reconstruction.” ICLR. Author manuscript: arXiv:2006.10742.
[16] Klenske, E. D., and Hennig, P. (2016). “Dual Control for Approximate Bayesian Reinforcement Learning.” Journal of Machine Learning Research, 17(127), 1–30.
[17] Liu, F., et al. (2010). “Feedback amplification of fibrosis through matrix stiffening and COX-2 suppression.” Journal of Cell Biology, 190(4), 693–706. DOI: 10.1083/jcb.201004082.



I enjoyed your thought-provoking post! We think a lot about these topics in our work testing human tumor tissue ex vivo and training models to predict response to multiple therapeutics. You touched on several things we’ve also found important in experimental and model design, particularly the value of perturbations and understanding response trajectories over time.
I agree that general-purpose foundation models have limits, since there is no single, task-independent “best” representation of a biological system. We see this empirically: our models generalize better to unseen patients when we train the representation and treatment-response predictor end-to-end.
One of the biggest challenges I see with clinical datasets is that a patient can only receive one therapy at a time, so responses to alternative treatments (at the same initial state) are never observed... This makes it difficult to disentangle patient biology from treatment effects and learn the patient-by-treatment interactions that ultimately matter for treatment selection. Experimentally perturbing the same patient’s tumor with multiple therapies gives us a very different kind of training signal.
If compensatory capacity is latent rather than directly observable, how would you operationalize its measurement in practice? Would it require repeated perturbations to identify the hidden state, rather than a single response–recovery trajectory?