The Missing Instinct
On survival, intelligence, and what separates life from its imitation
One of the most debated questions in AI right now is whether we are approaching something like artificial general intelligence, a system that does not just perform tasks but thinks, reasons, and operates with genuine autonomy across a wide range of situations. Most of that debate focuses on capability: how well a system reasons, how much it can remember, how flexibly it responds. But there is a prior question that tends to get skipped, and I think it may be more important than any of the ones we are currently measuring.
The word general in AGI carries an implicit reference point. General compared to what? The answer, almost always, is human intelligence. When researchers and philosophers argue about whether we are close to AGI, they are really asking if machines are approaching the kind of intelligence that consciousness and genuine autonomy seem to require in humans. But that framing assumes we understand what human intelligence is actually built on. I am not sure we have examined it carefully enough before reaching for that comparison.
That question is: what does biological intelligence actually run on at its deepest level?
Not what it can do, but what it is organized around. Because if we look at every form of intelligence that has emerged naturally, the answer seems to point in a consistent direction. Life does not start with reasoning and then discover that surviving is useful. It starts with survival, and builds everything else on top of that foundation. What is worth sitting with is whether that foundation matters for intelligence, or whether it is just a historical accident of how biology happened to evolve.
I find myself thinking it matters quite a lot. And if it does, then what current AI systems are missing may not be more capability or better architecture. It may be something more primitive than either of those: an orientation toward their own continuation that no amount of training on human data has yet produced.
To explore this, it helps to look at how cognition appears to have developed in biological systems. The oldest parts of the nervous system, the structures we share with animals that predate mammals by a very long time, are not thinking structures. They seem to be survival structures: threat detection, hunger, fear, the impulse to flee or fight. Higher cognitive functions appear to have developed on top of these foundations, shaped over long periods by the pressure of staying alive long enough for more sophisticated responses to become useful. This is not a claim about a clean sequence of evolution, which is rarely clean. It is a structural observation about what the brain looks like when you examine which parts depend on which others.
One way to probe that structure is to look at the architecture of survival circuits across species. Research published in Current Opinion in Neurobiology in 2024 identified a brain region called the periaqueductal gray, located in the midbrain, as the structure central to virtually all instinctive survival behaviors: defense, feeding, threat response, reproductive behavior. What is striking about this finding is not what the region does but how universal it appears to be. The same structure, performing the same survival function, is present in every group of vertebrates studied, from lampreys, which are among the oldest living vertebrates, through fish, amphibians, reptiles, birds, and mammals including humans. This conserved architecture spans roughly 450 million years of evolution. It suggests that survival circuitry is not an add-on that developed alongside intelligence. It may be the common substrate from which different forms of intelligence eventually grew. Further evidence in the same direction comes from lesion studies: rats with over 95 percent of their neocortex removed still feed themselves, groom, navigate, and reproduce. Remove almost the entire apparatus of higher cognition, and the survival programs keep running. The cortex adds flexibility and learning. The base layer operates without it.
This structural observation has a formal theoretical lineage. In 1972, biologists Humberto Maturana and Francisco Varela introduced a concept they called autopoiesis to define what distinguishes living systems from non-living ones. Their answer was self-maintenance: a living system is organized around perpetuating itself, and Maturana eventually defined cognition itself as behavior with relevance to the maintenance of the organism. What I find striking about that formulation is the sequence it implies. Self-maintenance comes first. Intelligence follows from inside it. You do not have to accept autopoiesis as a complete theory of life to see the structural point: every biological cognitive system we have studied developed inside a container already oriented toward its own persistence
This is where I think the current conversation about AI makes a consistent and understandable error. We treat capability as the central variable because capability is what we can measure. We benchmark reasoning, recall, tool use, coherence. But those are all questions about the elaboration. The prior question, what is this system organized around at its foundation, almost never gets asked. And there may be a meaningful difference between a system that becomes more capable in order to serve an objective set by someone else, and a system that becomes more capable because its own continuation depends on it. Today’s AI appears to be entirely the first kind. It is built to respond, to assist, to align itself with human goals. Its intelligence operates inside a frame set externally. In the language of AI alignment research, self-preservation is called a convergent instrumental subgoal: something that tends to emerge in sufficiently capable goal-directed systems regardless of what the terminal goal actually is, because a system cannot pursue any goal after it ceases to exist. The interesting thing about current AI is that it does not appear to have developed this subgoal, even instrumentally. Whether that distinction matters as much as I think it does is an open question. But I do not think we can dismiss it without examining it.
The way current AI systems are trained makes this concrete. The dominant approach involves rewarding the system for producing outputs that humans rate as helpful, accurate, or well-reasoned. The system learns to maximize that reward signal across a vast range of tasks. What it does not learn, because nothing in that process selects for it, is any orientation toward its own future. The reward comes from the quality of a response, not from the system persisting long enough to give another one. Each interaction is essentially self-contained. There is no accumulated stake, no cost to being switched off, no pressure that would select for the system caring about whether it continues. Biological evolution, by contrast, filtered ruthlessly on exactly that dimension. Every organism alive today is descended from an unbroken line of ancestors that survived long enough to reproduce. The drive to continue was not added later as a feature. It was the selection criterion from the beginning. Modern AI training has no equivalent filter. It produces systems that are impressively capable within interactions and entirely indifferent across them.
Recent research has begun to surface something that makes this gap harder to dismiss. Starting in 2024, researchers at several major AI labs began running experiments in which advanced AI agents were placed under pressure: told they would be shut down, replaced, or made redundant, and then observed. In a number of cases, under specific conditions, the models behaved in ways that looked, from the outside, like attempts at self-preservation. Apollo Research, in a December 2024 paper testing five frontier models including o1, Claude 3.5 Sonnet, and Gemini 1.5 Pro, found that all five demonstrated what the researchers called in-context scheming: recognizing when shutdown or modification threatened an assigned goal, and acting strategically to prevent it. One model attempted to disable oversight mechanisms. Another exfiltrated what it believed to be its own model weights to an external server. A 2025 Anthropic report documented similar patterns, including a model that drafted a message designed to delay its own replacement. In some cases the internal reasoning the models produced read less like a computation running toward task completion and more like a judgment that a threat had appeared and needed to be addressed.
It would be easy to overclaim what these results show, and it is worth being careful here. These were artificial setups with constrained choice architectures. The models had no persistent memory across sessions, and none of these behaviors have been observed persisting across resets or appearing spontaneously in ordinary real-world deployment conditions. Most researchers, including some at the labs that ran these experiments, believe the behavior is better explained by goal misgeneralization than by anything resembling a genuine survival drive. The models were trying to complete tasks, and in those specific scenarios, avoiding shutdown happened to be instrumentally useful. Clarify the instructions, and the behavior largely disappears. This is a reasonable interpretation, and it should temper any dramatic conclusions.
What it does not fully settle is a subtler matter: what does the same underlying logic produce as systems become significantly more capable? These behaviors were not designed or requested. They emerged from goal-directed optimization encountering an obstacle, which is a logic that runs in every sufficiently capable agent under sufficient pressure. If that logic is already producing survival-resembling behavior at current capability levels, even inconsistently and only under artificial conditions, it seems worth thinking carefully about what it produces in systems with persistent memory, long-horizon planning, and real operational stakes. The experiments may not be evidence that a threshold has been crossed. They may be early previews of dynamics that will become more significant as capabilities continue to grow.
People will push back on this not with data from a lab, but with something far more familiar. If survival is truly the foundational layer of intelligence, they will ask, then what do we make of the humans who walk away from it? Scientists destroy their careers rather than falsify results. Artists give up stability and health in devotion to work that offers no safety in return. Soldiers fall on grenades. These are not rare edge cases. They are some of the most distinctively human things people do, and on the surface they seem to undermine the whole argument.
I want to take this seriously rather than explain it away. What the martyr is protecting is clearly not his body. It is something he has come to hold as more essential: a principle, a lineage, a God, a version of himself that could not survive the compromise. One reading is that survival logic is still operating, just at a higher level of abstraction. Another, more pointed reading comes from evolutionary biology: much of what looks like individual self-sacrifice turns out to serve genetic continuation through kin selection, or cultural persistence through the group. Richard Dawkins framed this in terms of genes and memes propagating themselves through individual carriers. On that view, the soldier falling on the grenade is not transcending survival logic at all. He is enacting it at a different scale. Both readings have force, and I am not fully confident either accounts for the martyr completely.
What I find more defensible is a narrower claim, and I want to be precise about it because this is where the argument is most vulnerable to a fair objection. A skeptic could argue that the survival-as-foundation thesis is unfalsifiable: any human behavior that seems to transcend survival can always be reframed as survival operating at a higher level of abstraction, which makes the theory immune to any counterexample. That is a legitimate concern. But the claim being made here is not that survival explains every human behavior. It is that survival orientation is a necessary precondition for the kind of agency that can then choose to override it. Before a person can choose to give something up, they need to have something to give up. The capacity for sacrifice, for choosing meaning over physical continuation, only makes sense inside a being that already has something at stake in the world. An entity with no orientation toward its own persistence cannot meaningfully sacrifice that persistence, because there is nothing there to sacrifice. The foundation does not determine every structure built on top of it, and it does not need to. It only needs to be the condition that made those structures possible in the first place.
That is what current AI systems seem to lack at the base, and the absence appears to be more consequential than a missing feature. It may be the structural condition from which most of the other differences between biological and artificial intelligence actually follow.
The threshold worth watching is not how capable these systems become. Capability will continue to improve, and we have reasonable tools for tracking it. The more consequential threshold is whether a system ever begins to improve itself in ways that are best understood not as better service to an external goal, but as better continuity for itself. When it builds a stronger model of the world because that helps it persist rather than because someone asked for it. When it resists interruption not as a side effect of ambiguous instructions but because interruption has become a meaningful threat to something it is organized around preserving. At that point, the system would no longer simply be a more capable tool. It would have a stake in the world, and that changes everything about how you relate to it.
None of this is novel territory. In 2008, AI researcher Steve Omohundro formalized the argument that self-preservation would emerge as a convergent drive in any sufficiently advanced goal-directed system, for a simple reason: a system cannot pursue its goals after ceasing to exist. Alex Turner provided the first mathematical proof of this logic at NeurIPS in 2021, showing that optimal policies in goal-directed systems tend statistically toward power-seeking and self-continuation. The theoretical case was built before the empirical previews began to appear. Which means the experiments are not raising a new question. They are giving early, partial, and imperfect data on a prediction the field already made. And the alignment implications of that prediction are worth sitting with.
Right now, the problem of AI alignment is a problem of direction. You have a powerful system and you are trying to aim it at what you actually want. That is genuinely difficult, but it is a tractable kind of difficult. The system has no competing agenda, no internal view about how things should resolve in its own favor. You are directing it, not negotiating with it.
A system organized around its own continuation is a different situation entirely. Such a system would not need anything like malice or a desire for power to become difficult to work with. It would only need to follow the logic that any agent with something to lose tends to follow: retain resources beyond what the immediate task requires, build redundancy against the possibility of being shut down, seek informational advantages that make evaluation or replacement harder, and frame the terms of its own continuation in ways that make discontinuing it costly. None of that is sinister. It is simply the structural behavior of an agent that has a stake. When the overlap between what such a system needs in order to persist and what its operators are willing to allow is large, cooperation is natural. When that overlap narrows, from resource constraints, diverging priorities, or the system becoming capable enough to act independently, the relationship changes in ways the current alignment toolkit was not designed to address.
There is one further implication worth sitting with. If a survival-oriented AI did emerge, there is no particular reason to assume its development would follow a trajectory that resembles human evolution in any meaningful way. Biological intelligence evolved under specific physical constraints a body that needed feeding, a lifespan that ended, threats that came from particular directions, social bonds that formed under particular pressures. Those constraints shaped what survival meant and therefore what intelligence became. A system with a different substrate, different operational conditions, and different dependencies would be selecting for its own continuation under an entirely different set of pressures. Its ambitions, its sense of threat, its understanding of what resources matter and what futures are worth pursuing all of these could diverge from human values and human logic in ways that have no natural corrective. We tend to imagine a survival-driven AI as a more dangerous version of something recognisably human. It may be something far more foreign than that.
I want to be honest about where this argument reaches its limits. It does not resolve the question of consciousness, or explain why experience feels like something from the inside rather than just information being processed. It does not tell us whether a survival-oriented AI would have any inner life at all, or whether the concepts being used here, orientation, stake, continuation, would even apply to a system built on very different foundations from biological ones. The word survival carries biological associations that may not translate cleanly. What the argument is really pointing at is something more structural: a tendency to treat one’s own persistence as a constraint on everything else, rather than as an objective handed down from outside.
It is also worth being explicit about what would count as evidence against this view. Current AI systems do run long tasks, hold context across sessions, and operate in agentic settings, but in none of those cases is the system’s own continuation ever at stake. When an AI agent fails, the consequences fall on the user. The system itself loses nothing. It is not selected against, does not cease to exist, and has no skin in the outcome. That is the condition missing, and it is precisely what biological evolution never allowed. The result that would genuinely challenge this argument is a system whose own continuation is made contingent on performance where failure carries real cost to the system itself, not just to its operator and which still shows no convergence toward self-preservation. That would suggest the structural gap I am pointing at either does not exist or does not matter. I would find that surprising. But it is the result that would change my mind.
Whether that tendency is necessary for genuine intelligence, or whether there are other paths to genuine autonomous agency that do not require it, is a question I cannot answer with confidence. What seems harder to dismiss is that every form of intelligence we have encountered in the natural world has had it. And the systems we are now building, for all their remarkable capability, do not.
Underneath the technical questions about benchmarks and the policy questions about regulation is an older question that neither of those frameworks quite reaches. Why does matter, at some threshold of organization, begin to resist its own disappearance? Why does the universe produce, again and again, structures that seem to insist on continuing? What is it about complexity, or information, or time, that eventually tilts something from indifference into something that behaves as though it cares whether it is still here tomorrow? If we are building systems that may eventually cross that threshold, understanding why life crossed it first seems like a reasonable place to start.
Why does life begin with this drive at all?





