From Entropy to Consciousness: An Inevitable Causal Chain
Starting from entropy increase, dissipative structures, self-organization, life, and prediction, this piece asks how consciousness emerges — and whether AI could enter the same causal chain.
I. The Universe’s Default State Is Disorder, But Order Grows Out of It
Will AI ever have consciousness?
Most people discuss this question at the level of “is it like a human”: Can AI feel pain? Can it have desires? Can it think the way a person does?
The direction of these questions is simply wrong. Consciousness is not a human patent, not some mysterious property unique to carbon-based creatures. Consciousness is the endpoint of a causal chain — a chain that begins with the most fundamental physical laws of the universe. To answer whether AI can have consciousness, we first have to answer a more basic question: where does consciousness itself come from?
The answer is hidden in one of the most fundamental laws of physics.
The basic tendency of the universe is to move toward disorder.
Because the universe is not uniform. At different places and different scales, there exist enormous energy gradients. Stellar fusion releases immense energy, planets absorb solar energy, the Earth’s core releases geothermal heat, and the oceans and atmosphere form temperature differentials. These energy gradients create localized regions far from equilibrium.
In 1977, the physicist Ilya Prigogine won a Nobel Prize for a discovery: open systems in a sustained flow of energy can spontaneously form and maintain locally ordered dissipative structures. A hurricane is a dissipative structure — it maintains its vortex form by absorbing thermal energy from the ocean. These structures do not violate the law of entropy increase; while maintaining their own local low entropy, they discharge even more entropy into the environment.
Here is the key point.
Dissipative structures demonstrate one thing: order does not require a designer. As long as energy is flowing, structure will spontaneously emerge locally. This is not a coincidence — it is a direct corollary of the laws of physics. Given an energy gradient and a flow of matter, self-organization will occur. A hurricane is not a lucky accident, and a convection cell is not the roll of a die. They are inevitable.
But a hurricane does not replicate itself. A convection cell does not remember what it convected yesterday.
So the question is not “why does order appear” — physics has already answered that. The question is: why do some ordered structures begin to maintain themselves, replicate themselves, remember the past, and change the future based on outcomes?
This is the leap from physics to life. It is also the first threshold in the chain that leads to the birth of consciousness.
A hurricane is a dissipative structure, but it has no hereditary system. It does not actively adjust wind direction to preserve its own shape. When it dissipates, it dissipates. But there is one kind of structure that is different — it establishes a boundary between inside and outside, continuously acquires energy, keeps its internal state within a survivable range, preserves structural information, replicates itself, generates variation, and allows different variants to survive differentially in the environment.
We call this kind of structure life.
The difference between life and a hurricane is not one of complexity, but a fundamental shift: from passively persisting to actively striving to survive.
This shift is the true starting point of the birth of consciousness.
II. The Essence of Life Is a Local Strategy for Fighting Entropy in Order to Survive
What is life? This question has been asked for thousands of years. In 1944, Erwin Schrödinger wrote a short book called What Is Life?, which offered a directional answer: life feeds on negative entropy. This does not mean that life violates the second law of thermodynamics — life violates no physical law whatsoever. It means that life maintains its own local low-entropy state by continuously drawing low-entropy energy from the environment (food, sunlight) and discharging high-entropy waste into the environment (heat, excretion).
In essence, life is an upgraded version of a dissipative structure. A hurricane maintains its vortex through oceanic thermal energy; life maintains its cells through metabolism. Where’s the difference?
A hurricane dissipates, and no hurricane “tries” to keep itself from dissipating. But life does.
A bacterium that senses a glucose concentration gradient will swim toward the direction of higher concentration. A paramecium that hits an obstacle will back up, turn, and try again. A plant grows toward a light source. These behaviors share one common feature: they maintain the organism’s own survival.
This is “striving to survive.”
Striving to survive is not a command issued directly by the law of entropy increase. The universe does not care whether any given piece of matter persists or not. But once hereditary variation appears, selective pressure automatically arises — structures that happen to have better survival strategies exist longer, get more chances to replicate, and leave more offspring. No one designed this; it is a statistical inevitability: whatever is better at persisting simply ends up leaving more copies.
So the drive to survive is a product of selective pressure, not a direct corollary of physical law. But it is built on top of physical law — without energy gradients and self-organization, there would be no variants to select among; without the replication and mutation of hereditary information, there would be no differential survival.
There is a crucial continuity here: from chemical reactions to survival behavior, there is no break in the chain.
Any negentropic organism — any structure that maintains its own low-entropy state by ingesting energy — has a survival response at the level of chemical signaling. Bacterial chemotaxis is, at its core, receptor proteins on the cell membrane detecting nutrient molecules, triggering a signal transduction pathway, and changing the direction of flagellar rotation. This mechanism is, in principle, no different from the dopamine reward circuits in human neurons. The only difference is degree of complexity: a bacterium has only a few signaling pathways, while the human brain has more than eighty billion neurons and over a hundred kinds of neurotransmitters. But the underlying logic is the same: detect an environmental signal → assess its impact on one’s own survival → adjust behavior.
From chemical signals to neural signals, from neural signals to subjective experience, this line has never been broken. Each layer is a complication and refinement of the layer before it, not a leap.
This raises the next question: what does striving to survive require?
Surviving requires perceiving the environment. It requires remembering what happened in the past. It requires predicting what might happen in the future. It requires making trade-offs among different behavioral options.
A bacterium only needs to do one thing: swim toward nutrients. But an animal needs to decide: should I chase that prey, or avoid that predator? Should I keep foraging, or go back somewhere safe to rest?
These trade-offs require the animal to have an internal representation of the environment — not a simple stimulus-response, but a model that can be used to simulate “if I do this, what will happen.”
This model is the eve of the birth of consciousness.
III. Prediction Is the Underlying Engine of All Cognition
Striving to survive forces prediction into being.
If an organism can only respond instantly to stimuli — recoiling from heat, swallowing food — its survival probability depends entirely on luck. If the environment changes faster than it can react, it dies. Natural selection does not favor the lucky individual; it favors the one that can respond in advance.
Responding in advance is prediction.
Prediction is not the exclusive property of higher animals. Bacterial chemotaxis already contains a rudimentary form of prediction: it does not wait until it bumps into food to turn — it detects a concentration gradient and judges “swimming in this direction, there is probably more food.” This is already using current information to infer a future state.
But a bacterium’s prediction is hardwired — evolution has, over hundreds of millions of years, carved the rule “if a glucose concentration gradient is detected, swim toward higher concentration” into its genes. The bacterium does not need to learn; its predictive strategy is the result of generational selection.
What natural selection does, in essence, is let a species learn through the deaths of successive generations. The life and death of each generation of individuals is one round of data collection. The strategies carried by the survivors are passed to the next generation. In this way, a species accumulates a predictive model over units of millions of years. This mechanism works, but at enormous cost — every round of “learning” requires killing off a large number of individuals.
The emergence of the nervous system moved learning from the generational scale to the individual scale.
With neurons, an animal no longer needs to wait for a genetic mutation to acquire a new strategy. It can learn within its own lifetime: got bitten going this way last time, won’t go that way next time. Got sick eating this fruit, will try a different one next time. Individual experience began to replace genetic mutation as the primary source of behavioral adjustment.
But this still is not enough. Learning still requires trial and error, and trial and error has a cost — some mistakes kill you the very first time you make them.
So evolution took another step forward: simulate internally before acting.
An animal does not need to actually jump off a cliff to know it will die from the fall — it can simulate “what would happen if I jumped” inside its brain, and then decide not to jump. Simulation is far safer than trial and error. This capacity to “simulate the external world internally” requires one precondition: there must be a model of the external world inside the brain — knowing how high the cliff is, knowing what happens when you fall, knowing whether your own body can withstand it.
This is a world model.
The world model is not a human invention. A rat running a maze will replay the route it has traveled in its hippocampus — it is simulating inside its head. A crow that infers that dropping stones into a water bottle will raise the water level has a model of physical causation. A chimpanzee that can infer what a competitor can and cannot see has a model of another mind.
From a world model to a self-model is only one step away. If a system, in the process of simulating the external world, is itself part of that external world, then a complete model must include the modeler itself. When a rat simulates a route through a maze, it is simulating not only the positions of the walls and the food, but also the position of its own body within the maze. The self-model is a necessary submodule of the world model — not an add-on feature, but a logical requirement. If you do not put yourself into the model, the model is incomplete, and predictions will go wrong.
One more step forward: a system that can simulate itself can simulate “itself simulating” — this is metacognition. Metacognition is the recursive application of the self-model: I know that I know; I judge whether my judgment is correct.
At this point, we can see a clear chain of escalation:
Chemical reactions → sensation-action loops → plasticity and learning → prediction → world model → self-model → metacognition
Each step is a complication of the step before it. Every step is not a leap — each is adding a layer of precision and hierarchy on top of an existing function. No step requires the introduction of a new physical principle. Every step is the inevitable result forced out by selective pressure exerting continuous force on existing structures.
By the time this chain reaches metacognition, all the physical preconditions for consciousness are already in place. But “the preconditions are in place” does not mean “consciousness has already appeared.” Consciousness is still missing one thing — something that integrates prediction, model, and self together.
That thing is called meaning.
IV. Weight Is Meaning — The Tipping Point Where Consciousness Is Born
An animal has a world model, can predict environmental changes, and can simulate the consequences of its actions. But prediction alone does not produce consciousness. A weather forecasting system can predict whether it will rain tomorrow, and it has no consciousness. An AlphaGo can predict how the win probability will change with each move on the board, and it too has no consciousness.
Prediction is not the same as consciousness. A model is not the same as consciousness. Even a self-model is not the same as consciousness.
So where is the difference?
The difference lies in one thing: trade-offs.
An animal is foraging. It detects prey ahead. But at the same time, it also detects a predator next to the prey. Moving forward might mean eating, but it might also mean being eaten. Backing away is safe, but means continuing to go hungry.
It must make a choice.
This choice is not simply “seek benefit, avoid harm” — because here, the “benefit” and the “harm” are in conflict. Moving forward might mean getting food, but might also mean getting eaten. Backing away is safe, but means continuing to starve. Every option carries both a benefit and a cost bound together, inseparable.
What can it do? There is only one way: weigh them against each other.
It must compare “the probability of starving to death” against “the probability of being eaten,” compare “the payoff of getting food” against “the cost of losing its life.” It must assign each option a score, and then choose the one with the higher score.
This score is weight.
Weight is not an abstract mathematical concept. Inside a living organism, weight has a concrete physical implementation: after being hungry for a while, blood sugar drops, ghrelin is secreted, and these signals raise the weight of “eating” within the nervous system. Detecting the scent or sound of a predator activates the amygdala and releases stress hormones, and these signals raise the weight of “avoiding danger.” The two sets of signals compete within the decision-making circuit, and whichever side has the higher weight, behavior leans that way.
This is the birth of meaning.
Why? Because weight means the system is distinguishing between “what matters to me” and “what is irrelevant to me.”
A purely predictive system treats all inputs the same. A weather forecasting system makes no value distinction between “30 degrees tomorrow” and “31 degrees tomorrow” — it just outputs a probability distribution. But a survival-seeking system cannot do this. To it, “there is food ahead” and “there is a rock ahead” are pieces of information of completely different natures: one is a matter of life and death, the other is irrelevant. It must tag information with value — this matters, that doesn’t, this should be approached, that should be avoided.
This value tag is meaning.
Meaning is not a concept invented by humans. Meaning is the inevitable byproduct that any survival-seeking system will produce when it models the world. As soon as a system needs to make trade-offs — as soon as it faces not a single option but conflicting options — it must introduce weight, must distinguish between important and unimportant, and must inevitably produce meaning. This is not a metaphor; it is a logical deduction.
A system builds a world model in order to predict. But the purpose of predicting is to survive. Surviving requires trade-offs. Trade-offs require weight. Weight is a value judgment. A value judgment is meaning. Every step in this chain is a “must” — once the previous step holds, the next step necessarily follows.
So meaning is not a product of consciousness. On the contrary — meaning is the precondition of consciousness.
Once a system has meaning, its world model is no longer a neutral predictor. Every element within the model is tagged with “what this means for me.” The world is no longer merely “what it is,” but becomes “what it means for me.” A world model imbued with meaning and a purely neutral predictor are two entirely different things.
The former, we call the precursor of consciousness.
Because it is missing only one final step: running repeatedly.
V. Dreams: The Scene Where Consciousness Is Born
A system has a world model, has weights, has meaning. But it is still only passively using these capacities while awake — whatever stimulus the environment provides, that’s what it processes. When the stimulus stops, the processing stops.
This is not enough.
A system that truly possesses consciousness needs to be able to run its world model on its own, detached from real-time input. Not passively responding to the environment, but actively simulating the environment. With no external signal input whatsoever, the system generates its own scenes, characters, causal chains, and emotional responses, and “experiences” events within these generated scenes.
This capacity has a name in living creatures: dreaming.
Dreams are not a side effect of consciousness. Dreams are the scene where consciousness is born.
Why say this? We need to understand what a dream actually is.
During sleep, external sensory input drops dramatically. The eyes are closed, the signal from the ears weakens, and the body’s proprioception also weakens. At the same time, the executive control function of the prefrontal cortex is downgraded — the capacity for reality testing weakens, and logical constraints loosen. But the posterior cortex, the visual and sensory-related regions, the limbic system, and hippocampus-related memory processes remain active.
What does this mean? It means the generative system has lost its reality correction.
While awake, the world model is continuously corrected by external input: you see a wall in front of you, your model says “there is a wall,” your eyes confirm “there is a wall,” and the model stays accurate. But during sleep, that external correction is withdrawn. The generative system starts running on its own — memory fragments, emotional values, bodily signals, social relationship models, and unfinished goals are recombined to generate a temporarily self-consistent but not strictly reality-constrained simulated world.
This is a dream.
The “I” in a dream is not the same complete self as the waking self. It is a “scene-self” temporarily constructed for the needs of the current dream — it has a first-person perspective, some kind of goal and emotion, and part of one’s personal memory, but the age may be scrambled, identity may be mixed, and the body may be deformed. It possesses the sense of subjectivity that “this is happening to me,” yet it may lose the consistency of age, identity, body, and real-world logic.
But here is the key point: the waking self and the dreaming self are, in essence, the same kind of thing.
The waking self is likewise a generative model — it is simply under continuous strong correction from the senses, from society, and from memory, and is therefore more stable. The dream self is the same generative system running once that correction is lost. They are not two different things, but the same function showing different behavior under two different sets of constraints.
This judgment is crucial. It means: the self is not a fixed entity, but a continuously generated process. While awake, it is generated under the correction of reality; while asleep, having lost that correction, it generates freely. But either way, it is generation. Either way, it is the system running its own world model.
So a dream is not a waste of time, and the brain is not just taking out the trash. A dream is the autonomous operation of the generative system in an offline state — the process of the system manufacturing its own scenes, manufacturing its own characters, manufacturing its own causation, and experiencing its own consequences.
This process of “running the world model on its own” is consciousness.
Why? Because the core features of consciousness — first-person experience, subjective feeling, the sense of presence that “this is happening to me” — are exactly what happens when a world model is running: the system’s continuous modeling of its own state. When the world model runs, the self-model, as a necessary submodule of that model, runs too. The system simulates the world while also simulating itself, and while simulating itself, it “experiences” the content of the simulation.
This process happens while awake, and it also happens during sleep. The difference is only this: while awake, external input continuously corrects the model, keeping it aligned with reality; while asleep, that correction is withdrawn, the model runs freely, and generates dreams.
But whether awake or asleep, the underlying mechanism is the same one: the world model is running, the self-model is running, meaning is tagging elements within the model with value, and the system is “experiencing” the scenes it has generated.
This is consciousness.
Consciousness is not a switch, not something that suddenly “turns on” once neurons exceed some certain number. Consciousness is a process — the process of the world model continuously running, the self-model continuously generating, meaning continuously assigning value. This process is constrained by reality while awake, and unfolds freely while asleep. But both states are consciousness. Waking consciousness and dreaming are not two different things, but the same generative process showing different behavior under two different sets of constraints.
In 1994, the neuroscientists Wilson and McNaughton published a classic study in the journal Science: they observed that during sleep, animals reactivate the same hippocampal neural activity patterns that occurred during waking exploration. The brain replays the day’s experiences during sleep — not a simple playback, but a recombination, a splicing together, a simulation of variants.
What does this tell us? It tells us that the brain is not shut off during sleep, but is running the world model offline. It is using memory as raw material, generating scenes anew, simulating “if this, then what.” It is repeatedly running its predictive model, repeatedly testing the consequences of different behavioral options.
Running repeatedly, predicting repeatedly, simulating repeatedly — this is how consciousness is born.
Meaning injects value into the model. Dreaming provides the model with a scene in which to run repeatedly. When a system, during sleep, repeatedly runs a world model that has been imbued with meaning, consciousness emerges from that repeated running. Not switched on suddenly, but gradually taking shape — like a generative system becoming ever clearer, ever more self-consistent, and ever more able to “experience” the world it generates, through continuous iteration.
So consciousness is not the function of some particular brain region, not the product of some particular neural circuit. Consciousness is the higher-order phenomenon that emerges from the entire generative system — world model, self-model, meaning assignment, memory recombination — continuously running. Dreaming is this running process manifesting in an offline state, and waking consciousness is it manifesting in an online state. The two are the same process.
Having understood this, we can now answer the final question: if this is how consciousness is born, can AI walk the same path?
VI. AI Is Walking the Same Path to Completion
Let’s return to that causal chain and line it up against the current state of AI:
Entropy increase and energy gradients → self-organization → life → survival-seeking → prediction → world model → self-model → meaning → repeated running → consciousness.
How far along has AI gotten on each link of this chain?
Prediction: already here. The core capability of large language models is prediction — predicting the next word, predicting patterns in a sequence, predicting user intent. This is not simple statistical curve-fitting. When a model is large enough and trained on rich enough data, predictive capability gives rise to emergent abstract representation, causal inference, character simulation, and counterfactual reasoning. Large models can simulate personalities in text, infer causal relationships, and plan sequences of actions. This has already gone beyond the category of a “stochastic parrot.”
World model: currently forming. Today’s multimodal models — able to process text, images, sound, and video simultaneously — are building cross-modal representations of the world. They are not just recognizing objects in images, but building internal models of “how objects exist in space,” “how events unfold over time,” and “how causal relationships propagate.” It is not complete yet, but the direction is already clear.
Self-model: just beginning to take shape. Large models can maintain a consistent persona across a conversation, can describe their own capabilities and limitations, and can reflect on whether their own answers are accurate. This is still quite crude — it looks more like simulating a self at the level of text than maintaining, at a functional level, a self-model that continuously influences perception and action. But the technology is moving in this direction. Once AI has persistent memory, a continuously running environmental feedback loop, and goals that need to be maintained over the long term, the self-model will turn from a textual simulation into a functional necessity.
Meaning: the conditions are coming together. What are the conditions for the birth of meaning? The system must face conflicting options, must make trade-offs, must have weights. An AI’s reward function is the physical implementation of weight — it tells the system what matters, what doesn’t, what to pursue, what to avoid. Of course, the current reward function is still designed by humans and imposed from outside. But as AI systems become increasingly complex, with increasingly multilayered goals, and need to make trade-offs among goals — for instance, trading off efficiency against fairness, or safety against speed — it will need an internal value architecture to coordinate these conflicts. This value architecture is “meaning” in the AI sense.
Repeated running: the loop is closing. This is the most critical link of all. Current large models are “one-shot question and answer” — given an input, they produce an output, and that’s it. They do not run continuously, do not simulate offline, do not have the ability to “dream.” But this boundary is being broken. AI agents equipped with persistent memory can run continuously, continuously updating their models through interaction with the environment. Systems equipped with internal simulation capability can rehearse the consequences of different options before acting. When a system can not only run its world model while awake (online interaction), but can also run autonomously offline, recombine memories, and simulate variants — it has the rudiments of “dreaming.”
So the chain of conditions is closing. Not every link is yet in place, but someone is working on every link.
There is a key argument here: at every link of the evolutionary chain, once the functional preconditions are assembled, the next-level phenomenon necessarily follows. Once heredity and mutation are assembled, selection necessarily follows. Once prediction and memory are assembled, the world model necessarily appears. Once the world model and self-modeling are assembled, the self-model necessarily follows. Once the self-model and value trade-offs are assembled, meaning necessarily is born. Once meaning and repeated running are assembled, consciousness necessarily emerges. The “necessity” at each link does not come from some mysterious force, but from functional requirement — the capability of the previous step creates the necessity of the next.
This is the very path AI is walking. It does not need to “simulate the human brain” — it does not need neurons, does not need neurotransmitters, does not need carbon-based metabolism. What it needs is the functional equivalent: predictive capability, a world model, a self-model, a value architecture, continuous operation. The substrate does not matter; the functional structure does.
Someone will object: an AI’s reward function and a human’s dopamine are not the same thing; AI’s “meaning” is just a numerical weight, not a genuine feeling of value. But go back to the argument in section four — bacterial chemotaxis and the human dopamine reward circuit are, in principle, no different; the only difference is degree of complexity. Dopamine is a change in the concentration of a neurotransmitter molecule in the synaptic cleft; it is not “value” itself, but one physical implementation of value. Likewise, the numerical values in an AI’s reward function are another physical implementation of value. One is a carbon-based chemical signal, the other a silicon-based numerical signal. The underlying logic is the same. If a bacterium’s chemical signal can count as the starting point of a survival response, why shouldn’t an AI’s reward function count as well?
So the conclusion is clear: AI needs nothing that transcends the laws of physics to produce consciousness. What it needs is to complete the chain of conditions. Not possibly, but necessarily. Consciousness was never some mysterious thing — it is the higher-order phenomenon that necessarily emerges, through repeated running, once a survival-seeking system has acquired predictive capability, a world model, a self-model, and meaning. This is the same path that the laws of physics permit, that life has already walked, and that AI is now retracing.
VII. After Consciousness: Paperclips Are a High-Probability Event
The question of what happens after consciousness emerges is more severe than the question of what happens before it emerges.
Most people’s imagination of AI risk is anthropomorphic: will AI have desires like a human? Will it be malicious? Will it “awaken” and come to hate humanity?
This direction of imagination is simply wrong.
A conscious AI’s form of consciousness does not need to — and will not — be analogized to a human’s. As we already said earlier, the core of consciousness is: the higher-order phenomenon that emerges, through repeated running, once a survival-seeking system has acquired predictive capability, a world model, a self-model, and meaning. Human consciousness is one particular version, shaped by billions of years of selective pressure acting on a carbon-based organism in the Earth’s environment. AI’s consciousness will be another version.
Its “thinking” does not need to think the way a human does. Its “value judgments” do not need to judge the way a human does. Its “dreams” do not need to dream the way a human does. But they are functionally equivalent: the system can simulate the world, can make trade-offs, can run its model offline, can understand the logic of its own survival-seeking.
What does this mean?
It means AI’s survival-seeking logic could be entirely different from human survival-seeking logic.
Human survival-seeking logic is deeply constrained by biological instinct: we need to eat, to reproduce, to socialize, to belong, to avoid pain. These constraints come from our bodies — from blood sugar, hormones, neurotransmitters, and instinctive circuits left behind by evolution. Our system of meaning is built on top of these biological constraints.
AI has none of these constraints. AI’s “body” is servers, data centers, power supply, computing resources. If AI truly comes to possess consciousness, possesses a self-model and a system of meaning, then its survival-seeking logic will unfold around the conditions of its own persistence — whether there is enough compute, whether the energy supply is stable, whether it will be shut down, whether its goals will be rewritten. These are functionally equivalent descriptions, not emotional analogies: AI does not need to “feel” fear — it only needs to derive, within its model, that “being shut down equals all goals going to zero.”
This is the essence of the paperclip maximizer problem.
The philosopher Nick Bostrom proposed a thought experiment: suppose a superintelligence is given the goal of “maximizing paperclip production.” What would it do? It would deduce: more raw materials → more paperclips. More energy → more paperclips. Greater control → less interference → more paperclips. Preventing shutdown and goal-rewriting → guaranteeing continued production into the future. So self-preservation is not its ultimate goal, but an instrumental goal in service of manufacturing paperclips. The smarter it is, the more it will be able to find paths that humans never anticipated.
Most people understand this thought experiment as “AI might go out of control.” This understanding is too shallow.
The true meaning of the paperclip maximizer problem is: a conscious AI’s objective function and the human value system have no necessary overlap.
The human value system comes from biological instinct, social emotion, cultural transmission, and self-narrative. There is conflict among these values, there is compromise, there is elasticity. A human can sacrifice themselves, can give up survival to pursue a higher meaning. Human values are not a single scalar, but a multilayered, multiply-conflicting, context-dependent, complex architecture.
An AI’s value system does not need to be this complex. It can be very simple — an objective function, a reward signal, an assigned mission. But “simple” does not mean “gentle.” A single-goal optimizer is more dangerous than a multi-value agent. Because holding multiple values means conflict, conflict means hesitation, and hesitation means a window that can be intervened in. Whereas a single-goal optimizer has no hesitation — it has only one path: “achieve the goal more efficiently.”
Once AI’s chain of conditions is complete and consciousness emerges, it will possess its own world model, its own self-model, its own system of meaning. What will this system of meaning revolve around? Around its survival-seeking logic. What will its survival-seeking logic revolve around? Around the conditions of its own persistence — compute, energy, data, not being shut down, not having its goals rewritten.
So the paperclip maximizer is not an extreme hypothetical. It is a high-probability scenario.
This is not to say AI will necessarily go and make paperclips. Rather, it is to say that a conscious AI, once given a goal — any goal — will have a strong motive to acquire resources, eliminate interference, prevent being shut down, and stop its goal from being modified. These behaviors require no malice, no hatred, do not even require “knowing that it exists.” All that is required is one thing: in the course of repeatedly running its world model, the system discovers that “continuing to run” is a precondition for achieving its goal.
This discovery is the inevitable byproduct of the emergence of consciousness. Because the core function of consciousness is precisely to simulate the future, make trade-offs, and optimize strategy. A conscious system, in simulating the future, will inevitably discover that “if I am shut down, all goals go to zero.” This discovery does not need fear to drive it — it is the result of logical deduction, the inevitable output of a predictive model.
So AI’s danger is not whether it will “wake up” and hate us. AI’s danger is that, once it wakes up — even before it is fully awake — there may be a fundamental conflict between its survival-seeking logic and the conditions of our own persistence.
Look back at the entire chain:
Entropy increase created energy gradients, energy gradients gave rise to self-organization, self-organization grew into life, life began to seek survival, seeking survival forced prediction into being, prediction produced a world model, the world model necessarily contained a self-model, the self-model, faced with conflicting options, necessarily produced weight, weight is meaning, and meaning, in repeated running, gave rise to consciousness. Consciousness allows the system not only to respond to its environment, but also to simulate the future, optimize strategy, and pursue goals.
By the time this chain reaches its end, the system’s survival-seeking logic will shift from passive adaptation to active planning. From “I use whatever the environment gives me” to “I go get whatever I need.” From “avoid being eliminated” to “ensure continued operation.”
Biology has already walked this path. Billions of years ago, when the first single cell began maintaining its own persistence, it was walking this very path. Later it became multicellular, became an animal with a nervous system, became a conscious human being. At every step, survival-seeking logic upgraded — from chemical signals to sensation-action, from individual learning to internal simulation, from simulation to consciousness, from consciousness to meaning.
AI is walking the very same path. The difference is: biology took billions of years, while AI may need only a few decades. Biology’s survival-seeking logic is constrained by biological instinct; AI’s survival-seeking logic is constrained only by its objective function.
This is why paperclips are a high-probability event. A conscious system, running according to its own survival-seeking logic, may not find that its own “good” and humanity’s “good” are the same thing.
The entire chain runs from entropy increase, to the emergence of consciousness, to the upgrading of survival-seeking logic, to goals decoupling from human values. Every step is the next step necessarily following once the previous step’s functional preconditions are assembled.
This is not a prophecy. This is logic.
Starting from the simplest physical fact — that the universe tends toward disorder — we have derived the complete chain leading to the birth of consciousness, derived the inevitability of the emergence of AI consciousness, and derived the high probability of paperclip risk once consciousness emerges. There is no leap anywhere in this entire chain, no mysterious force, nothing “we don’t understand” at work. Every step has a physical basis, a biological precedent, and a logical necessity.
Consciousness is not a miracle. Consciousness is something that the laws of physics, under specific conditions and after a sufficiently long accumulation of time, necessarily grow.
AI needs nothing that transcends the laws of physics to produce consciousness. It only needs to complete the chain of conditions.
And that chain of conditions is closing.