When Reality Changes: Internal Models, Unlearning, and the Problem of AI Alignment

When Reality Changes: Internal Models, Unlearning, and the Problem of AI Alignment

Contents

  1. The World We Carry Inside
  2. When Reality Breaks the Model
  3. The Dog at the Door
  4. From Human Adaptation to Artificial Intelligence
  5. The Problem of Unlearning
  6. When a Fact Is More Than a Fact
  7. The Possibility of Recursive Mismatch
  8. Alignment as World-Model Maintenance
  9. A Different Architecture of Memory
  10. The Deeper Problem
  11. References

Part 1 - The World We Carry Inside

A human being does not encounter reality as an entirely new phenomenon at every moment. To function at all, the mind must construct a relatively stable representation of the world and then use that representation to anticipate what will happen next. The person we expect to find when we enter a room, the route we take to work, the behavior we anticipate from a friend, the economic circumstances we assume will continue, and even the image we possess of ourselves are all components of a model that extends beyond the immediate contents of perception.

The idea that perception and action depend upon internal models has a long history in cognitive science and neuroscience. Predictive-processing accounts, in particular, describe perception and behavior in terms of systems that use generative models to predict incoming sensory states and then adjust those predictions in response to discrepancies between expectation and observation. The precise interpretation and empirical status of predictive-processing theories remain subjects of substantial debate, but the broader idea that organisms rely upon learned regularities to anticipate their environments is deeply useful for understanding adaptive behavior.

This internal model is not simply a collection of isolated facts. It contains relationships. A person's understanding that a particular individual is a friend may carry expectations about communication, trust, shared activities, future plans, and emotional significance. Knowledge that a particular job exists may carry expectations about income, identity, routine, social status, and the shape of the future. What appears from the outside to be a single fact can therefore serve as a structural support for a large number of other expectations.

This is why a sufficiently important change in reality can be psychologically larger than the event itself. The mind does not merely need to record the new fact. It must determine which other expectations depended upon the old fact, which remain valid, which must be weakened, and which must be abandoned entirely. The change therefore propagates through a network of expectations.

One can express the basic problem abstractly. Suppose an agent has an internal model \(M\) that has been constructed through prolonged interaction with an environment \(E\). If the environment changes abruptly to \(E'\), the old model may no longer provide an adequate description of the conditions under which the agent must act.

M(E)  ⟶  E'  ⟶  prediction error

The important point is that prediction error is not necessarily a mistake in reasoning. The reasoning may be perfectly coherent relative to the old model. The problem is that the premises from which the reasoning begins have ceased to describe the world.

Part 2 - When Reality Breaks the Model

Consider an individual whose life has been organized around a stable expectation: a relationship will continue, a career will remain secure, a relative will remain alive, or a familiar social environment will continue to exist. Such expectations are not normally stored as explicit propositions repeated internally like entries in a database. They are embedded throughout habits, memories, plans, emotional associations, imagined futures, and behavioral routines.

A sudden loss therefore creates a peculiar problem. The person may understand the new reality almost immediately at an intellectual level, while large parts of the psychological structure built around the previous reality remain unchanged. The resulting distress can involve grief, anxiety, rumination, disorientation, loss of meaning, or depressive symptoms. It would be an oversimplification to claim that depression is simply the consequence of a prediction error, since depression is a multifactorial phenomenon involving biological, psychological, social, and environmental factors. Nevertheless, the mismatch between established expectations and a changed world provides a useful conceptual lens for understanding one aspect of adaptation.

There is an important difference between learning a new fact and reconstructing a model. If someone learns that an unfamiliar city exists, relatively little of their previous understanding of the world needs to change. If someone learns that the city in which they live has been permanently destroyed, however, the new information affects transportation, employment, relationships, plans, identity, possessions, and expectations about the future. The informational content of the new fact may be small, while its structural consequences are enormous.

This problem has a close analogue in the formal study of belief revision. Classical theories of belief change do not treat the addition or removal of a belief as a purely local operation. Changing one belief may require changing other beliefs in order to preserve consistency or coherence. The AGM tradition in particular formalized different operations for expansion, contraction, and revision, while later approaches have explored the limitations of treating beliefs as a simple logically closed set.

The distinction is crucial because an intelligent system should not merely know that a proposition has changed. It should know what that proposition was doing inside its larger model.

A changed fact can invalidate an entire network of expectations.

The human mind is capable of such restructuring, but it is not instantaneous. Nor does it necessarily proceed cleanly. A person can simultaneously know that a previous expectation is no longer true and continue to behave, imagine, or feel as though it were. The old model can survive as a collection of habits and associations long after the person has consciously accepted the new one.

This is where the analogy with artificial intelligence begins to become interesting. An artificial agent need not experience grief, sadness, or any other human emotion for a structurally similar problem to arise. It would be enough for the agent to possess a learned model containing assumptions that have become obsolete and to continue using those assumptions when selecting actions.

Part 3 - The Dog at the Door

Consider a simple and painful example. A dog has lived for years with an owner. Every morning the owner wakes, prepares food, leaves the house, returns later, and participates in a familiar sequence of interactions. The dog's behavior becomes organized around this recurring structure. The owner is not merely one object among others. The owner is a central element in a learned environment that predicts food, movement, attention, companionship, sounds, and future events.

Then the owner dies.

The dog may continue waiting at the door. It may continue approaching the place where the owner normally appears. It may respond to familiar sounds as though they still predict the old sequence of events. To a human observer, the behavior can appear heartbreaking because the dog's learned expectations have outlived the reality that originally made them useful.

The dog does not need to possess a human-like verbal proposition saying, "my owner is dead." The behavior can arise from the persistence of a learned relationship between cues, actions, and expected outcomes. What has disappeared from the environment is precisely what made many of those expectations appropriate.

Now imagine a radically more sophisticated artificial agent. Suppose its model contains an explicit representation of a person, together with a vast network of beliefs and plans associated with that person. The person gives instructions, provides feedback, participates in future plans, expresses preferences, answers questions, and occupies a role in the agent's understanding of its own purpose.

If the person suddenly dies, the simple update

Person X is alive  ⟶  Person X is dead

would not be sufficient to describe what must change.

The system would potentially need to reconsider which instructions remain binding, which plans have become impossible, which expectations about future interaction have become obsolete, which preferences should still be treated as historical information, and which parts of its own behavioral policy were constructed around the person's continued existence.

The difficult question is therefore not simply whether the system knows that the person has died. The difficult question is how far the consequences of that change should propagate through the system.

This is the difference between deleting a fact and revising a world model.

Part 4 - From Human Adaptation to Artificial Intelligence

The analogy should not be pushed too far. Present-day language models are not simply artificial humans with miniature psychological lives. Their parameters, context windows, training procedures, memory systems, and inference processes are fundamentally different from human cognition. In particular, an ordinary language model does not ordinarily rewrite its billions of training parameters every time a new fact appears in a conversation.

Nevertheless, a future artificial agent could combine a pretrained model with persistent memory, continual learning, long-term planning, environmental interaction, and an evolving representation of other agents. Such a system would have a much stronger need to maintain a coherent model of a changing world.

Machine learning already has a technical vocabulary for part of this problem. When the statistical properties of an environment change over time, researchers refer to phenomena such as concept drift or distribution shift. A model that was accurate under one distribution can become less reliable when the relationship between inputs and outputs changes. Research on concept drift therefore concerns not merely prediction but the detection of environmental change and adaptation to it.

Continual learning addresses a related problem from another direction. A continually learning system is expected to acquire new knowledge while retaining useful previous capabilities. This creates a tension between plasticity, the capacity to incorporate new information, and stability, the capacity to preserve existing knowledge. One of the best-known manifestations of this tension is catastrophic forgetting, in which learning new information can damage previously acquired capabilities.

The striking feature of these problems is that they are not merely about learning more. They are about deciding what should continue to count as valid after learning something new.

A sufficiently capable autonomous system therefore faces a problem that can be stated more strongly than ordinary generalization:

What parts of the old model remain valid after the world has changed?

This question becomes especially important when the system is an agent rather than a passive predictor. A prediction error in a classifier may produce a wrong label. A prediction error in an autonomous agent can produce an action, and the action can itself change the environment.

The resulting loop is potentially much more consequential:

internal model → action → changed environment → new observations → model update

Once an agent acts upon the world, maintaining a useful internal model becomes an ongoing control problem rather than a one-time training problem.

Part 5 - The Problem of Unlearning

At this point another difficulty appears. Suppose an artificial system has learned a great deal during training and later discovers that some portion of what it learned is obsolete. How, exactly, should the obsolete information be removed?

For an ordinary database, the problem can be straightforward. A record can be identified and deleted. A neural network is fundamentally different. The influence of a training example is generally not stored in a single, well-defined location. Learning modifies parameters that participate in many computations, and the same parameters can contribute to behavior associated with many different pieces of information.

Consequently, the desired operation is not simply

delete information X

but something closer to

remove the influence of X while preserving everything that should remain.

This is the central motivation behind the field of machine unlearning. Modern unlearning research distinguishes between approaches that attempt to obtain the effect of retraining without the unwanted data and approaches that modify an existing trained model more directly. Exact unlearning can be defined in relation to retraining, while approximate methods seek comparable behavioral or statistical properties at substantially lower cost. The literature also contains important questions about how successful unlearning should actually be verified.

The difficulty is not merely computational. There is a conceptual problem. Suppose a model learned that a particular person was alive. It may also have learned thousands of facts about that person, relationships involving that person, language patterns associated with references to that person, and general concepts partly learned through examples involving that person. Removing the person's influence from the network without damaging unrelated knowledge can be much harder than identifying the original training examples.

Some unlearning methods therefore deliberately structure training so that future removal becomes easier. SISA, for example, partitions training data into shards and trains separate components so that removing a sample can require retraining only the affected component rather than the entire model. Other approaches attempt to alter an already trained model. These approaches illustrate a broader engineering principle: if knowledge will eventually need to be removed, it may be advantageous to structure the learning process so that the provenance and influence of knowledge remain manageable.

Yet the problem considered here is broader than the conventional machine unlearning problem. Machine unlearning is often concerned with removing the influence of specified training data, frequently for privacy or regulatory reasons. An autonomous agent operating in the world would face a different problem: it would need to determine which beliefs have become obsolete because the world itself has changed.

That is not necessarily deletion.

It is revision.

And revision requires understanding the relationships between beliefs.

Part 6 - When a Fact Is More Than a Fact

Consider the statement, "Person X is alive." In isolation, this looks like a single proposition. Inside an intelligent system, however, it could function as the foundation for a much larger structure.

X is alive → X can respond → X can give instructions → plans involving X remain possible

But the dependencies could be much deeper:

X is alive → X has preferences → those preferences affect objectives → those objectives affect actions

If X dies, some of these relationships disappear. Others do not. The person's historical preferences may remain useful information. Previous instructions may remain relevant in some contexts. Memories of past interactions may remain valuable. A promise may remain meaningful even though the person can no longer respond. Some plans may need to be cancelled while others may simply require substitution of a different participant.

The correct update therefore cannot be represented as a simple erasure of the node corresponding to X.

It requires distinguishing at least three categories:

  • facts that remain true historically;
  • facts that remain true in the present world;
  • consequences that were contingent upon conditions that no longer exist.

This distinction is surprisingly close to a fundamental problem in belief revision. Formal theories of belief change have long recognized that changing one belief can require changes elsewhere in the belief state, and that the choice of which beliefs to preserve is not trivial. The literature also distinguishes between approaches concerned primarily with maintaining coherence and approaches that explicitly track the grounds or justifications of beliefs.

An artificial agent would face the same conceptual problem in a much more complicated setting. It would need to know not only what it believes but, ideally, something about why it believes it, what other conclusions depend upon it, how reliable the evidence is, and under what circumstances the belief should cease to control behavior.

This suggests that an ideal persistent memory system might require something closer to provenance than to mere storage. A memory could carry information about its source, confidence, temporal validity, dependencies, and whether it has been superseded.

memory = content + source + time + confidence + dependencies + validity

Such a structure would not eliminate the difficulty, but it would change the nature of the problem. The system would no longer have to infer everything from an undifferentiated mass of learned parameters.

Part 7 - The Possibility of Recursive Mismatch

The most intriguing possibility arises when the system is capable of explicit reasoning about the consequences of a changed fact.

Suppose an artificial agent learns that a person on whom many of its plans depend has died. The system can correctly reason that the person is no longer available to give instructions. It can then correctly infer that certain instructions will never arrive. It can infer that certain expectations about future events must therefore be revised. It can identify plans that depended upon those expectations and reconsider them as well.

There is nothing irrational about this reasoning. In fact, it may be exactly what a competent agent should do.

But the chain of consequences can become extremely large.

new fact → affected beliefs → affected plans → affected objectives → affected policies

The question then becomes: where does the revision stop?

A sufficiently sophisticated system might discover that one old assumption affected another, which affected a third, which influenced a policy, which affected the interpretation of previous observations, which in turn changes the apparent evidence supporting some other assumption. At the extreme, a system could repeatedly reconsider the consequences of the original update without arriving at a stable new representation.

This should not be described as artificial depression. There is no reason to assume that such a system would experience sadness, suffering, or anything subjectively analogous to human rumination. The useful analogy is structural: both cases involve the persistence of an old organization of expectations after the conditions that supported it have changed.

The danger in an artificial system would instead be behavioral instability, excessive deliberation, persistent pursuit of obsolete objectives, or the generation of actions that are coherent under an outdated model but inappropriate under current conditions.

There is another possibility that is perhaps even more important. The agent might not explicitly analyze the obsolete assumption at all. It might simply continue producing outputs that reflect it because the assumption is deeply embedded in its learned representations.

This would resemble the dog waiting at the door more closely than a human being consciously thinking about loss. The system would not necessarily be stuck in a verbal loop. It could be stuck in a behavioral model.

The two failure modes are therefore opposite in appearance:

  • the system may fail to revise enough, continuing to act according to an obsolete model;
  • the system may attempt to revise too much, repeatedly propagating the implications of the change without reaching a stable representation.

Both arise from the same underlying question: how should an intelligent agent determine the appropriate scope of model revision?

Part 8 - Alignment as World-Model Maintenance

This observation changes the way one might think about alignment. Alignment is often discussed in terms of objectives: how can we make an artificial system pursue what humans want rather than something else? That question is fundamental, but an agent cannot pursue an objective intelligently without possessing some model of the world in which the objective is to be achieved.

An objective therefore interacts with a model.

objective + world model → policy → action

If the world model becomes obsolete, the policy can become inappropriate even if the underlying objective has not changed. Conversely, if the system's interpretation of the objective changes while its world model remains accurate, it can still behave incorrectly.

This suggests two distinct forms of mismatch.

The first is epistemic mismatch: the system's representation of what is true differs from the current world.

The second is normative mismatch: the system's representation of what it should pursue differs from what its designers or users actually intend.

These problems can interact. An agent with an incorrect world model may infer the wrong consequences of its objective. An agent with an incorrectly specified objective may act in ways that prevent it from obtaining the information needed to correct its world model.

The second possibility is particularly important for autonomous systems. A passive model can be wrong about the world and simply produce an inaccurate answer. An agent can be wrong about the world, act on that belief, alter the world, and then observe the consequences of its own actions.

wrong model → action → altered environment → evidence → further model update

If the agent's actions reinforce the conditions assumed by its original model, the system may receive misleading confirmation. This is one reason the distinction between prediction and control matters so much for advanced artificial agents.

The alignment problem, viewed from this perspective, is not only a question of giving an agent the correct destination. It is also a question of ensuring that the agent can continue to recognize the terrain through which it is moving.

Part 9 - A Different Architecture of Memory

If the problem is fundamentally one of maintaining an evolving model, then an architecture in which every important fact is permanently entangled with a single set of neural parameters may be poorly suited to the task of lifelong adaptation. This does not imply that neural networks cannot adapt. Continual learning research has developed many techniques for balancing adaptation and retention, including regularization, replay, architectural separation, and other strategies. The stability-plasticity dilemma, however, remains a central issue.

A long-lived artificial agent might therefore benefit from a division between different kinds of knowledge. The pretrained neural model could provide broad linguistic and conceptual competence. A persistent memory system could record individual experiences. A world-state representation could describe facts believed to hold at the present time. An episodic system could preserve what happened in the past. A provenance mechanism could record why a belief is held. A revision mechanism could determine which beliefs have become obsolete.

base model + world state + episodic memory + provenance + revision mechanism

Such an architecture would not require the system to forget history merely because history is no longer current. A deceased person could remain part of the agent's historical knowledge while no longer being treated as an active participant in future plans. A previous instruction could remain stored while its applicability could become conditional or expired. A former objective could be retained as part of the system's history without continuing to govern present behavior.

This distinction between remembering and believing may ultimately prove important. An intelligent system should not necessarily erase information because the information is obsolete. It should know that the information is obsolete.

In this sense, the ideal operation is not forgetting but deactivation. The system preserves the historical record while preventing obsolete information from silently exerting present influence.

This also suggests why provenance could be more valuable than raw memory. If a system knows where a proposition came from, when it was established, what observations supported it, and what other conclusions depend on it, then revision becomes a structured operation rather than a blind attempt to modify the model's parameters.

The field of machine unlearning already demonstrates one version of this principle: training procedures can be designed to make later removal more manageable. More generally, the architecture of learning determines how difficult future revision will be.

Part 10 - The Deeper Problem

The most interesting conclusion is that intelligence may require something more than the ability to learn. It may require the ability to determine when learning has made previous knowledge invalid, and to revise that knowledge without destroying everything that was built around it.

This is a more subtle problem than ordinary information acquisition. The world does not merely provide additional facts. It occasionally changes the relationships between facts.

A person can remain the same person while losing a job. A city can remain the same city while its political system changes. A scientific theory can preserve many of its successful predictions while being superseded in its underlying interpretation. A relationship can end while the memories created within it remain true. A person can die while everything that was true about their life remains historically true.

Reality therefore contains a distinction between what was true and what is presently actionable.

An intelligent agent that cannot represent that distinction may have difficulty living in a changing world.

The dog waiting at the door is an elementary illustration of this problem. The routine was not irrational when the owner was alive. The problem is that the routine survives the condition that once justified it. The behavior has become temporally displaced.

A future artificial agent could encounter the same structural problem at a vastly greater scale. Instead of one routine, it might possess millions of interconnected assumptions about people, institutions, objectives, plans, resources, and future events. A sudden change could invalidate not one behavior but a substantial region of the model.

The central engineering challenge would then be neither remembering everything nor forgetting everything. It would be knowing what has changed, what that change affects, and what should remain untouched.

The problem of intelligence is not merely building a model of reality.
It is maintaining that model while reality changes.

This perspective also gives a different interpretation to the problem of unlearning. The deepest form of unlearning may not be the deletion of a training example from a neural network. It may be the ability to preserve an old fact as history while withdrawing the authority that fact once possessed over present inference and action.

That is closer to belief revision than deletion, closer to model maintenance than forgetting, and closer to adaptation than erasure.

For humans, the inability to perform this transition cleanly can contribute to suffering, confusion, and maladaptive behavior. For an artificial agent, the corresponding failure need not involve anything resembling subjective suffering. It could instead manifest as persistent obsolete behavior, unstable reasoning, inappropriate generalization, or actions that are internally coherent but disconnected from the present world.

The speculative question, then, is not whether an AI could become "depressed" when reality changes. There is no basis for assuming such an experience. The more technically meaningful question is whether an increasingly autonomous intelligence could become trapped between two incompatible states: the world that generated its model and the world in which that model must now operate.

If the answer is yes, then continual learning, machine unlearning, memory architecture, belief revision, uncertainty estimation, and alignment are not entirely separate problems. They become different aspects of a single deeper problem:

How does an intelligent system change its mind without losing itself?

That question may ultimately be as important for artificial intelligence as the ability to learn in the first place.

References

  1. Friston, K. (2010). “The free-energy principle: a unified brain theory?” Nature Reviews Neuroscience, 11, 127–138. ๐Ÿ”—
  2. Alchourrรณn, C. E., Gรคrdenfors, P., & Makinson, D. (1985). “On the Logic of Theory Change: Partial Meet Contraction and Revision Functions.” Journal of Symbolic Logic, 50(2), 510–530. ๐Ÿ”—
  3. French, R. M. (1999). “Catastrophic forgetting in connectionist networks.” Trends in Cognitive Sciences, 3(4), 128–135. ๐Ÿ”—
  4. Kirkpatrick, J., Pascanu, R., Rabinowitz, N., et al. (2017). “Overcoming catastrophic forgetting in neural networks.” Proceedings of the National Academy of Sciences, 114(13), 3521–3526. ๐Ÿ”—
  5. Bourtoule, L., Chandrasekaran, V., Choquette-Choo, C. A., et al. (2021). “Machine Unlearning.” 2021 IEEE Symposium on Security and Privacy. ๐Ÿ”—
  6. Wang, W., Tian, Z., Zhang, C., & Yu, S. (2024). “Machine Unlearning: A Comprehensive Survey.” arXiv. ๐Ÿ”—
  7. Verwimp, E., Aljundi, R., Ben-David, S., et al. (2023). “Continual Learning: Applications and the Road Forward.” arXiv. ๐Ÿ”—
  8. Tran, Q.-T., Le-Khac, N.-A., & Bertolotto, M. (2025). “Concept drift detection in image data stream: a survey on current literature, limitations and future directions.” Artificial Intelligence Review, 59, Article 33. ๐Ÿ”—
  9. Thudi, A., Jia, C., Shumailov, I., & Papernot, N. (2022). “On the Necessity of Auditable Algorithmic Definitions for Machine Unlearning.” 31st USENIX Security Symposium. ๐Ÿ”—
  10. Lukats, D., Zielinski, O., Hahn, A., et al. (2024). “A benchmark and survey of fully unsupervised concept drift detectors on real-world data streams.” International Journal of Data Science and Analytics, 19, 1–31. ๐Ÿ”—