Intelligence Without Continuity

Intelligence Without Continuity

Contents

  1. The Strange Programmer That Cannot Run Its Programs
  2. The Difference Between Memory and Working Memory
  3. Compression as the Substance of Human Working Memory
  4. Thinking at Many Levels at Once
  5. The Missing Experience of Procedure
  6. The Problem of Time and the Economics of Attention
  7. The Rabbit Hole
  8. Why Execution Changes Intelligence
  9. Toward a Persistent Hierarchical Mind
  10. The Strange Future of Programming
  11. References

Part 1 - The Strange Programmer That Cannot Run Its Programs

There is something deeply strange about the current generation of large language models when viewed through the eyes of a programmer. They can read enormous quantities of source code, explain unfamiliar architectures, recognize subtle bugs, generate implementations in languages they have never personally executed, reconstruct algorithms from incomplete descriptions, and move effortlessly between levels of abstraction that would require years of experience for a human developer to navigate. In many ordinary programming tasks, their ability to produce plausible and often excellent code is already remarkable.

Yet the same system may be unable to perform one of the most ordinary acts in software engineering: write the program, compile it, observe what happened, and continue from the evidence. It may know almost everything required to fix a bug while lacking the direct mechanism by which a programmer normally discovers that the bug exists. The human operator becomes an external sensory organ, running the program and returning the result to the model. The model is then asked to reason about the observation and propose another intervention.

This creates an unusual division of labor. The machine may possess extraordinary capability at symbolic manipulation while remaining dependent on a human for the closed loop that gives programming its empirical character. The programmer does not merely write code. The programmer acts upon a world, observes its response, updates a model of that world, decides what deserves attention, and acts again.

The distinction matters because intelligence in practical work is not exhausted by the ability to produce a correct local answer. A programmer who can write excellent functions but cannot determine which problem deserves attention, cannot remember why previous approaches failed, cannot notice that an optimization has become irrelevant, and cannot decide when to stop working on a detail is not equivalent to an experienced engineer. The missing ingredient is not necessarily raw reasoning ability. It is continuity.

The most interesting limitations of contemporary language models may therefore not be understood as simple deficiencies in intelligence. They may instead be limitations in the organization of intelligence: how information is compressed, maintained, retrieved, revised, and connected across time and across levels of abstraction.

Part 2 - The Difference Between Memory and Working Memory

It is tempting to describe the limitation as a problem of memory. Large language models have context windows containing hundreds of thousands, and in some systems vastly more, tokens. A human programmer cannot consciously retain anything remotely comparable to such a quantity of textual information. On a superficial comparison, therefore, the machine appears to have the obvious advantage.

But raw storage is not the same thing as working memory. A database can contain millions of facts without possessing anything resembling a human train of thought. The important question is not how much information is available, but how information is represented while an activity is underway.

A human programmer does not normally hold the source code of an entire project in consciousness. Instead, the programmer carries a compact and continuously updated model of the project. A particular subsystem may be represented as something like "the asynchronous cache invalidation path," even though that phrase stands for thousands of lines of implementation, a collection of design decisions, several known invariants, historical bugs, and expectations about how the subsystem behaves.

That representation can be expanded when necessary. The programmer can move from an architectural idea to a component, from a component to a function, from a function to a particular branch, and from the branch to the exact expression under examination. Just as importantly, the programmer can move in the opposite direction. A tiny implementation detail can suddenly cause a reconsideration of the component design, the product requirement, or even the business assumption that motivated the feature in the first place.

This makes human working memory hierarchical and addressable. It is not merely a temporary container of facts. It is a structure of compressed representations that can be selectively decompressed.

This distinction is important because simply increasing the context window of a language model does not necessarily reproduce it. A model can be given every relevant file and every previous conversation while still failing to maintain a stable, appropriately abstract representation of what matters. More context can even make the problem worse when the model must repeatedly reconstruct the important structure from a large undifferentiated mass of information.

The problem is therefore not adequately described as "the model cannot remember enough." In some circumstances the model has access to vastly more information than the human. The deeper problem is that it does not necessarily possess a human-like mechanism for deciding what should remain actively represented, at what level of abstraction it should be represented, and how that representation should be expanded again when circumstances require it.

Part 3 - Compression as the Substance of Human Working Memory

Human expertise is in large part a technology of compression. An experienced programmer does not necessarily perceive more details than a beginner. In many situations the opposite is true. The expert has learned which details can be discarded because they are consequences of deeper structures already understood.

Consider the difference between reading a piece of unfamiliar code and reading a piece of code belonging to a system one has maintained for years. The text on the screen may be identical, but the experienced programmer sees additional structure. A ten-line function may immediately evoke assumptions about its caller, the invariants of the data structure, the reason it was written in its present form, the sort of failures likely to occur around it, and the parts of the system that would be affected by changing it.

Much of that knowledge does not need to be explicitly recalled at every moment. It is latent inside compressed representations. A high-level concept acts as an index into a much larger body of knowledge. The programmer can remain at the high level until a reason appears to descend into greater detail.

This gives human thought an unusual economy. A person can carry a very large effective model while maintaining only a small amount of explicit material in conscious attention. The rest exists as structured potential, ready to be reconstructed.

This may explain why a human can appear to have an enormous working memory despite the well-known limitations of conscious attention. The apparent memory capacity comes not from retaining every constituent fact, but from retaining the right abstractions and the relationships between them.

The distinction can be expressed informally as follows. Raw memory answers the question, "How much information is available?" Hierarchical working memory answers a different question: "How much structure can I keep immediately available, and how cheaply can I recover the detail represented by that structure?"

For software engineering, the second question may be considerably more important than the first.

Part 4 - Thinking at Many Levels at Once

The hierarchy becomes even more significant because human thought does not operate at only one level of abstraction. During an ordinary development session, a programmer may move continuously between implementation, component architecture, product behavior, user experience, organizational constraints, business priorities, and long-term strategy.

Imagine a programmer investigating a small performance problem. At the lowest level, the question might concern the number of allocations made by a particular function. At the component level, it might concern the design of a cache. At the product level, it might concern whether users actually notice the latency. At the business level, it might concern whether improving that latency could make a feature commercially distinctive. At the organizational level, it might concern whether the complexity introduced by the optimization will increase maintenance costs for the engineering team.

These are not separate conversations. They are different resolutions of the same activity.

A small technical discovery can propagate upward through the hierarchy. A seemingly trivial optimization may reveal that an architectural assumption is wrong. An architectural discovery may alter the user experience. A user experience change may affect product positioning. Conversely, a business constraint can propagate downward and make a technically elegant solution irrelevant because it solves a problem that no longer matters.

The important property is therefore bidirectionality. Human cognition can descend into implementation detail while retaining some representation of the higher-level purpose that gives the detail meaning. It can also climb upward again without necessarily reconstructing the entire lower-level state from scratch.

A language model can certainly discuss all of these levels. It can write a marketing plan, design an API, optimize a function, and explain an algorithm. The harder question is whether those representations are continuously maintained as parts of one evolving internal model while the project develops.

This distinction becomes especially important over long tasks. If the model must repeatedly rediscover the relation between a local technical decision and the broader purpose of the project, then the enormous amount of information available to it does not necessarily produce the continuity enjoyed by an experienced human developer.

local implementation ↔ component ↔ architecture ↔ product ↔ organization ↔ strategy

The human mind can move through this hierarchy opportunistically. A language model, by contrast, is usually asked to produce a response from a context. The difference between these two arrangements may prove more important than the difference between their respective quantities of stored information.

Part 5 - The Missing Experience of Procedure

There is another limitation that follows naturally from this picture: the difference between knowing facts about programming and possessing experience of doing programming.

A modern language model has been exposed to an extraordinary quantity of programming knowledge. It has seen implementations, documentation, tutorials, bug reports, discussions, design patterns, algorithms, code reviews, and countless examples of programmers explaining how something works. It can therefore possess a very rich statistical model of programming practice.

But a large amount of programming expertise consists of procedural judgments that are difficult to reduce to isolated facts. An experienced programmer knows that some bugs should be attacked immediately while others should be left alone until more information is available. They know that an ugly implementation is sometimes preferable to a premature abstraction. They know when a test is worth writing and when an exploratory change is sufficient. They know when a debugging hypothesis has become unproductive.

These judgments are learned through trajectories rather than snapshots. They come from having pursued wrong explanations, watched projects become complicated, missed deadlines, maintained code written years earlier, discovered that a supposedly important optimization had no measurable effect, and experienced the consequences of making the wrong decision at the wrong time.

One might call this a procedural training set: not merely examples of correct answers, but histories of action and consequence.

The distinction is subtle. A model can know that premature optimization is usually undesirable. It can even explain why. But knowing the proposition is not identical to possessing the intuition that says, in the middle of a real project, "I have already spent too much time on this."

The latter judgment depends on accumulated experience with trajectories. It depends on knowing what tends to happen next.

This is one reason human expertise can appear irrational when reduced to rules. An experienced engineer may abandon an investigation without being able to produce a formal proof that the investigation is unproductive. The decision is based on a compressed model of many previous situations. Experience has become a policy.

Part 6 - The Problem of Time and the Economics of Attention

Time introduces another dimension that is easy to overlook when thinking about language-model intelligence. Human activity takes place under a continuously experienced budget. There is always some implicit answer to the question: "Is this worth doing now?"

The answer depends not only on the intrinsic value of an action but also on what else could be done with the same time. An engineer debugging a small performance regression is therefore not solving an isolated optimization problem. The engineer is choosing how to allocate scarce attention among competing possibilities.

Perhaps the regression can be fixed in an hour. But during that hour, another feature could have been completed, an architectural problem could have been investigated, tests could have been written, or the engineer could simply have stopped and returned to the problem after more evidence became available.

Humans experience this opportunity cost continuously. Deadlines, fatigue, boredom, anticipation, curiosity, social expectations, and memories of previous projects all contribute to the perceived value of continuing or abandoning a line of work.

A language model operating in a conventional inference setting has no equivalent experience of elapsed time. It does not become tired because it has spent three hours examining an obscure function. It does not become anxious because a release deadline is approaching. It does not feel that the afternoon has been consumed by an investigation that should have taken fifteen minutes.

These observations should not be confused with claims about consciousness or emotion. The relevant issue is functional. Human motivation contains mechanisms that continuously transform time, effort, and opportunity cost into decisions about what to do next. A conventional model does not automatically possess an equivalent mechanism merely because it can reason about time when asked.

This distinction creates a potentially serious failure mode. A highly capable model can optimize the objective currently in front of it while neglecting the larger allocation problem in which that objective is embedded.

The result can be an agent that is extraordinarily good at solving the wrong problem.

Part 7 - The Rabbit Hole

Consider a programmer who discovers a two percent performance regression. An intelligent but poorly managed coding agent might identify an interesting interaction between allocation behavior, cache locality, and a particular data structure. It could formulate increasingly sophisticated hypotheses, inspect more code, develop benchmarks, derive alternative implementations, and eventually discover a genuinely elegant optimization.

The work might also have been completely irrational.

Perhaps the regression is invisible to users. Perhaps the relevant code is scheduled for replacement next month. Perhaps the optimization introduces significant maintenance complexity. Perhaps the same engineering time would have produced a feature worth far more to the product.

The model's local reasoning can be impeccable while its global behavior is poor.

This is a particularly interesting limitation because increasing raw reasoning ability does not necessarily eliminate it. A more capable model may become better at exploring a rabbit hole. It may generate more convincing hypotheses, conduct more elaborate analyses, and discover increasingly subtle solutions, while still failing to ask whether the rabbit hole deserves exploration.

The problem is therefore not simply optimization. It is meta-optimization: deciding what deserves optimization in the first place.

Human programmers often develop this ability implicitly. After years of experience, they acquire a sense for the shape of a productive investigation. They recognize when a promising clue is becoming a distraction. They learn to accept temporary imperfection. They develop thresholds for when to benchmark, when to refactor, when to ask another person, and when to ship.

Such behavior can be understood as a policy over workflows rather than a collection of programming facts.

solve the problem → evaluate the value of solving it → decide what deserves attention next

The third step is easy to underestimate because it is rarely visible in the finished artifact. Yet it may account for a substantial fraction of the value of human expertise.

Part 8 - Why Execution Changes Intelligence

The inability to execute code is therefore more profound than an inconvenience. Execution supplies an external source of reality against which internal reasoning can be corrected.

A programmer who writes code without running it is forced to reason about what the program will do. A programmer who runs the code receives evidence. The difference is enormous. The compiler catches syntactic mistakes. Tests reveal violated assumptions. Profilers expose bottlenecks. Logs reveal unexpected states. Production behavior reveals interactions that no static analysis anticipated.

The feedback loop can be represented simply:

hypothesis → implementation → execution → observation → revised hypothesis

Without the observation step, intelligence is forced to operate against an imagined world. It can still be remarkably effective because programming languages and software architectures are highly structured domains. But every unobserved assumption creates an opportunity for divergence between the model and reality.

More importantly, execution provides memory. The failed test is not merely a negative result. It is an externally stored fact that narrows the space of possible explanations. A benchmark becomes a durable constraint on future reasoning. A crash log records a state that would otherwise have to be reconstructed mentally.

This means that tools do not merely extend the model's capabilities. They can change the character of its reasoning. A compiler, debugger, profiler, repository, issue tracker, test suite, and deployment environment together form an external cognitive system.

An autonomous programmer would therefore not simply be an LLM with access to a terminal. It would be an intelligence participating in a persistent environment that continually produces evidence, stores consequences, and changes the space of future decisions.

Part 9 - Toward a Persistent Hierarchical Mind

If the limitations described above are manifestations of a common problem, then simply extending context windows may not be the ultimate solution. What may be needed is a persistent hierarchical state that is continuously revised as work progresses.

Such a system might maintain multiple representations of a project simultaneously. At the highest level it would retain the purpose and constraints of the project. Below that would be product behavior, architecture, components, implementation details, active hypotheses, known invariants, and the immediate task.

Each level would be compressed, but each would provide pathways into more detailed representations. A high-level product objective could constrain a technical decision. A low-level implementation discovery could invalidate an architectural assumption. A new benchmark could alter the priority of an entire feature.

The state might conceptually resemble:

goal ↔ product model ↔ architecture ↔ subsystem ↔ implementation ↔ current detail

But the hierarchy alone would not be sufficient. The system would also need temporal state: what has been attempted, what failed, what remains uncertain, why a decision was made, how much effort has already been spent, and which questions have become less important as the project evolved.

A useful persistent state might therefore contain not only facts but also uncertainty and history. It could distinguish established facts from working hypotheses, record abandoned approaches, preserve the reasons behind important decisions, and track the expected value of further investigation.

Most importantly, such a system would need to decide at what level to think. If the immediate task concerns a single expression, the model should be able to descend into implementation detail. If that expression reveals an architectural problem, the model should be able to climb upward. If the architectural problem has implications for the product, it should be able to climb further without losing its place in the original investigation.

This is closer to a navigable hierarchy of representations than to a larger context window.

The distinction may ultimately prove fundamental. A context window answers the question of how much information can be presented to the model. A persistent hierarchical memory answers the question of what the model believes the information means, how that meaning is organized, and how it should be reconstructed when needed.

Part 10 - The Strange Future of Programming

It is possible that the current generation of language models will eventually be remembered as having occupied an unusual intermediate stage. They already possess extraordinary abilities in local reasoning and code generation, yet they lack several mechanisms that make human intelligence effective over long periods of real activity.

This should not be interpreted as evidence that human programmers are simply "smarter" in every relevant sense. The comparison is more interesting than that. The machine can have enormous advantages in knowledge retrieval, symbolic manipulation, consistency, and rapid generation while the human retains advantages in continuity, prioritization, embodied time, compressed experience, and the ability to move naturally among multiple levels of a changing situation.

The resulting difference is not well captured by the phrase "short context." A human can forget the exact syntax of a function while retaining an extremely rich model of why that function exists. A language model can be given the exact syntax while lacking the same persistent representation of why the function matters now.

Nor is the difference well captured by saying that humans have more memory. Humans may have vastly less raw accessible information, but their memories are organized around abstractions that can be expanded and revised as needed. They are not merely storing a project. They are maintaining a compressed model of the project's evolving significance.

Nor is the difference simply motivation. A programmer's experience of time, effort, opportunity cost, boredom, urgency, and anticipated consequences forms part of the control system that decides what happens next. An intelligence without those mechanisms can remain remarkably capable while making profoundly poor decisions about where to spend its capability.

The central challenge, then, may be to transform an intelligence that produces excellent local continuations into one that can maintain a coherent trajectory through a world.

Such an intelligence would need to remember not only what is true, but what it has been trying to accomplish. It would need to remember not only what it tried, but why it tried it and what that result changed. It would need to represent the same project at several levels simultaneously and move between those levels without losing continuity. It would need to learn not only solutions, but the economics of pursuing solutions. It would need to recognize when a local improvement is globally irrelevant and when an apparently tiny detail has implications several levels above it.

Most of all, it would need to become capable of treating time as part of the problem rather than merely as an external parameter.

The deepest difference may therefore be summarized not as memory versus no memory, or intelligence versus no intelligence, but as continuity of organized thought. Human beings carry forward a changing compressed model of what they are doing. That model can be refined, expanded, abandoned, reconstructed, and connected to other models while activity continues. The current language-model paradigm is extraordinarily good at generating thought from a presented state, but it is much less naturally equipped to preserve and manage the evolving state itself.

If that limitation can be overcome, the consequences for programming will be considerable. The decisive advance may not be a model that writes code more beautifully than today's models. It may be a system that knows which code to write, which problem to postpone, which experiment to perform, which hypothesis to abandon, which detail to compress, which abstraction to revisit, and when the work is finally good enough.

At that point, the distinction between an AI that can program and an AI that can do programming will begin to disappear.

References

  1. Miller, G. A. (1956). “The Magical Number Seven, Plus or Minus Two: Some Limits on Our Capacity for Processing Information.” Psychological Review, 63(2), 81–97. 🔗
  2. Baddeley, A. D., & Hitch, G. (1974). “Working Memory.” In G. A. Bower (Ed.), The Psychology of Learning and Motivation, Vol. 8, 47–89. 🔗
  3. Ericsson, K. A., & Kintsch, W. (1995). “Long-Term Working Memory.” Psychological Review, 102(2), 211–245. 🔗
  4. Newell, A., & Simon, H. A. (1972). Human Problem Solving. Prentice-Hall.
  5. Hutchins, E. (1995). Cognition in the Wild. MIT Press. 🔗
  6. Clark, A., & Chalmers, D. (1998). “The Extended Mind.” Analysis, 58(1), 7–19. 🔗
  7. Vaswani, A., et al. (2017). “Attention Is All You Need.” Advances in Neural Information Processing Systems, 30. 🔗
  8. Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction, 2nd ed. MIT Press. 🔗