Intelligence and the Limits of Context: What Language Models and Sleep-Deprived Minds Reveal About Thinking

Intelligence and the Limits of Context: What Language Models and Sleep-Deprived Minds Reveal About Thinking

Contents

  1. The Paradox of Increasing Capacity
  2. When More Information Produces Less Understanding
  3. The Human Mind Under Cognitive Strain
  4. The Architecture of Effective Intelligence
  5. The Hidden Cost of Accumulation
  6. Compression, Abstraction, and Forgetting
  7. Sleep as a Condition for Cognitive Renewal
  8. The Difference Between Knowing and Understanding
  9. Toward a Theory of Effective Intelligence
  10. The Importance of Knowing What Matters
  11. References

Part 1: The Paradox of Increasing Capacity

One of the more curious observations about large language models is that their performance can deteriorate as the amount of information available to them increases. A model may be given a lengthy conversation containing everything necessary to solve a problem, including definitions, constraints, previous conclusions, and explicit instructions, yet produce an answer that is less coherent or less accurate than it would have produced from a much shorter prompt. The relevant information has not necessarily disappeared. It remains present in the context, sometimes in precise and unambiguous form. Nevertheless, the model may fail to use it effectively.

This phenomenon presents a peculiar paradox. If additional information is potentially useful, and if the model has been given access to that information, why should the availability of more context sometimes make the model perform worse? The intuitive expectation is that a larger information budget should expand the possibilities of reasoning. It should allow the model to preserve more details, connect ideas across longer conversations, and solve problems that would otherwise exceed its reach. Yet practical performance does not always follow this expectation. The quantity of information available to a system and its ability to reason with that information are not equivalent.

The same distinction appears in human cognition. A person can possess extensive knowledge and still be unable to apply it effectively under conditions of fatigue, distraction, or mental overload. An experienced physician can overlook an important symptom after a sleepless night. A mathematician can lose track of an assumption in a proof that would normally be straightforward. A programmer can spend an unreasonable amount of time searching for an error whose location would have been obvious under better cognitive conditions. In each case, the knowledge necessary to perform the task may remain intact, while the ability to coordinate and apply it deteriorates.

This suggests a broader question about the nature of intelligence. Perhaps intelligence cannot be adequately described by the quantity of information a system stores, the number of facts it can retrieve, or even the complexity of the problems it can solve under ideal conditions. Perhaps we must also consider the mechanisms by which information is selected, organized, prioritized, and integrated into a coherent representation of the problem at hand.

The comparison between large language models and the human mind should not be taken literally. Neural networks and biological brains operate through substantially different mechanisms, and the limitations of one cannot simply be assumed to explain those of the other. Nevertheless, their contrasting forms of cognitive failure provide an illuminating starting point. Both invite us to distinguish the possession of information from the effective use of information, and both suggest that increasing the amount of available material does not necessarily increase the quality of thought.

The central thesis of this essay is that effective intelligence depends not only on informational capacity but also on the organization and control of cognitive processes. From this perspective, the deterioration of a language model under an excessively demanding context and the deterioration of human reasoning under sleep deprivation are not identical phenomena, but they exhibit a structurally similar problem: the resources required to manage information may become inadequate relative to the demands placed upon the system.

Part 2: When More Information Produces Less Understanding

A context window determines how much text a language model can process as part of a given interaction. Increasing this window allows the model to consider longer documents, preserve more conversational history, and examine relationships among a greater number of statements. It is tempting to regard this capacity as a direct measure of the model's ability to reason over long sequences of information. However, a larger context window establishes only how much material can be made available to the model. It does not guarantee that every relevant relationship within that material will be identified or used correctly.

One source of difficulty is the competition between relevant and irrelevant information. A long conversation may contain the central facts of a problem alongside repeated explanations, abandoned hypotheses, digressions, contradictory instructions, and details that were important earlier but no longer matter. Although these elements may occupy very different positions in the logical structure of the conversation, they all become part of the material presented to the model. The result can be a context in which the signal remains present but becomes harder to distinguish from the surrounding noise.

Another difficulty arises from the distribution of relevant information. A critical condition may appear near the beginning of a conversation, while its implications become apparent only much later. To answer correctly, the model must preserve that condition, recognize its continued relevance, and integrate it with subsequent information. The problem is not necessarily that the original statement has been forgotten in the ordinary human sense. Rather, the model may fail to give the statement sufficient influence during the computation that produces the answer.

Research on long-context language models has documented variations in performance depending on where relevant information appears within an input. In some settings, models perform better when useful information appears near the beginning or end of a context than when it is located in the middle. This behavior has been described as the "lost in the middle" phenomenon. It illustrates an important distinction between providing information to a model and ensuring that the information can be used reliably across different positions and configurations.

The difficulty becomes more pronounced when a conversation contains competing interpretations of the same issue. Suppose a model initially considers three possible explanations for a problem, later rejects two of them, and eventually reaches a conclusion that depends on a subtle distinction between the remaining alternatives. If the conversation preserves all the earlier material without clearly marking which ideas have been discarded, the model may continue to be influenced by possibilities that are no longer relevant. The accumulation of intermediate reasoning can become a liability rather than an asset.

This does not mean that every long context causes deterioration, or that longer contexts are inherently undesirable. Additional information can substantially improve performance when it supplies necessary evidence, clarifies constraints, or enables connections that would otherwise be impossible. The difficulty arises when the marginal value of additional information is outweighed by the demands of identifying and coordinating what matters. Context length is therefore not simply a measure of informational power. It is also a source of complexity that the model must manage.

The analogy with a library is useful here. Imagine a library in which every relevant book is available, but in which the catalog becomes increasingly unreliable as the collection expands. The presence of a book does not ensure that a researcher will find it, recognize its relevance, or correctly integrate its contents with other sources. A larger collection expands the range of possible discoveries, but it also increases the importance of indexing, selection, and synthesis. Without those functions, abundance can produce confusion rather than understanding.

The implication is that a sufficiently large context window cannot, by itself, solve the problem of long-range reasoning. Increasing capacity may postpone certain limitations, but it does not eliminate the need for effective information management. A system must not only accommodate information; it must also determine which information deserves to influence its next step.

Part 3: The Human Mind Under Cognitive Strain

Human cognition presents a related but biologically distinct example of the difference between possessing information and using it effectively. People do not process everything they know with equal intensity at every moment. Conscious reasoning depends on limited capacities for attention, working memory, and executive control, while the broader brain continually performs many processes outside conscious awareness. The ability to solve a problem therefore depends not only on what a person knows but also on whether the relevant knowledge can be maintained, retrieved, and manipulated under the prevailing cognitive conditions.

Working memory allows us to maintain and manipulate information over short periods. It enables us to compare alternatives, follow the steps of an argument, remember the conditions of a problem while evaluating a proposed solution, and combine separate pieces of information into a coherent whole. Its limitations help explain why tasks that are simple in isolation can become difficult when they must be performed simultaneously. A person may understand every individual step of a procedure yet struggle to complete it when too many intermediate results must be held in mind at once.

Executive control contributes another essential function. It helps maintain goals, resist distractions, switch between tasks when appropriate, and inhibit responses that are irrelevant or premature. These functions are crucial to reasoning because intelligent problem solving requires more than the activation of useful knowledge. It requires the ability to organize that knowledge around a goal while preventing competing impulses and irrelevant associations from dominating the process.

Sleep deprivation can compromise these functions. Attention becomes less stable, reaction times become more variable, working memory becomes less reliable, and the ability to monitor one's own performance may deteriorate. A sleep-deprived person can experience brief lapses in attention even while attempting to remain engaged with a task. Errors may emerge not because the person lacks the necessary knowledge but because the cognitive processes required to apply that knowledge consistently are no longer functioning at their usual level.

Consider a mathematician working through a difficult proof. The proof may depend on a condition introduced several lines earlier, and its correctness may require maintaining a sequence of relationships across many intermediate steps. Under normal conditions, the mathematician can keep track of these dependencies, recognize contradictions, and revise an argument when necessary. After substantial sleep loss, the same person may overlook a condition, repeat a failed line of reasoning, or accept a conclusion without adequately checking whether it follows from the premises. The underlying mathematical knowledge may remain largely intact, but the coordination required to use it has become less reliable.

A similar effect occurs in everyday decision-making. Fatigue can make it more difficult to evaluate competing possibilities, anticipate consequences, regulate emotional responses, and distinguish important details from incidental ones. The resulting behavior may appear unintelligent, even when the individual possesses considerable experience and expertise. What has deteriorated is not necessarily the person's general intellectual capacity but the conditions under which that capacity can be expressed.

This distinction matters because intelligence is often judged through observable performance. When someone reasons poorly, we tend to infer that they do not know enough, have insufficient ability, or have failed to understand the problem. Yet performance is also sensitive to the state of the system performing the task. The same individual can produce markedly different results depending on fatigue, stress, distraction, motivation, and the complexity of the immediate demands.

Sleep deprivation therefore provides a powerful illustration of a general principle: the effective use of knowledge depends on the integrity of the processes that regulate attention, memory, and reasoning. However, it would be misleading to conclude that the human brain becomes impaired simply because it has accumulated too much information. Sleep loss affects multiple biological and cognitive processes, and the total quantity of stored memories is not equivalent to the length of a language model's context. The analogy is most useful at the level of function rather than mechanism.

Part 4: The Architecture of Effective Intelligence

The comparison between long-context language models and sleep-deprived humans becomes clearer when we distinguish three aspects of cognition: stored information, active processing, and coordination. These categories are not a complete theory of intelligence, but they provide a useful framework for understanding why information alone is insufficient.

The first aspect is stored information. In a language model, this includes the information available in the current context, together with patterns represented in the model's learned parameters. In a human being, it includes long-term memories, acquired knowledge, learned skills, and accumulated experience. Stored information establishes the resources that may be drawn upon, but it does not determine whether those resources will be accessible or useful in a particular situation.

The second aspect is active processing capacity. A language model must compute relationships among tokens, representations, and contextual elements within the limits of its architecture and available computational resources. A human being must maintain and manipulate information through working memory and related cognitive processes. In neither case should the relevant capacity be confused with the total amount of stored information. A system can possess an extensive body of knowledge while having a much more limited ability to manipulate all the relevant elements simultaneously.

The third aspect is coordination and control. The system must determine which information is relevant, how separate elements relate to one another, which hypotheses remain plausible, and which conclusions follow from the evidence. In humans, executive functions and attentional processes contribute to this organization. In language models, the interaction of attention mechanisms, learned representations, contextual cues, and inference procedures contributes to the way information influences the generated response. The mechanisms differ, but both cases demonstrate that the organization of processing matters as much as the availability of material.

These three aspects can be represented conceptually as follows:

Stored Information → Active Processing → Coordination and Control → Effective Performance

This sequence should not be interpreted as a literal description of a biological or computational pipeline. It is a conceptual model that emphasizes the difference between possessing resources and using them. In practice, the processes interact continuously. Attention influences what information is processed; processing changes which relationships become salient; and the resulting interpretation can alter what the system considers relevant.

The model also reveals why increasing any single capacity may fail to produce proportional improvements in performance. More stored information is of limited value if the system cannot retrieve the relevant parts. Greater processing capacity may not help if irrelevant information continually captures attention. Improved retrieval may still produce poor conclusions if the system cannot distinguish established facts from abandoned hypotheses. Effective intelligence depends on the interaction among these functions, not merely on maximizing one of them.

This helps explain why the pursuit of ever-larger context windows is only one possible route toward better language models. Other improvements may involve retrieving relevant information selectively, summarizing previous interactions, separating active constraints from historical discussion, maintaining explicit representations of goals, and checking whether a proposed answer remains consistent with established facts. Such techniques do not simply increase the amount of information available. They attempt to improve the structure through which information becomes useful.

The human equivalent is not to eliminate knowledge or artificially restrict thought, but to create conditions in which the mind can use its resources effectively. Adequate sleep, focused attention, deliberate prioritization, and the external organization of complex tasks can reduce the demands placed on working memory and executive control. In both cases, the goal is to make the relevant information easier to identify and integrate.

Part 5: The Hidden Cost of Accumulation

Modern information systems are often designed around the assumption that more information is better. Larger databases, longer records, expanded memory, and more comprehensive archives are treated as unambiguous improvements because they preserve options that might otherwise be lost. There is considerable truth in this assumption. Information that has been discarded cannot always be recovered, and a system deprived of essential context may produce shallow or incorrect conclusions. Nevertheless, accumulation has costs that become visible when the system must act on what it has collected.

Every additional piece of information introduces the possibility of another relationship that might matter. Some details reinforce existing conclusions, while others contradict them, qualify them, or make them irrelevant. A long record can contain statements that were accurate at different moments but cannot all be treated as simultaneously valid. The system must distinguish between current and outdated information, between facts and speculation, between decisions and possibilities, and between conclusions that remain authoritative and those that have been superseded.

This creates a distinction between informational complexity and informational usefulness. A document can become longer without becoming more informative in any practical sense. Repeated statements may increase its size while adding little new content. A conversation may preserve every intermediate idea yet obscure the final decision. A collection of observations may contain abundant detail without making the underlying pattern easier to identify. Accumulation increases the available material, but organization determines how much of that material contributes to understanding.

Human beings encounter a related problem in the form of cognitive overload. A person managing multiple projects may continually revisit unfinished tasks, remember unresolved decisions, anticipate future obligations, and respond to new demands. These concerns do not necessarily occupy working memory continuously, but they can compete for attention and reappear when circumstances trigger them. The resulting experience is often described as mental clutter, a condition in which too many competing demands make it difficult to establish a clear priority.

The important distinction is that mental clutter is not simply the possession of many memories. Human memory is distributed across different systems, and much of what we know remains inactive until it becomes relevant. Cognitive overload arises more directly when demands on attention, working memory, and control exceed what can be managed effectively in a particular situation. The difficulty is therefore less about how much the mind contains in total than about how much it must coordinate under current conditions.

Long-context language models face a different but related challenge. They may be exposed to large quantities of text that contain a mixture of active instructions, historical discussion, redundant details, and conflicting statements. The presence of all this material can complicate the task of identifying what should govern the current response. In this sense, a poorly organized history can become a liability even when every sentence is individually understandable.

The lesson is not that information should be minimized indiscriminately. Rather, accumulation must be accompanied by mechanisms that establish hierarchy and relevance. Without them, the system inherits the burden of repeatedly deciding which parts of its history matter. What appears to be an increase in memory can become an increase in the work required to use memory.

Part 6: Compression, Abstraction, and Forgetting

If accumulation can impair the effective use of information, an obvious response is to compress it. In the context of language models, summarization can transform a lengthy conversation into a shorter representation containing the essential facts, current objectives, important constraints, and conclusions that remain valid. Instead of repeatedly processing every abandoned possibility and explanatory detour, the model can work from a more organized account of the conversation's present state.

A good summary is not merely a shorter version of the original text. It is a selective representation of the information most likely to matter in future reasoning. It distinguishes enduring facts from temporary observations, final decisions from tentative proposals, and active requirements from historical details. Ideally, it also preserves uncertainty and unresolved questions, preventing an ambiguous suggestion from becoming an apparently established fact through compression.

This process resembles abstraction, one of the central operations of intelligent thought. Abstraction allows a system to preserve a pattern while ignoring details that do not affect the problem being solved. A physicist does not need to represent every microscopic interaction when constructing a useful model of planetary motion. A mathematician can replace a long sequence of examples with a general theorem. A skilled chess player does not evaluate every piece independently but recognizes familiar configurations and their strategic significance.

Abstraction makes complex reasoning possible because it reduces the number of details that must be handled explicitly. It allows a system to represent a large body of experience through a smaller set of relationships that retain explanatory or predictive value. Rather than repeatedly reconstructing a problem from every individual observation, the thinker can operate on a compact representation of its structure.

Yet compression creates a difficult trade-off. The process that removes irrelevant information may also eliminate a detail that later turns out to be essential. A summary that preserves a conclusion but omits the assumptions supporting it can make subsequent reasoning brittle. A model that compresses a conversation too aggressively may lose a seemingly minor qualification that determines whether a recommendation is appropriate. A human being who forgets the circumstances surrounding an experience may retain the broad lesson while losing the context needed to apply it correctly.

The challenge is therefore not compression at any cost, but selective preservation. Effective compression must retain the information that governs future decisions while discarding details whose contribution is negligible. It must preserve relationships, not merely isolated facts, because the meaning of a fact often depends on the conditions under which it was established.

Human memory offers a particularly interesting parallel. We do not ordinarily retain a perfect record of every experience. Memory is reconstructive, selective, and influenced by subsequent knowledge and context. Forgetting can make room for generalization by preventing every incidental detail from remaining equally prominent. It can help us extract patterns that apply beyond a single event. However, forgetting is not always beneficial, and human memory does not function as a perfectly optimized summarization system. Important information can be lost, distorted, or made inaccessible, while irrelevant information can remain unusually persistent.

Sleep contributes to the consolidation and reorganization of memories, helping stabilize some information and alter how it is represented over time. Research suggests that sleep supports the selective strengthening of certain memories and the integration of new learning with existing knowledge. These processes are more complex than a simple deletion of irrelevant details, and their effects vary according to the kind of learning, the stage of sleep, and other conditions.

The broader principle remains important: intelligent behavior does not require the equal preservation of everything. It requires representations that retain the distinctions and relationships necessary for future use. A system that remembers every detail but cannot identify the important ones may be less effective than a system that preserves a carefully structured account of what matters.

Part 7: Sleep as a Condition for Cognitive Renewal

Sleep is often understood as a period of inactivity, a temporary suspension of the work performed by the conscious mind. From the perspective of cognition, however, sleep is not simply an absence of wakefulness. It is an active biological state associated with changes in neural activity, memory processing, metabolic regulation, and the functioning of multiple physiological systems. Its importance becomes especially apparent when sleep is insufficient and the quality of waking cognition begins to deteriorate.

Sleep deprivation affects sustained attention, vigilance, working memory, decision-making, and the regulation of behavior. Some effects are immediately noticeable, while others are more insidious. A tired person may recognize that they feel sleepy but fail to appreciate how much their performance has declined. Brief lapses in attention can occur without a deliberate decision to disengage, and the ability to detect and correct errors can become less reliable. The individual may continue working while the quality of that work deteriorates.

The relationship between sleep and memory is especially relevant to the comparison with language models. Learning is not complete when information has merely been encountered. Newly acquired material must be stabilized, integrated with existing knowledge, and made available for later use. Sleep contributes to these processes, although the precise mechanisms depend on the type of memory and the circumstances of learning. Inadequate sleep can therefore interfere with the transformation of experience into durable and usable knowledge.

This distinction separates the mere accumulation of experiences from the effective incorporation of those experiences into a cognitive system. A person may spend an entire day reading, studying, working, and making decisions, yet derive less lasting benefit from those activities if sleep is insufficient. The day has produced a large quantity of input, but the biological processes supporting learning, regulation, and recovery have not been adequately maintained.

The analogy with a long-context language model is tempting. One might imagine the waking day as an expanding context in which every experience, concern, and unfinished task competes for attention, while sleep acts as a period of consolidation and reorganization. In this metaphor, the mind does not simply need more capacity to hold its experiences. It also needs conditions under which those experiences can be integrated and made useful.

The metaphor must nevertheless be handled carefully. The brain does not literally accumulate a textual context window during the day, and sleep is not a general-purpose operation that clears cognitive clutter. Nor is there evidence that ordinary sleep deprivation impairs cognition because the brain has reached a maximum capacity for accumulated experience. Sleep loss produces a range of physiological and cognitive effects, many of which cannot be reduced to information overload.

The more defensible comparison concerns the conditions necessary for effective processing. A language model can have access to information that it fails to use appropriately because of limitations in its processing architecture or the organization of its context. A human being can possess the knowledge necessary for a task but fail to use it reliably because fatigue has impaired attention, working memory, or executive control. In neither case is the available information alone a sufficient explanation of performance.

Sleep therefore highlights a feature of intelligence that is easy to overlook: cognitive ability is not simply a fixed quantity that remains equally accessible under all conditions. Its expression depends on the state of the system, the demands of the task, and the processes that sustain attention and coordination. An intelligent system must not only acquire information and perform computations. In biological organisms, at least, it must also maintain the conditions that make effective cognition possible.

Part 8: The Difference Between Knowing and Understanding

The distinction between information and effective cognition leads to a deeper question: what does it mean for a system to know something? If a fact is present in a model's context but does not influence its answer when it should, in what practical sense has the model successfully used that fact? If a person has learned a principle but cannot apply it when circumstances demand, how much of that knowledge is available in the form that matters?

These questions do not require us to settle the philosophical problem of whether language models possess understanding in the same sense as human beings. They arise at the more modest level of functional competence. Information becomes useful when it can be brought to bear on the task for which it is relevant. Merely storing a statement is different from recognizing its implications, connecting it with other statements, and allowing it to constrain a conclusion.

Consider a simple example. A model may be told that a particular recommendation is appropriate only when three conditions are satisfied. Later in the conversation, the user supplies information that establishes two conditions but leaves the third unresolved. If the model recommends the action without acknowledging the missing condition, the relevant rule may have been present in its context, yet it has not been applied correctly. The failure lies not in the absence of information but in the failure to enforce the relationship between the rule and the available evidence.

A human expert can fail in an analogous way. A doctor may know that a treatment is contraindicated under a particular condition but overlook evidence that the condition is present. A lawyer may understand the general principle governing a case yet fail to notice that an exception applies. An engineer may remember a safety requirement but neglect to incorporate it into a calculation. In each instance, knowledge exists in some form, but its influence on the current decision is inadequate.

This suggests that practical understanding involves more than the possession of correct propositions. It also involves the ability to preserve dependencies, recognize when a principle applies, detect conflicts, and revise conclusions when the evidence changes. Understanding is expressed through the organization of knowledge into a structure that can guide inference and action.

The point becomes especially important when information is distributed across a long history. A conclusion may depend on an assumption introduced much earlier, a qualification stated in passing, or a distinction that was clarified only after several rounds of discussion. To reason correctly, the system must preserve not merely the individual statements but the logical relationships among them. A collection of facts is not yet a coherent model of the situation.

This is why a concise but well-organized representation can sometimes outperform a complete transcript. The transcript preserves the historical record, while the organized representation makes the current structure of the problem explicit. It indicates which facts remain relevant, which conclusions have been established, which assumptions require verification, and which questions are still open.

The distinction also explains why a system's ability to reproduce information should not be confused with its ability to reason reliably. Recall is important, but reasoning requires the appropriate use of recalled material. The decisive question is not only whether the system can retrieve a fact, but whether it can determine when the fact matters and what follows from it.

Part 9: Toward a Theory of Effective Intelligence

The preceding observations suggest a conceptual model in which effective intelligence depends on the relationship between informational resources and the demands of processing those resources. Such a model would not attempt to reduce intelligence to a single numerical quantity. Instead, it would distinguish the availability of information, the capacity to manipulate it, the ability to select relevant material, and the reliability with which the system checks and revises its conclusions.

For heuristic purposes, we can express this relationship as:

\[ I_{\mathrm{eff}} = f(K,P,A,C,R), \]

where \(I_{\mathrm{eff}}\) denotes effective intelligence in a particular task, \(K\) represents available knowledge, \(P\) represents processing capacity, \(A\) represents attentional selection, \(C\) represents coordination and control, and \(R\) represents the reliability of retrieval and evaluation. The function \(f\) is deliberately unspecified. The expression is a conceptual framework, not a validated psychometric formula or a claim that intelligence can be calculated directly from these variables.

The value of this framework lies in its refusal to treat knowledge as the sole determinant of performance. A system may have a high value of \(K\) but still perform poorly if relevant information is difficult to identify, if processing demands exceed its effective capacity, or if the mechanisms responsible for checking conclusions are unreliable. Conversely, improvements in organization, retrieval, and control may produce substantial gains without requiring a proportional increase in the total amount of stored information.

For language models, this perspective supports the development of architectures and workflows that do more than enlarge the context window. A system might retrieve only the portions of a long history relevant to the current question, maintain a concise record of active constraints, distinguish established facts from hypotheses, or periodically summarize decisions and unresolved issues. It might also verify that the final answer remains consistent with information that should constrain it. These methods address the organization of cognition rather than simply expanding its informational capacity.

Such approaches have limitations. Retrieval can omit the very detail that would have changed the answer. Summaries can erase uncertainty or introduce distortions. A compressed record may conceal the reasoning behind a conclusion, while an incorrectly maintained state can preserve an outdated assumption with unwarranted confidence. Better information management therefore requires not just compression but methods for preserving provenance, uncertainty, dependencies, and the ability to return to original evidence when necessary.

For human cognition, the same framework emphasizes the importance of managing demands on attention and working memory. Complex tasks can be broken into stages, intermediate results can be recorded externally, distractions can be reduced, and priorities can be made explicit. Such techniques do not increase every underlying cognitive capacity, but they reduce the amount of information that must be maintained internally at any given moment. They make it easier for the mind to devote its limited resources to the relationships that actually determine the outcome.

Sleep belongs to this broader picture because it supports the biological conditions required for reliable cognitive functioning. It cannot be replaced by better note-taking or more efficient task organization, just as a larger context window cannot eliminate every limitation of a model's architecture. Different failures require different remedies. The usefulness of the comparison lies in recognizing that the quality of reasoning depends on a system's internal organization and operating conditions as well as on the information it possesses.

One further implication concerns the evaluation of intelligence. If a system's performance varies substantially with context organization, fatigue, distraction, or task structure, then a single successful demonstration may provide an incomplete picture of its capabilities. A model that solves a problem when all relevant facts are presented together may fail when the same facts are dispersed across a long interaction. A human expert who performs brilliantly under rested conditions may make serious errors after prolonged wakefulness. Reliable intelligence must therefore be assessed not only by peak performance but also by robustness across the conditions in which reasoning is required.

This does not mean that all failures should be explained as limitations of context or control. Errors can arise from false beliefs, insufficient knowledge, flawed representations, poor algorithms, mistaken assumptions, or inadequate reasoning strategies. Effective intelligence is a multidimensional problem, and no single analogy can account for every failure. Nevertheless, the distinction between information and its effective use provides a valuable corrective to the assumption that adding more information must always improve performance.

Part 10: The Importance of Knowing What Matters

The pursuit of intelligence is often imagined as a pursuit of expansion. We seek larger memories, broader knowledge, faster computation, longer context windows, and the ability to consider more possibilities at once. These ambitions are understandable, since many genuine limitations arise from insufficient information or inadequate processing capacity. Yet expansion is not the whole story. As the quantity of available information increases, the ability to determine what matters becomes progressively more important.

An intelligent system must distinguish what is essential from what is incidental, what is current from what has been superseded, what is established from what remains uncertain, and what should influence a decision from what should be ignored. These distinctions are not decorative additions to reasoning. They are part of what makes reasoning effective. Without them, information can accumulate faster than it can be integrated into a useful representation of the world.

The same principle appears in the relationship between expertise and attention. An expert does not necessarily outperform a novice because the expert consciously considers more details. Often, expertise allows the individual to recognize the structure of a situation quickly, identify the variables that matter, and disregard information that is irrelevant to the immediate problem. Experience has been transformed into patterns and abstractions that reduce the burden of explicit computation.

This form of selectivity should not be confused with carelessness or premature simplification. Important details can be overlooked precisely because they initially appear insignificant. Good abstraction must preserve the features that determine the answer, and good attention must remain responsive to evidence that challenges an established interpretation. The objective is not to reduce complexity indiscriminately, but to represent complexity at the level necessary for sound judgment.

The comparison between long-context language models and sleep-deprived humans ultimately reveals a common question about the conditions under which intelligence becomes effective. The model may possess the information needed to answer correctly but fail to integrate it across an unwieldy context. The human being may possess the knowledge needed to solve a problem but fail to apply it reliably when attention and executive control have deteriorated. The underlying mechanisms are different, but both examples challenge the assumption that available information translates directly into successful reasoning.

Perhaps, then, the most useful way to think about intelligence is not as the ability to contain the largest possible amount of information, but as the ability to establish a productive relationship with information. This relationship requires memory, but it also requires selection. It requires computational or cognitive capacity, but it also requires organization. It requires the preservation of experience, but it also requires abstraction and the ability to disregard what is no longer relevant. In biological systems, it further requires the physiological conditions that sustain attention, learning, and self-regulation.

This perspective changes how we think about the future of artificial intelligence. Progress may depend not only on models that can accommodate increasingly long histories, but also on models that can organize those histories into reliable, revisable representations. A system that can distinguish the essential structure of a problem from the incidental details of its history may reason more effectively than one that merely retains a larger quantity of text. Greater capacity remains valuable, but its benefits depend on the mechanisms that transform capacity into usable intelligence.

It also changes how we understand our own minds. When we are tired, distracted, or overwhelmed, the answer is not necessarily to acquire more information or exert more effort. Sometimes we need to reduce competing demands, externalize intermediate steps, recover the capacity to concentrate, or allow time for learning and experience to be consolidated. These measures do not replace knowledge. They create better conditions for knowledge to guide thought and action.

There is, finally, a subtle reversal hidden in the original analogy. We tend to regard forgetting, interruption, and rest as obstacles to intelligence because they appear to reduce the amount of information immediately available to us. Yet selective attention, abstraction, and sleep are integral to the functioning of human cognition. Their value lies partly in how they shape the use of information, not merely in how much information they preserve. Similarly, the answer to the limitations of language models may not always be to give them more context, but to improve the ways in which they identify, organize, and use the context they already have.

The deepest lesson is not that too much information makes a mind unintelligent. It is that information and intelligence are different things. Information provides possibilities; intelligence establishes relationships among them, determines which are relevant, and draws conclusions that remain coherent under the constraints of the task. A system can know a great deal and still fail to think effectively if it cannot manage the demands of using what it knows.

The ultimate challenge, for artificial intelligence as well as for human cognition, is therefore not simply to remember more. It is to develop increasingly reliable ways of deciding what deserves attention, what must be preserved, what can be compressed, and what should guide the next thought or action. The measure of intelligence may lie less in the size of the library than in the quality of the understanding that can be drawn from it.

References

  1. Baddeley, A. (1992). "Working Memory." Science, 255(5044), 556–559. 🔗
  2. Lim, J., & Dinges, D. F. (2010). "A Meta-Analysis of the Impact of Short-Term Sleep Deprivation on Cognitive Variables." Psychological Bulletin, 136(3), 375–389. 🔗
  3. Rasch, B., & Born, J. (2013). "About Sleep's Role in Memory." Physiological Reviews, 93(2), 681–766. 🔗
  4. Walker, M. P., & Stickgold, R. (2006). "Sleep, Memory, and Plasticity." Annual Review of Psychology, 57, 139–166. 🔗
  5. Cowan, N. (2001). "The Magical Number 4 in Short-Term Memory: A Reconsideration of Mental Storage Capacity." Behavioral and Brain Sciences, 24(1), 87–114. 🔗
  6. Miller, G. A. (1956). "The Magical Number Seven, Plus or Minus Two: Some Limits on Our Capacity for Processing Information." Psychological Review, 63(2), 81–97. 🔗
  7. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). "Lost in the Middle: How Language Models Use Long Contexts." Transactions of the Association for Computational Linguistics, 12, 157–173. 🔗
  8. Vaswani, A., et al. (2017). "Attention Is All You Need." Advances in Neural Information Processing Systems, 30. 🔗
  9. Tononi, G., & Cirelli, C. (2006). "Sleep Function and Synaptic Homeostasis." Sleep Medicine Reviews, 10(1), 49–62. 🔗