Catastrophic Forgetting — The Brain Learns by Forgetting; the Machine Stores Where It Learns, and Collapses
Take a neural network that, trained on 100 photos of dogs, picked out dogs flawlessly, and teach it cats. Within a few rounds of training it learns cats. But the dogs vanish. It collapses as if falling off a cliff — the classifier that was dead-on a moment ago can no longer recognize a dog.
Strange. People, too, shed a little of the old as they learn the new — but they don't lose it wholesale like this. Stranger still: the brain forgets far more aggressively than a neural network. Every time it sleeps it prunes the weak connections. And yet it remembers better. You'd think the more you hold, and the longer you hold it, the better — so why does the side that forgets well win?
Let us plant one frame of reference we will use throughout. Call a neural network's memory a slab of wet cement. While the slab is soft, it takes new writing sharply — but the writing hand, passing over, smears the old writing beside it. Once the slab hardens, what was carved is preserved, but no new writing gets in. The power to carve and the power to keep sit on one and the same surface. To get ahead of the argument: the brain's trick with this slab is to leave it neither soft all over nor hard all over. It keeps only a chosen part soft, and it erases the faint marks. So the real question is not "how do you remember everything" but "what do you forget, and when, and where."
Learn the New and the Old Collapses
The phenomenon has a name. Catastrophic forgetting is the learning failure in which a neural network, learning tasks in sequence, sees performance on an old task collapse off a cliff the moment it masters a new one.
In 1989 two researchers first gave it a name. What they called it was catastrophic interference — memories stored one on top of another interfering with each other. A later paper put "forgetting" in its title and that name stuck; it is the same phenomenon under a different emphasis.
First naming · McCloskey & Cohen (1989) first framed the problem as catastrophic interference; French (1999) put catastrophic forgetting in a paper title and that name stuck.
The first demonstration was simple. A small network learning single-digit addition was taught the 'add-one' problems first, then trained on the 'add-two' problems; a single training pass on 'add-two' was enough to wreck the already-learned 'add-one'. The weights here are the internal parameters a network adjusts through learning — the place where knowledge sets. Their conclusion: as long as new learning touches those weights, a minimum of interference is unavoidable.
In human terms it would go like this. A person who, on memorizing the room layout of a new house, has the structure of the old house wiped out entirely. The whole memory holding an old phone number collapses the instant a new one is learned. People call this "AI forgetting the way humans do." A convenient analogy, and that is all it is. Humans do not forget like this.
The received view actually at stake lies elsewhere. First, people assume machines do not forget — that what you put in stays where you put it. Second, they assume that even if such a thing once happened, it is now old news: today's models absorb new information on the fly inside the context window, the block of input a model holds in view at once, so surely the problem is dated. What we mean to show is that both assumptions invert. Machines lose far more brutally than people do. But what summons that collapse is not the absence of forgetting; it is a design in which the place that learns and the place that stores are the same place. Having no apparatus to choose is a second deficit layered on top of that one — unable to choose what to protect, it loses more; unable to choose what to let go, the learner hardens. The context-window assumption we will settle separately, later.
Take the Engine Apart
Why does it collapse off a cliff rather than gently? Take the engine apart and out comes a single chain of interlocking parts.
First, a neural network holds knowledge as a distributed representation. A representation is the form in which the network keeps what it has learned; distributed means it is stored not in one neuron or one weight but scattered and overlapping across many weights. Back to the slab: a piece of knowledge is not stamped at one point on it but carved broadly, sharing ground with the marks other knowledge has left.
This overlap is not a defect. Because it stores things overlapping, a neural network generalizes even to inputs it has never seen. As was already pointed out in 1999, the very "single set of shared weights" that granted generalization is itself the source of interference.
Shared weights · French (1999) — a review pinning the main cause of catastrophic forgetting on the overlap of distributed representations in the hidden layer.
Second, a neural network learns by gradient descent — a learning rule that looks only at the error on the task it is solving right now and pushes every weight in the direction that reduces that error. The direction is found by sending the error back through the network, and that computation is backpropagation.
The crux is that this rule never looks at what past tasks encoded. It is a piano tuner who turns the pegs by ear to the piece being played right now, and never hears that yesterday's tuning hangs on those same pegs.
Now the two interlock. Old knowledge lies overlapping across many weights, and the instant gradient descent pushes those overlapping weights without looking at the past, the learning of the new task overwrites the old representation indiscriminately. Because one task's knowledge is scattered across a great many weights, when those weights are freshly updated the old task crumbles not bit by bit but all at once. Without overlap there would be nothing to overwrite; had it looked at the past it would have steered clear — but neither holds. So the collapse comes off a cliff. This is what catastrophic means.
Follow the chain to its end and one vital point remains. The place that learns and the place that stores share the same weights. The power to carve the new and the power to keep the old sit physically on the same parameters, so raising one lowers the other. This is the slab of wet cement.
Sharing alone is not enough. Most accounts slip right here. Even the same shared-weight network does not collapse if you interleave the old task and the new and feed them together in random order. Set the condition precisely and it is the conjunction of shared weights and sequential learning. The cliff appears only when tasks arrive one block at a time, in order — only when you finish learning the old and move on to the new.
If the cliff needs both conditions, then switching off either one stops the collapse. One is to sever the sharing. Give each new task its own site and leave the old sites untouched, and there is nothing to overwrite. Progressive Networks take this road. The other is to break the order. Mix old data back into the new learning and the sequentiality disappears.
These two conditions are often confused with a storage-capacity problem. A camera, however many new photos you take, never erases the old ones — because the place that stores and the place that processes are separate. A neural network collapses not because its storage is small. A larger model collapses just the same.
What happens when you scale up is not yet settled. There is an observation that forgetting worsens as scale grows, and a contrary observation that a larger pre-training makes a model forget less, so the direction splits by setting. Either way the conclusion is the same. The problem is not the amount of capacity but that the place that learns and the place that stores share the same weights while tasks arrive in order.
What confirms this diagnosis most strongly is an objection. "Large models in practice train on data shuffled together wholesale. Isn't catastrophic forgetting a lab phenomenon that shows up only in the artificial order of task A then task B?" True. Shuffling really does eliminate catastrophic forgetting. That is exactly the interleaving named above. And that is precisely what reveals the vital point to be 'shared weights and sequentiality' — remove the sequentiality and the collapse goes away. Why shuffling is a privilege, though, and what the objection leaves standing, we will settle later.
How the Brain Sidesteps the Vital Point
The brain sidesteps this vital point head-on. That is how Complementary Learning Systems reads the brain — on the brain's side there is no measured value for the size of this collapse. The brain does not receive its experience shuffled either. Yesterday and today arrive in order, and past events cannot be lived through again. The conditions are the network's; the outcome is not. The method is the two switches above, unchanged. Sever the sharing, break the order. Except the brain got there first.
The brain splits the place that learns and the place that stores into two outright. Complementary Learning Systems theory formalized this. The hippocampus stores individual experiences fast and without overlap. The neocortex stores slowly, overlapping, generalizing across many experiences. It is a division into a notepad scrawled hastily by day (the hippocampus) and an archive transcribed in a fair hand by night (the neocortex).
The theory itself set out from the motive of explaining the neural network's catastrophic interference. It flipped the network's 'fundamental flaw' into a lever for reading the organizing principles of brain tissue.
Complementary Learning Systems (CLS) · The theory formalized by McClelland, McNaughton & O'Reilly (1995), out of the motive of explaining the neural network's catastrophic interference.
The neocortex is slow because each experience is only one sample of the world. The statistics settle only by gathering many samples in small steps. Interleave new information with old and integrate it a little at a time, and the existing representations are disturbed less. Correct everything at once and they are badly damaged — as a network's are in sequential learning. The neocortex's slow, interleaved learning is the biological version of breaking the order.
When does that interleaving happen? During sleep. In 1994 an experiment observed that hippocampal cells that had fired together while a rat was awake and active fired together again in later deep sleep. The day's experience is replayed at night. That replay is the process that slowly moves the hippocampus's recent memories into the neocortex for long-term storage — memory consolidation. The name of an ML technique we will meet later, EWC, comes from synaptic consolidation, this consolidation's cell-level counterpart.
Hippocampal replay · First observed by Wilson & McNaughton (1994) in ensembles of place cells in the rat. This re-firing is what is called replay, and the ML replay that follows took the name from it.
Sleep does not only transcribe. The synaptic homeostasis hypothesis holds that synapses strengthen across the board while an animal is awake and that sleep downscales them again, pruning the weak ones. The picture is still contested, but there is a measurement behind it. Measure mouse motor and sensory cortex under an electron microscope and the synaptic contact interface after sleep is about 18% smaller than in the waking state — with the large synapses preserved and only the weak ones selectively shrinking. It is like pruning the weak branches overnight to make room for the next growth.
This forgetting is no passive decay. As shown most clearly in the fruit fly, the brain has a dedicated dopamine circuit that actively erodes memory traces. Researchers in this line argue that this intrinsic forgetting may be the brain's default state.
Pruning · During development, microglia — the immune cells that handle cleanup and maintenance inside the brain — literally engulf weak synapses to refine the circuitry. This pruning is a slower, separate process from the active forgetting above.
The point is not to extol the brain. By contrast, the brain shows why ML's vital point is vital. A system that pulls fast learning out of the shared site — the brain, its sites split between hippocampus and neocortex — does not collapse. A system that does not keep the order — the brain, replaying old and new mixed together with every sleep — does not collapse either. Two devices break the two conditions of collapse, one each: structural separation, and sleep replay with interleaving.
Pruning is not on that list. Erasing selected weak traces neither divides the sites nor breaks the order. What that forgetting does lies elsewhere. It does not avoid interference; it keeps the slab from hardening — and the second illness, in the next section, is where it belongs. The brain learns by forgetting — not forgetting in order to lose less, but forgetting in order to go on learning.
None of which makes the brain's synapses and a network's weights the same thing. What resembles is the problem faced and the structure of the solution, not the material.
The Same Dilemma, a Different Illness
So the answer looks simple. Freeze the old weights and keep them — why not? And right here the second illness comes out.
That the power to learn the new (plasticity) and the power to keep the old (stability) sit on the same parameters and form a trade-off — this is the stability–plasticity dilemma. Deep learning did not coin the term. It is a concept defined in neural modeling in 1980, decades ahead of deep learning. Brain or machine, any system that means to learn anew without indiscriminately erasing the old meets this dilemma.
Back to the wet cement. The two illnesses are two states of the slab. While the slab is soft, carving smears what lies beside it — catastrophic forgetting. Once it hardens, nothing gets smeared but nothing gets carved either — this is the second illness. Except this hardening does not come from anyone deliberately freezing weights. Loss of plasticity, reported in 2024, is a separate failure: with no special freezing at all, merely by learning on and on for long enough, a neural network gradually loses the very capacity to learn anything new. Units die, the rank of the representation collapses, and the learner hardens of its own accord.
Loss of plasticity · The original report is Dohare, Sutton et al. (2024), Nature 632. Units dying means that some of the individual neurons making up the network stop responding to any input at all; the rank of the representation collapsing means that the outputs of the surviving neurons overlap one another until a handful of neurons reproduce the whole layer, and the number of directions the network can hold apart shrinks.
The two must not be lumped together as "AI forgets." What is lost differs. Catastrophic forgetting is the loss of old knowledge; loss of plasticity is the loss of the capacity to acquire new knowledge. The split is not clean between old and new, though. In the original report the hardening network, its units dying and its representational rank collapsing, was closer to a degeneration that lost old and new alike. Researchers still drew an explicit line between the two, and for a reason. What is more fundamental to continual learning — learning in sequence while trying to keep the old — is the capacity to keep learning the new, and loss of plasticity is that capacity hardening and dying.
On one continual-learning benchmark the network's accuracy falls from about 89% initially to about 77% by around the 2000th task. That is the level of a shallow model with no stacked layers. Far from growing smarter the longer it trains, the learner hardens.
Benchmark · The 89→77% figures above were measured on Continual ImageNet — a problem the authors built by reworking ImageNet into a continual-learning setting, one that hands the network an endless run of binary-classification tasks.
This hardening can be worked on too. The same study kept plasticity alive by adding to standard backpropagation a device that, at every step, revives a fraction of the useless units with fresh random numbers — Continual Backprop. What the device did was erase what had gone useless and clear the ground. A forgetting device bolted on from outside held off the hardening. It did not do away with the dilemma itself.
Here is where forgetting finds its place. When we say the brain learns by forgetting, that forgetting is not a device for keeping the old. It is a device against hardening. What sleep gains by downscaling synapses is not the avoidance of interference either; it is removing the weak ones to improve the ratio of signal to noise and free up headroom for the next round of learning. Forgetting is not a device that prevents forgetting; it is a device that prevents hardening. This is exactly what standard backpropagation lacks.
This is the place to split the terms. Choosing what to protect and choosing what to let go are two different jobs. The first reduces how much is lost; the second keeps the slab from hardening. When this piece says the brain learns by forgetting, the forgetting is the second.
So the dilemma does not yield to a single knob. Stability and plasticity do push each other away: make it forget less and it cannot learn; make it learn well and it forgets. But the two failures do not sit at the two ends of that knob. Loss of plasticity does not come from turning the knob toward stability; it comes with the knob turned nowhere at all, from long training alone.
The root is still one. Go back to the wet cement and you see it. Carving and keeping sit on one and the same surface — both illnesses come from there. When the carving hand smears the old marks beside it, that is catastrophic forgetting; when the surface, carved over and over, hardens and takes no more, that is loss of plasticity. One comes from sharing the surface, the other from the surface being used without pause. The same condition — there is only one slab — forks in two. That is what it means to call them different illnesses born of the same dilemma. So no adjustment of a single 'how much to forget' fixes either one. This is why the brain's solution was not to tune a knob but to split the structure in two and, separately, to choose what to let go.
Engineering Rebuilds the Brain
Lay out the continual-learning techniques and one picture emerges. A good many of them are engineering reconstructions of the brain's devices from the last section.
Families that do not map onto the brain · Not all of them do. Some techniques map poorly onto the brain — those that avoid overwriting by pushing the new task's updates only in directions orthogonal to the ones the old task used, or those that tie the model to imitating the old model's answers while it learns the new task. The first family, though, partitions the subspace of update directions rather than the weights. It is closer to a continuous version of severing the sharing.
First, replay. Mix old data back into the new learning and break the sequentiality. The root of the idea is the brain. The brain came first, and ML followed. The 1995 proposal that brought rehearsal into neural networks, and Complementary Learning Systems theory with it, explicitly likened a network's rehearsal to hippocampal replay and sleep consolidation. The experience replay used by DQN — DeepMind's reinforcement-learning model that solved Atari games at human level — was itself drawn from hippocampal replay, and turning that replay off makes performance plummet.
The lineage of rehearsal · The source for network rehearsal is Robins (1995) on pseudorehearsal; its counterpart on the brain side is the Complementary Learning Systems of McClelland, McNaughton & O'Reilly (1995).
Second, EWC — Elastic Weight Consolidation. It picks out the weights that matter to the old task and holds them firmly in place. Importance is measured once, at the moment the old task has just finished: how much the old task's answers wobble when a given weight is nudged is computed right then. The holding force grows the further a weight strays from its old value, and the penalty rises with the square of that distance — a quadratic penalty. From then on it neither stores nor relearns old data. Unlike replay, which keeps the data, it is a regularization that hangs a spring on the important weights so they move less.
Why is 'consolidation' in the name? As the authors say in the paper, they took the neuroscience concept of memory consolidation — name, idea, and all.
EWC naming · The original paper's wording: "we develop an algorithm analogous to synaptic consolidation which we call elastic weight consolidation."
Third, Progressive Networks. For each new task it appends a new column and freezes the weights of the old columns — they are never updated again. Because the old sites are never touched, forgetting is zero by design. It ports to a machine the brain's structural separation of hippocampus and neocortex. In exchange, parameters grow as tasks accumulate. That is the cost the authors spelled out: forgetting was not removed for free but traded against capacity.
Two of the three stand on the same ground. Progressive severs the sharing; replay breaks the order. Each switches off one of the two conditions of collapse — the same ground the brain's two devices stand on.
축 위 자리는 순서만 나타낸다(거리·정도 아님). EWC는 이 축에 없다 — 두 조건 어느 것도 끄지 않고, 무엇을 지킬지 고르는 별개의 완화 층에 선다. 뇌 항목은 계산 수준의 대응이지 기질이 같다는 뜻이 아니다.
EWC is not on that axis. It neither severs the sharing nor breaks the order. It uses no old data, so the sequentiality stands; the weights stay shared and keep being updated. A spring merely holds them back. So EWC cannot remove the conditions of collapse. Nor is it wasted work. To the degree it holds the important weights in place, less of the old task is lost. It does not switch a condition off, it reduces how much is lost — which is to say it works on a different layer.
That second layer is the mitigation axis: choosing what to protect, and how much overlap there is. If the first layer is on and off, this one is a matter of degree. The better the choosing and the less the overlap, the less is lost. It is also where the observation mentioned earlier — that a larger pre-training forgets less — finds its explanation. Which settings produce that direction is still disputed, but when it appears, this is the layer doing the work. Less overlap in the representations means less gets buried. Overlap is not present-or-absent but a degree.
The two layers do not quite meet at the seam. If overlap is a degree, then the first layer's claim — sever the sharing and it stops — is a statement about the extreme value of that degree pushed all the way, and at values in between the collapse does not vanish but shrinks. The same holds for order: how much old data you mix back in is not a switch but a degree. And the erasure that prunes the weak is not this layer at all. It is clearing, not keeping; it works on hardening, not on how much is lost. The brain uses all three. It divides the sites and breaks the order to avoid the collapse, manages the overlap to lose less, and prunes weak traces to keep from hardening.
Overstatement is forbidden here. This correspondence does not mean 'the brain is a neural network.' Synapses and weights, dopamine and the replay buffer, are different materials. The materials differ, but the computational problem faced is the same. The two systems, facing that one problem — sequential learning over shared storage — arrived at structurally similar solutions (separation, replay, protection), and experts on both sides borrowed each other's formulations productively. The cross-referencing in which EWC cites consolidation and the latest edition of CLS cites deep networks is the trace of it.
Limits of the analogy · Huszár (2018) noted that EWC's original formulation, which stacks a separate penalty for each task, may be inaccurate as an approximation. EWC is an analogy to synaptic consolidation and an approximation of it, not an exact reproduction.
A good many of the studies that built this bridge came out of a single lab: DeepMind. The side that wrote the narrative linking brain and machine is, in that sense, also the author of that narrative. Not something to read as a conspiracy — the mutual citations are public fact — but not something to over-trust either.
Then is catastrophic forgetting an old problem already solved? There is replay, there are giant models, there is RAG. The mitigation techniques that have been adopted really do work. That much is granted.
And yet the fundamental dilemma is not solved. Replay that reuses the raw data has to keep storing and replaying old data, so it bears storage and privacy costs. A variant mimics old data with a generative model to avoid the storage, but then it bears the cost of maintaining that generator. Progressive grows parameters without bound. RAG — retrieval-augmented generation, the approach of finding and attaching external documents on the fly — is a detour that leaves the weights untouched and looks things up in an external memory. The mitigation is real, but it is a transfer of the cost rather than a dissolution of the dilemma. Unbounded sequential learning that keeps all of the old while learning all of the new is unsolved even now.
Not Being Able to Forget Is an Illness Too
This time the objection comes from the other side. If forgetting is so good, shouldn't someone who remembers everything be superhuman? The brain has both an illness of not being able to forget and an ability to remember everything — so isn't 'forgetting is good' a romanticization?
Half of the objection survives. That much is granted — there is no evidence that people who remember everything are unhappy. But there is no evidence that they learn better, either. What has actually been measured points closer to the opposite.
People with total autobiographical memory — hyperthymesia — can retrace the ordinary days of their lives date by date. Yet under controlled testing their standard memory was not significantly better than a control group's. Only the autobiographical memory is extraordinary; the ability is domain-specific, not general. A near-perfect memory is not an all-purpose memory. Which is not to say these people fail elsewhere.
But being truly unable to forget anything comes at a cost. The mnemonist S., observed for about 30 years by one neuropsychologist, forgot almost nothing — and it was that same indelible memory that got in the way of abstraction and of grasping the gist. He had trouble understanding metaphor and poetry. This is a single case, though, and it is entangled with his synesthesia — letters and sounds arrived with colors and tastes attached — so it is an illustration rather than evidence.
The two ledgers · The standard memory-test comparison is LePort et al. (2012), N=11. The record of some 30 years observing the mnemonist S. is Luria (1968).
One null result and one illustration. That is the whole of the ammunition, and the two point the same way: the problem is not the size of the storage but the ability to choose what to release.
Not being able to forget can turn straight into illness. PTSD is explained by the hypothesis that a traumatic memory, under the influence of stress hormones, sets too hard — overconsolidated — and remains in a state that is excessively strong and cannot be erased. Not being able to erase it is the core of the suffering.
So the thesis is not that "forgetting is unconditionally good." Forgetting indiscriminately (catastrophic forgetting) and being unable to forget anything (total hypermnesia) are both failures. The ammunition differs on the two sides, though. The first is backed by benchmark measurement; the second amounts to a null result, an illustration, and a hypothesis pointing one way. The way they point is this. The condition for learning is a designed, selective forgetting — pruning the weak traces and consolidating the important ones. What a standard neural network lacks by default is these two choosings. Choosing nothing to protect, its new gradient descent buries without discrimination; choosing nothing to let go, the slab hardens as it stands. The techniques above bolt those choosings on from outside, one at a time.
The Boundary, and What to Forget
Today's transformers absorb new information on the fly within the context window, so isn't catastrophic forgetting an outdated problem? No. Context-window learning does not change a single character of the weights. It is a temporary activation inside the context — a working memory that vanishes when the session ends. In brain terms it is not even storage, but a note held briefly in hand. Catastrophic forgetting still occurs in the real continual learning that actually updates the weights. Large language models collapse the same way when fine-tuned. The context window is finite, so it cannot be a substitute for unbounded continual learning.
That today's language models lack continual learning — the ability to keep learning after deployment and accrue it into the weights — was taken up in our piece on LLMs. Catastrophic forgetting is another face of why that absence is fundamental. Because old knowledge collapses the moment you update the weights in sequence, learning-in-deployment is not merely a 'feature not yet switched on.'
In environments where the data can be shuffled together wholesale, catastrophic forgetting is largely sidestepped. Many of today's flagship models train that way. That much is granted. In that setting catastrophic forgetting is not a problem of practice, and nothing in the diagnosis so far changes anything there. As long as shuffling works, the illness does not exist.
But shuffling is not an escape hatch; it is a privilege. Streams flowing in online, learning that runs on the device, privacy settings that require discarding old data — none of these can avoid sequentiality. There, catastrophic forgetting is not a lab artifact but a wall that really stands. And even an environment that has removed the order by shuffling is not safe. Train long enough and this time a different illness — loss of plasticity — is waiting. Dodge the first wall and another appears. The system that lives in a world where experience is never shuffled, the brain itself, carries that sequentiality by dividing its sites and remixing them every night — and forgets on top of that, so as not to harden.
Back to the wet cement. The working principle of catastrophic forgetting is one sentence. In a network where the place that learns and the place that stores are one and the same, update the weights with tasks that arrive in order, and the new gradient descent buries the old representation. One slab, and a hand passing over it again and again — a hand that does not look at what is carved beside it. Three things follow.
One. Capacity alone cannot fix it. Neither a larger model nor more memory touches the two conditions. There are at least three places to work: the structure that stops the collapse by cutting the sharing or the order; the selectivity that reduces how much is lost by choosing what to protect; and the erasure that keeps the slab from hardening by choosing what to let go.
Two. One knob cannot solve it. Forgetting and loss of plasticity are not the two ends of one axis but different illnesses born of the same dilemma, so 'how much to forget' alone fixes neither. The answer is a system that divides its sites, as the brain does, and then chooses separately what to protect and what to let go.
Three. So forgetting is not a bug to eliminate but a feature to design. On the machine side, measurement backs it: the learner with a forgetting device bolted on kept learning, and the learner without one hardened. The ammunition on the human side is thinner, but it points the same way. The condition for intelligence is not perfect preservation but selective erasure. The real question was, from the start, not "how do you remember everything" but "what do you forget, and when, and where."
We began from the fact that we expect a perfect memory from machines. But if forgetting well is the very condition of learning, then that is also where human cognition parts from the machine. Not knowing what to hold, but knowing what to let go.
Sources
- > The primary and secondary literature this piece rests on. All links accessed: 2026-07-12.
- McCloskey & Cohen (1989), "Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem" — Psychology of Learning and Motivation 24. experts.illinois.edu
- Ratcliff (1990), rapid forgetting and degraded discriminability in sequential learning — Psychological Review 97. PubMed
- French (1999), "Catastrophic forgetting in connectionist networks" — Trends in Cognitive Sciences 3(4). Cell01294-2)
- Grossberg (1980), stability-plasticity dilemma / Adaptive Resonance Theory — Psychological Review 87. Scholarpedia
- Kirkpatrick et al. (2017), "Overcoming catastrophic forgetting in neural networks" (EWC) — PNAS. ar5iv
- Robins (1995), pseudorehearsal (rehearsal) — Connection Science 7(2). IngentaConnect
- Rusu et al. (2016), "Progressive Neural Networks" — ar5iv
- Dohare, Sutton et al. (2024), "Loss of plasticity in deep continual learning" — Nature 632. ar5iv
- Huszár (2018), a note on the EWC approximation — PNAS. PNAS
- Ramasesh, Lewkowycz & Dyer (2022), "Effect of Scale on Catastrophic Forgetting in Neural Networks" — ICLR. OpenReview
- Luo et al. (2023), catastrophic forgetting in LLM fine-tuning — arXiv:2308.08747
- Mnih et al. (2015), DQN · experience replay — Nature 518. Nature
- Brown et al. (2020), GPT-3 few-shot (in-context learning) — arXiv:2005.14165
- McClelland, McNaughton & O'Reilly (1995), Complementary Learning Systems (CLS) — Psychological Review 102. Semantic Scholar
- Kumaran, Hassabis & McClelland (2016), CLS update — Trends in Cognitive Sciences 20. PDF
- Wilson & McNaughton (1994), hippocampal replay during sleep — Science 265. PubMed
- Tononi & Cirelli, synaptic homeostasis hypothesis (SHY) — PubMed
- de Vivo et al. (2017), synaptic contact interface shrinks by about 18% during sleep — Science 355. PMC
- Davis & Zhong (2017), "The biology of forgetting" (active forgetting) — PMC
- Paolicelli et al. (2011), microglial synaptic pruning — Science 333. PubMed