The Organisational Forgetting Curve
Focusing on preventing forgetting can be as powerful as learning new things
Some learning is hard. Some learning is easy. Some learning is hard and then easy. What dynamics shape the way we learn personally? What about at an organisational level?
César Hidalgo, a physicist with an primary interest in the theory of networks, set out to answer these questions. But before we look at learning, let’s think about what is being learned.
In 2015, Hidalgo wrote the book Why Information Grows, where he explores why economies are structured in the way they are.
For many economists, capital is everything. For Hidalgo, information is everything. And information is instantiated primarily in the complexity of a product. He argues that you can determine long-run economic growth by looking at what a country exports. It is important to distinguish long-run growth from short-run growth, which is much more volatile.
He uses two measures of these exports - diversity and ubiquity. Most countries export pretty similar things to other countries, but the difference is in the number of things that they export. For instance, Hidalgo cites this example:
“Consider the exports of Argentina, Honduras, and the Netherlands. Of the 50 products that Honduras exported in 2008, Argentina exported 25 (50 percent) of them and the Netherlands 48 (96 percent). Of the 227 products that Argentina exported in 2008, the Netherlands exported 213 (94 percent). These overlaps tell us that the exports of Honduras are—statistically speaking—a subset of those of Argentina, and those of Argentina and Honduras are in turn a subset of the exports of the Netherlands.”
Why is this the case? Why don’t countries export massively different things? As Hidalgo asks:
Why is it easier to bring the lithium atoms that lie dormant in the Atacama Desert to Korea than to bring the knowledge of lithium batteries that resides in Korean scientists to the bodies of the miners who populate the Atacaman cities of Antofagasta and Calama?
Why can’t those miners learn to manufacture batteries themselves?
Knowledge and Information
Four concepts underlie Hidalgo’s work:
Knowledge.
Knowhow.
Entropy.
Information.
Knowledge involves relationships between entities. These relationships can be used to predict the outcomes of events without having to act them out.
Knowhow involves the capacity to perform actions. Knowhow is often tacit, in the sense that we know how to walk but we usually don’t know how we walk. It is tacit computational capacity that allows us to perform actions.
Entropy is often described as a measure of disorder. Specifically, entropy is a measure of the number of microstates that a particular system can occupy. Imagine two bags which can contain one marble each. There are four possible states; both bags empty, one bag with one marble, the other bag with one marble, and both bags with a marble in them.
If we increase the number of bags and the capacity of each bag marbles by two each, there are suddenly 256 states. This grows exponentially with more bags and greater bag capacity. Entropy is about the whole menu of states a system could be in, weighted by how likely they are.
Information is about the significance of learning which one state the system is actually in. In the words of Manfred Eigen, to be “completely informed” means:
having the ability to distinguish between different alternatives, or, in physical language, to discern between different microstates.
These concepts are almost infinitely extendable and mutable. For instance, information, or relative entropy, can also determine Bayesian surprise.
In Bayesian surprise, the importance of information is in how much it changes our posterior predictive distribution from our original set of priors, or in layman’s terms, how much a new piece of information matters for our world model. If you learn something new, does it change the way you see the world? And the natural consequence of this is that most things you learn will be relatively unsurprising.
One of Hidalgo’s central ideas is that while all of these forms of knowledge may be transmitted between people, in a process we call learning, this process is very lossy. Knowhow, composed primarily of tacit knowledge, is difficult to transmit, even if people are in the same room.1
One way to think about Hidalgo’s argument is that understanding must almost be embedded into someone’s nervous system by consistent practice. This means that it is very difficult to transfer these ideas quickly, because knowledge and knowhow transfer are slow. If practice isn’t consistent, forgetting also occurs, making long-term retention difficult.
Decay of Knowledge
In Hidalgo’s 2025 book, The Infinite Alphabet, he discusses the decay of knowledge. I’ve written about forgetting curves before, and it is no surprise to most people that knowledge decays rapidly. People are constantly frustrated that something learned one day is forgotten the next. Similar principles apply to manufacturing and to organisations.
During World War Two, the United States produced cheap mass-produced service ships known as Liberty ships. Shipyards could only develop experience producing liberty ships once they opened, and some shipyards were established before others.
In 1965, Leonard Rapping used the production figures of these shipyards to measure changes in output. He concluded that the increase in productivity observed in these shipyards was not coming from economies of scale or changes in technology. He instead concluded that ‘the secular improvement in output per man-hour during the war is attributable to learning’.
As you can see in the chart below, initially ship production was very slow, taking thousands of hours. As shipyards acquired more experience, the time taken to produce a ship decreased around eightfold, falling from ~2000 hours to ~250 hours.

But Rapping also found that this learning curve decayed quickly. If you stopped making ships, you forgot how to make ships (imagine creeping to the left on the chart above).
Hidalgo writes that:
In 1990, Linda Argote, Sara Beckman and Dennis Epple concluded that knowledge decayed quickly, at a rate of about 15% to 25% a month. These estimates were later improved by Peter Thompson, who adjusted them by controlling for other factors, and concluding that knowledge depreciated at a slower monthly rate of 3-6%. This implies that a shipyard that stops producing ships would lose half its knowledge in about a year.
I’ve taken the liberty of plotting this information in a ‘forgetting’ curve. Starting at 100 imaginary ‘knowledge units’, you can see that at higher rates of forgetting, an organisation runs out of ‘knowledge units’ completely after about 12 months, and no longer has the right to have an owl in their logo.
This has some implications. For instance, if you want to build a manufacturing sector in a country that has previously discarded its manufacturing sector, it will take time to slowly build it up. And if you have certain forms of infrastructure that need almost constant production in your country (roads, underground train networks, bridges), you should constantly have people producing them, or the knowledge of how to make this infrastructure will wither.
The rate of forgetting is not flat. It depends on where your knowledge and knowhow are embedded. For instance, if they are embedded in people, they decay faster than if they are embedded in technology. Surely, the answer is obvious then: just embed all of your practices into technology.
The catch is twofold. Firstly, people are required to embed knowledge and knowhow into technology, and secondly, technology is not flexible. It can’t respond on the fly to new problems or new strategies by competitors. You need people to do that.
Embedding Knowledge Into Things
Knowhow can be moved by people, but not by technology. Hidalgo gives the example of Samuel Slater, a spinning frame operator who went to America and reconstructed the water frames and carding machines that were being used by Richard Arkwright and others for textile manufacture back in the United Kingdom.
One New England spinning frame operator, Moses Brown, had acquired a machine built by two Scotsmen who had seen the Arkwright system, but didn’t know how to operate the machines. Merely having a machine doesn’t allow you to operate it. And besides, the machine had flaws that needed repairing. Brown had to wait for Slater to arrive and fix his machines for proper textile production to get underway.
And once it got there, it spread slowly, through New England, by people who could pass the knowhow onto others. Hidalgo’s point here is that it is hard to move knowhow around. Merely giving someone a machine means they have to spend a significant amount of time working it out. Teaching someone to use it is a faster process.
Even then, it is generally quite hard to help people acquire new skills even if you have people with those skills directly educating others. In the words of Dave Snowden, “We always know more than we can say, and we will always say more than we can write down.” Tacit knowledge is difficult to elucidate, let alone to then embed into technology.
This all means that is very difficult for organisations to learn. And knowledge transfer is a two-way process. In 1990, Wesley Cohen and Daniel Levinthal coined the term absorptive capacity, referring to the ability of a firm “to recognize the value of new, external information, assimilate it, and apply it to commercial ends”.
So when you lay people off, you often lose more expertise than you might suspect from a surface analysis of performance or metrics. This is the point Howard Yu made recently with regard to AI, where many people assume that AI can replace personnel in their business without paying attention to what those people actually offered.
Organisations also contain expertise in relationships between people, tools, and tasks, and removing nodes from such circuits dissolves the joint expertise buried in the links connecting that node to others. All of these linkages also decay generally, even without layoffs, if they are not refreshed.
Curves, Curves, Curves
There are three different types of curves that matter here. The first is a forgetting curve. The second is a learning curve. The third is an experience curve.
There are three stories that he tells. I discussed the forgetting curve above, when I talked about the Liberty ships and the decay of knowledge.
The second curve, the learning curve, refers to an individual’s gradual improvement at a specific skill. A classic example of Louis Thurstone’s typing class in the 1930s, where people gradually improved in the number of words that they could type per minute (WPM). In this taxonomy, learning curves are limited to individuals, and so have a hard cap. Even if you’re trying to win speed stenographer of the year, there is a limit to the amount of WPM you can manage.2

In Thurstone’s curve, improvements happen quickly initially and they then hit a plateau.3 And we can expand the Thurstone curve to the organisational level. For instance, in the 1930s, Theodore Wright’s work on the production of planes showed that the cost of producing a plane decreased as a power-law.
Wright used the fact that planes are produced in batches to study the cost of the last plane in a batch. The cost of the last plane in a ten-plane batch was about half that of a plane made in a batch of one. And like Thurstone, Wright also found improvements petered out, with most of the learning accumulating at the beginning of each batch. In a twenty-plane batch the cost was about 41 percent of the original cost, and in a thirty-plane batch, the cost dropped only a few additional percentage points, to about 38 percent.
Hidalgo argues that is a combination of these learning curves that produces runaway Moore’s law type effects.
This idea has important implications for starting a new business. If you want to enter a technological space as a less established producer, you need to make sure that space is at the start of the learning curve. You can’t compete with the end of the learning curve from the start of the learning curve. But you can compete if you’re on a different learning curve.
Let’s say you wanted to build a software company that might compete with Anthropic, Google and OpenAI. Building an large language model using the known transformer-attention architecture is going to be too slow. Experience is critical to success, and you’re not going to catch-up.
You might be better off trying to get in early on designs that solve fundamental problems that the current transformer paradigm has.4 Or trying to work out where demand might shift if AI companies were forced to offer their compute at a lower subsidy.
The idea that each doubling in experience leads to the same relative price decline is now known as Wright’s Law, and can be seen in the chart below of the cost of solar panels, which fall by around 20% in cost for every doubling of installed capacity.
What is the difference between using Wright’s Law and Moore’s Law to understand an industry? Wright’s law is a way of showing that costs fall exponentially as production expands, while Moore’s law says that costs fall exponentially with time. Both laws are similar, because production itself tends to grow roughly exponentially over time for many technologies.
If costs fall exponentially with time and cumulative production also rises exponentially with time, then you can transform one law into the other. So we could call it Mright’s Law. Or Woore’s law. One law is conditional on time and the other is conditional on production, so the relationship only breaks down if production changes.
The linear increase in production under Woore’s Law tends to materialise in industries like manufacturing where there are standardised, high-volume products that need to be produced. You can get quick feedback, and there are competitors pushing you to lower your costs. Transistors, solar panels and the explosion of compute with the recent development of AI are all obvious examples.
But different goods will have different learning rates, and we can infer from above that we have to be continuously installing such capacity, or the relationship will start to break down. Even in cases where we are installing such capacity, it doesn’t necessarily create a virtuous learning curve. Fossil-fuels haven’t experienced the same kind of experience-curve cost decline that renewables do, which is a main reason why renewables have become cheaper than fossil-fuels so fast.
Measuring Forgetting
Wout Broekema argues that when it comes to organisational learning, we should adopt a cyclical rather than a cumulative perspective. The cumulative perspective is often implicit, and believes that all the lessons an organisation learns are only ever accumulated.
The cyclical perspective assumes that organisations learn in cycles that include both learning and forgetting. Learning only seems as though it is cumulative because learning is primarily an active, visible process, whereas forgetting is much less visible.
Organisations often expand naïvely. Learn more, produce more, grow more. But organisations should also be trying to monitor where they forget. Where is expertise going to be lost over the next year? Where do we need to put practices into place to cement our knowledge and knowhow?
I’ve plotted the forgetting curve again here, with more intuitive numbers.
If you look again at the forgetting chart, you can see that small differences in forgetting rates create huge differences in the amount of ‘knowledge units’ an organisation retains. Obviously, there are tradeoffs. If we’re spending time reconsolidating old knowledge, we’re not learning new things. But it is important to be aware of tradeoffs, and it is also easier to reconsolidate information than it is to learn new information.
If an organisation can decrease the forgetting rate from 10% to 5%, it retains about three times as much capacity after a year. This is likely to be faster than many organisations can grow.
Preventing forgetting can be just as powerful as learning new things. It’s just less visible.
Thanks to Andrew, Sue and Izzy for notes on this essay.
For instance, the SECI model explores how information moves between tacit and explicit knowledge in an organisation.
The story isn’t always an asymptotic curve to a ceiling. People can see non-linear gains, and there are also significant offline gains in ability at a task.
Transformer architectures have representations of information that are entangled, hard to interpret, and that are not well aligned with known symmetries. There are other networks that solve these problems. For instance, compositional pattern-producing networks can encode regularities such as symmetry, but currently cannot match transformers as scalable general-purpose models.





Also, there is a shrine in Japan which has destroyed and rebuilt itself every 20ish years since, maybe, about 4BCE.
“The shrine buildings at Naikū and Gekū, as well as the Uji Bridge, are rebuilt every 20 years as a part of the Shinto belief in tokowaka (常若), which means renewing objects to maintain a strong sense of divine prestige in pursuit of eternity, and as a way of passing building techniques from one generation to the next.”
One can see that as a deliberate ritual of remembering, and keeping process learning alive.
https://en.wikipedia.org/wiki/Ise_Shrine