The Recorded Intelligence Manifesto

There's structure in the data, and it's teaching the machine

A generated technical texture used as a placeholder hero.

Recorded Intelligence: The Other Half of AI

Opening: Debunking The Founding Myth Of Modern Artificial Intelligence

The past decade of progress in modeling language has set off an avalanche of interest in new applications, algorithms, and hardware for running the newly minted AI technology of LLMs.

In addition to using and extending this powerful technology, it has also become important to understand the causes of this sudden technological advancement. This understanding is not only of intellectual curiosity, but also of vital importance to those managing policy or allocating capital. [needs stylistic revision]

The common narrative explaining this recent phenomenon is “backpropagation and Moore’s law.” In other words, simple but general computer algorithms for learning from data were coupled to an exponential increase in computational horsepower, and suddenly the AI started getting usefully smart.

Admittedly, this narrative is quite compelling in explaining the rise of LLMs circa 2020 to 2023 (in particular, the GPT-3 to GPT-4 arc). However, by 2024, the LLM paradigm already revealed the cracks forming in this explanation.

One of the most compelling chinks in the “backprop and Moore’s law” narrative was the trend of smaller models that leveraged differently shaped “thinking” sequences pushing far ahead of their bigger but “non-thinking” counterparts (e.g. OpenAI o1 being stronger than GPT 4.5).

We also find the backpropagation and Moore’s law narrative lacking when it comes to explaining the substantially different trajectories taken by non-language learning in the past decade. If the algorithms and GPUs are so good, then where’s my self-driving car running on deep learning?

The comparatively different trajectory of advancement in other popular applications like visual models and time series forecasting suggests that the common narrative is missing something big.

If it’s not simple learning plus compute, though, then what is responsible for all of this progress? Though no pithy explanation can fully pin down the messiness of reality, the addition just one more dimension affords us a much more fundamental view of what’s happening. This extra factor is the presence of intelligence in certain types of data, a phenomenon that we shall call Recorded Intelligence from here onward.

Recorded Intelligence represents a reframing that shifts our perspective from machine learning to machine teaching. We go from thinking that the intelligence is in the algorithms to acknowledging that the intelligence is actually out there in the world. When we remember that intelligence is being pushed into the algorithms through the data we feed in, rather than being an emergent phenomenon of the algorithms themselves, we suddenly see a much fuller and less brittle picture of how the cutting edge of AI is developing.

Once you see it, it’s hard to unsee. Just like after seeing David Graeber point out in Debt: The First 5000 Years that there is anthropological evidence of debt predating money, the convenient origin story of modern financialization never feels quite as solid. AI isn’t advancing just from throwing more GPUs at some abstract pile of data. It’s advancing from feeding more and more intelligence into the models we’re training on these GPU superclusters.

The Recorded Intelligence Thesis

Undoubtedly, machine learning has unlocked a lot of progress in artificial intelligence, taking us well beyond expert systems and logic programs. However, progress in machine-learned intelligent programs has clearly arrived not just by advances in algorithms and compute, but also by advances in the data input to these algorithms. In some ways this influence, though not as explicitly tracked in citation history or publicly disclosed through model releases, has been the more important driving force than algorithmic advances.

Here’s a simple equation: Intelligence = Learning(Examples). Learning is a function (the mathematical term for a behavior) that maps examples (something observed, recorded, of experienced) into intelligence. Intelligence in this equation itself is a function that can be applied to new inputs to produce new outputs. There is a more specialized case of that equation for computer-based intelligence: ArtificialIntelligence = MachineLearning(DigitalExamples). To run intelligence on computers, we also need to run learning algorithmically on computers, and we need to put teacher examples digitally into our computers digitally.

At the core, though, both equations show us that if the goal is intelligent behavior, both learning and examples are equally critical. Unfortunately, though, academically and professionally the data side of this equation has been underappreciated. It is not uncommon for the preparation of training data to be seein as scut work or tedium in academia and industry. Heck, it was explicitly written out from the entire theoretical framework in the case of Valiant’s influential founding structure for PAC learning back in 1984. In academia, the same data sets have been reused for years as the same raw fuel pouring into buried learning algorithms, and architecture/algorithm/objective function tweaks have a very straightforward map to publishing to be done all on pre-existing data and have been a surefire way to advance one’s academic career. You can also write proofs about algorithms in ways that are difficult about data.

Though an entire field of data science sprung up over the observation “garbage in, garbage out”, there is still a pervasive oversight when it comes to who gets the Nobel prize or billion-dollar funding round from venture capitalists. Historically we saw huge hype for AutoML, a field that was trying to automate away the thinking about the data part of the ML practice, only for the companies pioneering this technology to lose much of their shine and valuation as the promise never really panned out in a general sense.

In industry, maybe recorded intelligence is underloved because of academic hiring, but it’s also possible that it’s just not as sexy to realize that you’re sort of doing manual work, having humans create data, and kind of patching behavior piece by piece. It is pretty sexy to imagine computers teaching themselves and being intelligent and us humans just tweaking the clever equations and algorithms needed to make that happen.

Much ink has been spilled on the matter of Tycho Brahe being actually the OG behind the Copernican revolution, although he obtains much less fame than Kepler or Copernicus did when they wrote their laws. And all of science really advances with observations about the natural world prompting humans to think about the natural world more than they advanced from just ideas in a vacuum.

Like Brahe’s observations before Kepler’s laws, the bottleneck in AI has often been what we have managed to record from the world, not how clever our learning machinery is in the abstract. And this recording is not just a number of bits – it’s all about their content.

Defining recorded intelligence

I have spoken a lot about Recorded Intelligence without pinning down a real definition – let’s try to rectify that now. Recorded Intelligence can be loosely defined as knowledge or behavior that is exactly transmittable, either physically or digitally. This is a broad definition (e.g. LLMs themselves are artifacts recorded intelligence because they can be seen as compression of a lot of language or behavior data). Physical books also count because they record language losslessly, even though they’re not really directly compatible with ML programs. A distinction between digital and analogue Recorded Intelligence may be helpful but is not necessary to get to the heart of things.

Data Is Not Recorded Intelligence

I’m permissively including printed books and LLMs under the Recorded Intelligence umbrella, but critically Recorded Intelligence is not actually as broad as one might think. For one, Recorded intelligence is not just data – data is too broad of a concept. Take, for example, video data. You can shoot a video at raw 8K at 60 frames per second and get a movie that has way more bytes of data than shooting the same thing at 1080p at just 30 frames per second with MP4 compression. However, the actual difference in what you’ve captured is not nearly as massive as the difference in the bytes you’ve stored on disk. This example shows the difference between data and Recorded Intelligence. You can also collect data of random radioactive decay or data – radioactive decay data is of the natural world, but it doesn’t show how to intelligently navigate the natural world the way that, say, a video of how to open a door by turning the door handle embodies that intelligent behavior.

An analogy here is how geologists study the earth in a way far more precise than just looking at “dirt” as a homogenous and boring substance. Science can quantify and understand what is actually in the earth and the important bits of it, and a lot of it is just dirt and there isn’t really much of interest. There are different types of things in the earth and we dig them up and there are specific things that are incredibly valuable, like rare earth metals or precious metals or otherwise. Viewing recorded information through the lens of Recorded Intelligence is like viewing the earth beneath our feet though the lens of geology – we can articulate differences between certain types of substance, and these differences are so critically important that concepts like “earth” and “dirt” become more distracting than helpful. Completing the analogy, “data” is perhaps as useless a concept to Artificial Intelligence as “dirt” is to geology.

Information Is Not Recorded Intellgience

“Information” (in the mathematical sense of information theory), is also not Recorded Intelligence. Information theory actually describes how many bits of digital messages can be sent over a finite bandwidth, and this includes things like sending random data. In fact, random data holds the most information according to the theory, whereas information that has structure and actually means something to intelligent beings contains provably less information due to its semi-predictable structure.

The related field of algorithmic complexity gives us something a little bit more useful in the principle of minimum description length and the ideas of Kolmogorov complexity, but it still doesn’t really quite answer, like, what is the complexity of a textbook. These are immeasurable quantities, which don’t really capture the same notion as intelligence does.

[MISSING CONNECTION – need to stitch it together]

Historical continuity

Once we see AI through the lens of Recorded Intelligence, we can track the full arc of AI progress all the way from the 1956 Dartmouth summer conference to the present with clarity.

In that famous inaugural conference we saw the creation of the term artificial intelligence and a synthesis between fields like cybernetics, neuroscience, and formal logic. Rosenblatt’s Perceptron work in the early 1960s brought in 20x20 pixel grids of data representing silhouetted targets in aerial photos for military applications. Expert systems from the 1960s through 1980s brought to bear knowledge bases with many hundreds and thousands of logical-information-dense facts and rules, which were (post AI winter) themselves supplanted by even richer bodies of recorded intelligence like the hours of millisecond-accurate-annotated audio and phoneme data that was used to train the Hidden Markov Models that delivered the first working automatic speech recognition technology. What we see is that expert systems advanced because they were able to take in more and more knowledge from the real world through painstaking efforts to program them directly into programs and have knowledge bases and knowledge banks storing more and useful patterns extracted from real information about the world. We see machine learning really succeeded not just because it was a better algorithm, but because it was able to put a lot more intelligence to work. We were able to take many observations, often more than human programmers had patience to program by hand, and inject them into the output algorithms.

By looking at the full arc of AI history, we see that Recorded Intelligence explains far more than the three-year blip of LLM scaling laws. Those laws capture what happens when we scale compute and data without changing the kind of teaching signal, while the longer arc progress has been driven by new ways of recording intelligence for the machine.

Around somewhere in 2020 to 2023, LLMs were the preeminent driver of progress in artificial intelligence, and the zeitgeist in the era was that we were scaling towards a God model. We just wanted to feed it everything and scale the number of parameters up as high as possible, use a massive scale of compute, and achieve a predictable scaling towards God-like intelligence and create artificial superintelligence. However, there was an interesting subtext. Folks, maybe in the 2022 to 2024 overlapping window, noticed that retrieval-augmented generation could produce more useful artificial intelligence systems than pure scaling of models. They noticed that there was something special about extracting concrete known excerpts of data and augmenting models with them. And this gave way to what we have now in 2024 to present, which is the era of agents. Agents put the data into a different structure. They trained on “smart” data (also called “thinking” data), and they learned to produce it and train themselves on it through reinforcement learning.

“And here we find is that when you put data into the tools that these systems use and change the structure of the data so that it’s a richer, different sequence of tokens, you get much more useful behavior and the ability to solve problems that were unsolvable even with massive scale-ups of compute through the God model view with the same data distribution.”

[Key argument. Later shorten and clarify.]

“And so we see this data-dependent teaching matters more than learning view that kind of de-emphasizes the algorithm and emphasizes the intelligence itself, which is a potentially measurable quality of the data involved.”

[Very important. Later polish into a thesis sentence.]


Concrete examples: economic desperation around recorded intelligence

You can imagine that if algorithms were all that were needed for intelligence, we would just simply be running reinforcement learning from day one the same way that you train AlphaZero to play Go. However, in contrast to this we have numerous lawsuits against massive tech companies who have amassed incredibly valuable assets. For example, Meta has been accused of going to Z-Library. Others have been accused of going to Anna’s Archive to get tens of billions of e-books and torrenting as much as a petabyte of raw data. Google Books has a massive war chest of over 40 million scanned titles, and it is plausible that this data is somehow an input into the Gemini training pipeline.

The list goes on. Anthropic has had to admit their Project Panama exists through lawsuits when they digitally and destructively scanned millions of titles of books. We also see the gray zones, like, for example, Reddit was very open for a while, but then LLM training use of the Pushshift.io archives pushed Reddit to lock down their API access because of how valuable the data has been in teaching conversational agents.

There are also open beacons of recorded intelligence, which have been widely used in machine learning and are arguably responsible for enormous benefits to all of humanity in and beyond artificial intelligence. Things like Wikipedia contain a lot of the world’s knowledge and in refined forms that demonstrate intelligent reasoning from facts and organization. Common Crawl wins the award for scale, having amassed over 10 petabytes of raw HTTP traffic from crawling public web.

“Internet Archive has made available and free and computationally accessible massive troves of academic research.”

[Good example, though Internet Archive and academic research claims should be phrased carefully.]

“Stack Exchange has ensured that everything is openly licensed to Creative Commons licensing and it serves as a wonderful place to get data.”

[Good example.]

“I want to have the message here that the reader should understand there’s a level playing field aspect and that these open beacons are fighting back against big labs who are trying to close the door behind them to intelligence.”

[Strong policy/equity framing.]

“There are also some interesting lawsuits and exposé articles, for example, YouTube subtitles being in violation of terms of service scraped by Apple and others in addition to the Meta lawsuits and Project Panama lawsuits over torrenting and book scanning.”

[Useful list, but final version should pick a few and cite them.]

“And these are some of the interesting counterpoint to sort of academic culture that suggests algorithmic advances are responsible for everything.”

[Good return to core thesis.]

“Or Jensen Huang’s law, exceeding Moore’s law, are taking credit for advances in AI.”

[Good counterpoint to compute-centric narratives, but later make the phrasing cleaner.]

“Here what we see is that large companies are willing to take massive risks and spend lots of money settling or avoiding legal issues to acquire the same data they already had pirated.”

[Strong claim; needs careful legal wording in final.]

“And this just gives evidence to how valuable this is in practice from the people who are actually on the inside of pushing the frontier.”

[Good concluding sentence for this section.]

“Frontier Labs have spent billions on GPUs per year and less on large-scale recorded data, but they are not shying away from spending that kind of amount if need be.”

[Good numerical/economic point.]

“Anthropic is willing to spend $1.5 billion to settle a lawsuit, is willing to take the risk.”

[Strong concrete example; needs verification and citation.]

“The free web era masked the importance of data, but now we have scarcity with Reddit and other things and Cloudflare blocking bots, we have exposed the importance of the recorded intelligence from the web and other sources.”

[Excellent concluding observation for the economic/scarcity section.]


What this reframing changes

“What does this reframing change?”

[Good section title or rhetorical opener.]

“One thing is that large language models haven’t really invented a new form of intelligence.”

[Strong claim.]

“They’ve just given us a more efficient shovel or a shovel that we didn’t have before to dig up the intelligence that we’ve been recording for millennia.”

[Excellent metaphor. Keep.]

“As a corollary to this, LLMs are not human intelligence in the sense of the general intelligence of a single person experiencing the world.”

[Good distinction.]

“They’re actually crowd-level or civilization-level intelligence and are kind of a different form.”

[Important philosophical payoff.]

“LLMs are not individual minds that are kind of having general intelligence.”

[Good direct version.]

“They are crowd or civilization intelligence.”

[This should be a memorable line.]


Recorded intelligence beyond AI

“Just because recorded intelligence is useful beyond AI doesn’t mean it’s not also a critical driving force for AI.”

[Good transition to general-purpose-technology framing.]

“The analogy here is that we used to fund GPU development from gaming or crypto mining revenue, but then neural network demand later came in and outstripped it in funding.”

[Strong analogy.]

“And so a similar story could be playing out for recorded intelligence, where search engines or libraries gave way to LLMs, but the core value of the general-purpose technology of recorded intelligence has been there before even AI is the main consumer.”

[Excellent broader framing. Later clean up syntax.]


Future-looking note: compute and storage abundance

“The recorded intelligence Industrial Revolution being digital recording.”

[This is more of a phrase than a sentence, but it could become a subsection title.]

“Just like how we now have enough mechanically harnessed energy that it’s not a bottleneck for most production, and we have other issues stopping the progress of global industry, we now also have enough storage for a ton of recorded intelligence, and we may be coming up on having enough FLOPs for our algorithms to actually do a really good job modeling that intelligence.”

[Interesting future-looking argument. Later compress and place near conclusion.]

“Are we going to hit the top part of some sigmoid curve on storage?”

[Good provocative question.]

“It’s probably already happened.”

[Use carefully; may need evidence.]

“You can look at hard drives and GPUs may also happen.”

[Needs cleanup in final.]

“And that’s sort of an interesting contrast to this forever exponent idea that folks have around GPUs.”

[Good contrast to infinite-scaling narratives.]

“It’s not that we have the recorded intelligence abundance, it’s that we may actually be having an abundance in the other kind of critical ingredients for making recorded intelligence what it is today, and those other things are GPUs and storage.”

[This clarification is important.]

“We may actually hit the abundance there where those aren’t the limiting factor in the same way that fossil fuels are no longer the limiting factor for producing cars.”

[Good analogy. Later tighten and check if fossil fuels are the right comparison.]


Conclusion / call to action

“To close, we would probably do a bit of a call to action for a recap of how you should now be thinking about intelligence and learning and artificial intelligence and artificial machine learning.”

[Use as instruction, not final prose.]

“We’re trying to find a more fundamental view that’s less specific to a certain sub-era of progress.”

[Good near-conclusion sentence.]

“Instead of all this artificial intelligence and machine learning, you can just think about how real intelligence and learning can be kind of modeled.”

[Good broader frame.]

“If we had a scientific model of it that didn’t specifically sit in this kind of ML space, that would be a freeing way of actually seeing the core fundamentals.”

[Good closing aspiration.]

“Intelligence in the data explains so much more in our progress.”

[Strong final-line candidate, or near-final-line candidate.]