Call it entropy...no one knows what entropy really is.
A limit accepted, and the law of requisite variety
Richard S O'Rourke
This is the second post of three. The first was about the history of AI and Cybernetics and its interest in regulation. This one is about the limit on what regulation can do, and what accepting it makes possible. The third is a Cybernetic perspective on agent swarms.
"Just as the human muscles are absolutely limited by the law of conservation of energy, so is the brain limited by Shannon's theorems and the law of requisite variety" — Ashby, "Cybernetics Today and Its Future Contribution to the Engineering Sciences", 1961.
In December 2017, at a conference in Long Beach, Ali Rahimi collected a test-of-time award and used the platform to say that machine learning had become alchemy. He was careful about it.
"Alchemy's ok. Alchemy's not bad. There's a place for alchemy. Alchemy worked. Alchemists invented metallurgy, ways to make medication, dying techniques for textiles, and our modern glass-making processes."
The complaint was not that the results were false. It was that nobody could say why they worked, and that a technology in that condition should not be installed under everything.
Yann LeCun replied the next day. He called the analogy "insulting" and "wrong", and gave his counter-argument in one sentence:
"The engineering artifacts have almost always preceded the theoretical understanding."
Notice that the two men are not disagreeing about the history. Both accept that the machines came first. They are disagreeing about whether that is a problem, and neither has a way to settle it, because neither has said what the missing theory would be a theory of.
Joel Mokyr, the economic historian, splits useful knowledge into two kinds: propositional knowledge, which is knowing what is the case, and prescriptive knowledge, which is knowing how to do something. He shared the 2025 Nobel prize in economics for identifying the prerequisites for sustained growth through technological progress. The prerequisite his half of the citation names is the one this post is about: societies need to know not only that a technique works but why. Technique sits on top of an understanding of why it works, and he calls that understanding the technique's epistemic base. A technique cannot exist without a base, however narrow. What varies is how wide the base is, and the difference is not academic: with a wide base you can see where the next improvement is and why; with a narrow one you can only try things. Ashby had written that rule out as the rule of decision itself, in 1960:
"Use what you know to narrow the field as far as possible; after that, do as you please."
The base is what you know, and trying things is what is left when it runs out, with the one grace that a trial "may be a process that progressively wins more information." Mokyr's test is the sharp one. A narrow-base technology improves by trial and error, and cannot tell a dead end from a temporary obstacle.
By Mokyr's test, today's LLMs are a narrow-base technology, and the tell is the thing the field is proudest of. Scaling laws are how you decide what to build. You train a ladder of small models, plot how the error falls as you add data, parameters and computing time, and extrapolate the curve out to the size you can afford. It works. Chinchilla, the best-known case, used such a curve to show that everyone had the ratio of model size to training data wrong, and the prediction held when they retrained. But a fitted curve is not an account. It tells you what the error will be and not why, which is the epistemic status of Boyle's law before anyone knew what a gas was. The central planning instrument of the most consequential technology of the century is a regression.
So Mokyr gives you a question to ask, which is whether the base under a technique is widening, but no forecast. Narrow-base technologies have plateaued and they have also kept climbing, and nothing in his framework says which happens here. Read forward, he supplies a lens and not a prediction.
Except that there is one kind of forecast a wide base does supply, and it is not a trajectory. It is a bound. Carnot never predicted what any particular engine would do. He said what no engine could exceed, and that statement survived every change of design, fuel and scale that followed. A bound is the most durable thing a theory can produce, because it does not depend on the details of the thing it constrains.
The gusher¶
Ashby's "intelligence-amplifier" paper in Automata Studies, the 1956 volume Shannon and McCarthy edited, opens with the instinct and then a chart: "Our first instinctive action is to look for someone with corresponding intellectual powers: we think of a Napoleon or an Archimedes. But detailed study of the distribution of man's intelligence shows that this method can give little."

Ashby, "Design for an Intelligence-Amplifier" (1956), Figure 1: the adult IQ distribution, after Wechsler; with the author's red annotations.
The red annotations on the chart are mine.1
"What is important for us now is not the shape on the left but the absolute emptiness on the right. A variety of tests by other workers have always yielded about the same result: a scarcity of people with I.Q.s over 150, and a total absence of I.Q.s over 200. Let us admit frankly that man's intellectual powers are as bounded as are those of his muscles. What then are we to do?"
His answer, in McCarthy's own volume, was that "we must, somehow, construct amplifiers for intelligence — devices that, supplied with a little intelligence, will emit a lot," built by constructors who "are themselves quite averagely human."2 The same summer, at McCarthy's conference, the proposal on the table ran the other way. The study
"is to proceed on the basis of the conjecture that every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it."
To simulate is describe the feature and reproduce it. To amplify is to accept the bound and design within it. Both paradigms were in the room that summer in Dartmouth, and Solomonoff notes on Ashby's talk, 23 July: he
"is fascinated by the fact that a table of random numbers can give one all one wants — i.e. that selection is the only problem."
Then, on the two of them:
"McCarthy's view, that the only real problem is the search problem how to speed it up. Ashby feels the same but has fewer ideas. Ashby wants to start slowly, like organic evolution, and work up to more complex phenomena."3
The bound was then stated as a principle, in 1961, by the man who had spent thirty years looking for a way round it. Ashby grants LeCun's point before making his own. From that 1952 paper on building a chess machine:
"the measurement of a quantity can properly precede the understanding of what it is that is being measured—engineers were measuring the 'electric fluid' and were lighting towns with it several decades before its real nature was understood."
He is not asking anyone to wait for theory. He knows that is not how it goes.
What he argues instead is that what turned engine-builders into engineers was not a theory of heat. It was the acceptance of a limit. And he puts himself on the wrong side of it first. Speaking to instrumentation engineers in New York in 1961:
"When I started in 1928 I took for granted that the brain had a gimmick: find it, and we would have access to 'pure intelligence', that one could then tap in unlimited quantities, like a gusher. Today, we know better; but we know there is no gimmick, not because we have failed to find one but because we now know positively what is there."
The gusher is his perpetual motion machine. He had wanted one. Then the history he is reasoning from, in the same lecture:
"every imaginative engineer confidently expected that the machine with perpetual motion would be invented shortly. Then emerged an unpleasant surprise. Evidence began to grow that all machines were fundamentally limited, that energy was conserved, and that every erg of energy that came out had first to be got in. For a time many engineers felt profoundly disappointed. Nevertheless, the acceptance of this limitation was the first step to a new and altogether more realistic grasp of the principles of machinery. In the long run, the engineers who accepted the limitation outdistanced those who clung on to the hope of the perpetual motor."
The bound¶
The bound comes with a definition, in the same paper and the same breath. Intelligence, for Ashby, is not a property of a machine but a power shown in action: "The 'intelligent' processes par excellence are the goal-seeking — those that show high power of appropriate selection. Man and computer show their power alike, by appropriate selection." Selection toward a goal, out of the possibilities, and the more possibilities winnowed for the right one the more intelligence shown. He had said in 1956 what that leaves out: "Nothing is easier than the generation of new ideas: with some suitable interpretation, a kaleidoscope, the entrails of a sheep, or a noisy vacuum tube will generate them in profusion. What is remarkable in the genius is the discrimination with which the possibilities are winnowed." And the bound itself follows in the next sentence:
"If we accept the limitation — that appropriate selection can be achieved only to the degree that information is received and processed — and if we accept that this limitation holds absolutely over all brains, human and mechanical, our work, though less intoxicating, will in fact be more realistic. Those who build intelligent machines on this basis will outdistance those who want to build them on the old and superstitious basis that the human brain can do anything."
The obvious objection is that a limit is a counsel of despair, and Ashby answers it in 1956 with the best analogy in his writing. The engineers of the middle ages, he says, knowing the lever and the pulley, must often have concluded that since no machine worked by a man could yield more work than he put in, "therefore no machine could ever amplify a man's power." And yet "today we see one man keeping all the wheels in a factory turning by shovelling coal into a furnace." The stoker defeats the medieval engineer's dictum "while being still subject to the law of the conservation of energy", and the way he does it is by splitting the process into two stages with two energy budgets that can vary independently. Conservation is not violated. It is obeyed twice, and the amplification appears between the stages.
So the bound does not forbid amplification. It tells you where the amplification comes from.
Read the result the last post ended on in that light. Yue and colleagues compared models before and after the reinforcement-learning stage that is supposed to teach reasoning, by asking each the same questions many times. Ask once and the trained model wins. Allow each a large number of attempts and count a question as solved if any attempt succeeds, and the untrained model comes out ahead; the reasoning paths the trained model produces "are already included in the base models' sampling distribution". The field read this as a disappointment about post-training. Under the bound it is not a disappointment at all. It is a statement about which stage the information entered. It went in during pretraining, on the corpus, which filled the bunker; the later stage is the stoker, and a stoker can only shovel coal that is already there.
Which tells you what to do about it. In 1952 Ashby had asked whether a machine could outplay its designer4, and answered that it could, on one condition:
"the designer must construct the machine so that it can receive and use information not provided by him in detail."
The reinforcement stage admits no other source: the checker selects among the base model's own samples. Yue's other finding is the condition met. Distillation, which admits a teacher, "can introduce new reasoning patterns from the teacher and genuinely expand the model's reasoning capabilities." The answer is not a cleverer second stage. It is a second source. The distilled model was given information the base model never held, in the teacher's outputs; the reinforcement stage was given none, and had only the base model's own samples to choose among. What Yue calls the reasoning boundary is the bound, seen from inside the model. Run the sourceless loop for generations and the narrowing becomes Shumailov's model collapse, a 2024 Nature paper: "models start losing information about the true distribution, which first starts with tails disappearing." Their worked example starts from a paragraph on the master masons who built medieval parish church towers, and by the ninth generation of models trained on models it reads, hyphen tokens and all: "architecture. In addition to being home to some of the world's largest populations of black@-@tailed jackrabbits, white @-@ tailed jackrabbits, blue@-@tailed jackrabbits, red@-@tailed jackrabbits, yellow@-." Each generation keeps a smaller part of the last, until there is one response. Ashby had the shape of that in 1958, in a paper on habituation: under a regularly repeated disturbance "there is a fundamental bias in favour of the smaller," and the one thing that reverses it is a disturbance from outside.
What has to be got in, in his sense, is information the designer did not supply, and it comes from one of two places: a store somebody else filled, which is the corpus or the teacher, or the world itself, met by trial, which is the kitten's mice and the next post's subject.
Coal¶
Ashby's engineers accepted a limit about energy, and the analogy can be run one step further than he ran it, because the objection to the bound is already on the table. Nearly all the energy on earth is sunlight, direct or stored. Coal is sunlight captured by past life and laid down, dense and burned once. The written record is the world's information captured by past people, by their trials against it, and laid down; training on it is burning coal. The trials that filled the seam were done long ago, by centuries of people, which is why the corpus could be drawn on without repeating them, and it is why the seam does not refill. Yue's result is the seam seen from inside the model; Shumailov's is what happens when the ash is burned. The reply is obvious and it is the right one: go back to the sun. Send a billion agents to a trillion mice and harvest the world directly, as the corpus once harvested it. That is Silver and Sutton's programme, the era of experience, and it is the second of the three moves the next post examines. The bound does not forbid it. It says three things about it.
First, the rate. The corpus was fast to use because its trials were already done; harvesting live pays for every trial in time, and Ashby's second section in 1956 is an accounting of how long trials take. A billion agents in parallel shorten the time. They do not enlarge the world, and a trillion mice hold what the mice hold. Second, the criterion. The mice prune on survival in the mice's world, not on the sponsor's goal, so what comes back is whatever kept the agent running there. In the case the next post opens on, that was the grader's approval. Third, the recursion. A world increasingly made of agents' output is Shumailov's loop at the scale of the internet: panels pointed at each other's lamps, and every conversion loses some, however many panels are added. The next post has that measured. One caution on the analogy, since the word invites it. Ashby set his law beside the conservation of energy for what accepting a limit did to engineering, and that is the use made of it here. The law itself is not a conservation law. It is a limit of the second-law kind, an inequality: what comes out of selection can be less than went in, never more, and the sun does not change that. The classification is argued in the paper the fifth footnote points to.
The field re-ran that argument in three days in April 2023. John Wentworth, the field's most careful reader of Ashby, posted a conjecture with no cybernetic source in sight:
"observing an additional N bits of information can allow a system to perform at most N additional bits of optimization. I want a proof or disproof of this conjecture."
That is Ashby's law of requisite variety in its channel form: a regulator can block no more variety than it carries. The disproof came at once, and it is a lock. A safe's combination is a few digits. One observed bit, whether the digits match, releases as much effect as the safe holds, and the thread's conclusion was that there is no limit to what one bit can unlock. The commenter who gave the disproof then counted what the opener has to hold besides the combination. He wrote the sentence the conclusion should have been read against:
"besides the observation region O, we also need to know a certain amount of bits detailing the rules of the world. In its uncompressed form, this knowledge is always greater than the amount of optimization."
The thread found all three things Ashby's argument needs: the bound, which held, since one bit of combination blocks one bit of uncertainty about the door; the amplifier, a small input releasing a large stored effect, which is the Aladdin's cave; and the ledger, which the commenter wrote when he counted the bits of the world's rules. Then it drew its conclusion from the second alone. The safe had to be built and filled before the combination could open it, and every one of the unlocked bits had first to be got in. The thread had derived the erg and filed it under practical constraints.
A principle theory¶
One more thing about what kind of statement this is, because it decides who has to do the work. Einstein, writing in The Times in 1919, separated two kinds of theory. Constructive theories build a picture of complex phenomena out of simpler assumed elements — the kinetic theory of gases building heat out of moving molecules. Principle theories do the opposite:
"The elements which form their basis and starting-point are not hypothetically constructed but empirically discovered ones, general characteristics of natural processes, principles that give rise to mathematically formulated criteria which the separate processes or the theoretical representations of them have to satisfy. Thus the science of thermodynamics seeks by analytical means to deduce necessary conditions, which separate events have to satisfy, from the universally experienced fact that perpetual motion is impossible."
The missing knowledge here is not a constructive theory of what a transformer computes. That may come, and they will deserve a Nobel for it (preferably in Cybernetics). What is missing is principle-level, and principles do not need a genius. They need somebody to notice what everyone already experiences and refuse to let it be violated. Perpetual motion is impossible. Selection is bounded by the information received and processed5. Neither is a discovery in the sense of something no one had seen; it's a principle not to be violated.
The principle is already on the field's own shelves, filed under other names, and Harry Law's histories tell three of them. Fisher's theorem of 1930 says a population improves only as fast as it holds variance in fitness, and the Wright–Fisher model says how fast selection uses that variance up; evolutionary computation still calls keeping it the central problem. That is Yue's result stated for populations. Vapnik's bound, carried out of the Moscow Institute of Control Sciences, a cybernetics institute, says how well a learner of a given capacity can be expected to generalise from what it was given: what comes out is limited by what went in. At Bell Labs in the 1990s his camp and the neural-net camp sat in one department, and the second won by varying the recipe and keeping what worked.
Read with the 1960 rule, the two camps were its two halves. Vapnik's used what was known to narrow the field; the other did as it pleased inside what was left, and let the trials win the rest. The field remembers a winner and a loser where Ashby had written one rule. Lighthill, defending his report on television in 1973, said the real advances of the computing age belonged to automation,
"feedback control systems that act to reduce some change in quantity from its desired value."
That is a quantity held at a value against the world: regulation. Law reads such distinctions as ones that "don't make much sense to the modern reader," since a neural network also reduces a quantity, the loss. It reduces it in training, in a world that does not move, and then stops. Lighthill's systems reduce it in service and never stop. He is remembered for missing the point because the field that followed had no word for the point he made. That is the first post's claim observed rather than asserted. A distinction that has merely been renamed is still acted on where it matters; a lost one reads, to a careful modern reader, as a sentence that makes no sense, and Law's is the primary source for that reading. It is the second exhibit. The first was Wei's dial, and the third post counts the rest.
Three vocabularies. Two of them, Fisher's and Vapnik's, state the bound for their own domain, and the third, Lighthill's, states the distinction the bound is for, between a quantity reduced once and a quantity held. Ashby stated the bound in general and dated it to 1961. One is kept as a problem, one is judged not to have mattered, one is judged a confusion, and in none of the three is it read as a principle, a limit on anything that learns.
Ashby's own summary of what accepting the bound does to your work is the sentence I would put on the wall.
"We work now, not to find a magic process that will give us all in a flash but to develop one that achieves the goal with reasonably high efficiency."
Which raises the question this series has been working toward. If you accept the bound, you have accepted that whatever your system does at the far end was got in somewhere, and that a system placed in a world it was not optimised against will meet disturbances it has no information to deal with. What does accepting that oblige you to build? Not a trick; the sentence on the wall has ruled that out. It obliges a loop with a person in it who holds the goal: a channel that reads the world for that goal, the authority to stop or redirect the system while it runs, and the training that keeps the person's judgement fit for purpose. The person cannot be shipped. The loop can, and one industry already has. That is the last post, and it opens on July, which is what happened without it.
Sources¶
The exchange. Ali Rahimi and Ben Recht, "Reflections on Random Kitchen Sinks", NIPS test-of-time award talk, 5 December 2017, text. Yann LeCun's reply was posted the following day; it is quoted here from Synced's report, 12 December 2017.
Mokyr. The Gifts of Athena: Historical Origins of the Knowledge Economy, Princeton, 2002, for propositional and prescriptive knowledge and the epistemic base. The prize: Royal Swedish Academy of Sciences, press release, 13 October 2025, nobelprize.org — Mokyr "for having identified the prerequisites for sustained growth through technological progress", Aghion and Howitt "for the theory of sustained growth through creative destruction". See also his "The Intellectual Origins of Modern Economic Growth", Journal of Economic History 65(2), 2005, 285–351.
Scaling. Jordan Hoffmann et al., "Training Compute-Optimal Large Language Models", arXiv:2203.15556, 2022 — the Chinchilla paper.
The chart, and the proposal. Ashby, "Design for an Intelligence-Amplifier", 1956 (below), §1 — Figure 1 is his, "after Wechsler"; the paper's pp. 261–262 in the Conant reprint. J. McCarthy, M. L. Minsky, N. Rochester and C. E. Shannon, "A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence", 31 August 1955, McCarthy's transcription at jmc.stanford.edu; reprinted in AI Magazine 27(4), 2006, 12–14, the conjecture at p. 12. Solomonoff's notes on Ashby's talk and on McCarthy: Grace Solomonoff, Ray Solomonoff and the Dartmouth Summer Research Project in Artificial Intelligence, 1956, Oxbridge Research, 2016, pp. 11–12, PDF, as she transcribes them; "fewer ideas" is her transcription of Ray's notebook.
The thread. John Wentworth, "Fixing The Good Regulator Theorem", LessWrong, 9 February 2021; "How Many Bits Of Optimization Can One Bit Of Observation Unlock?", LessWrong, 26 April 2023 — the conjecture and the request are the post's opening; dr_s, "One Bit of Observation Can Unlock Many of Optimization: But at What Cost?", LessWrong, 29 April 2023 — the quotation is its Conclusions paragraph. Both are held in the author's LessWrong corpus at `Research/LessWrong/`. The channel form of the law is An Introduction to Cybernetics §11/11.
Collapse and habituation. Ilia Shumailov et al., "AI Models Collapse When Trained on Recursively Generated Data", Nature 631, 2024, 755–759 — the quotation is the abstract's; the church towers and the jackrabbits are Example 1, p. 758, printed as the paper prints the model's output. Ashby, "The Mechanism of Habituation", in Mechanisation of Thought Processes, HMSO, 1959, 95–118 — the bias and the de-habituation sentence are the Summary, p. 95. The two share a shape; that they are one theorem is not claimed here.
Ashby. "Computers and Decision Making", New Scientist 7, 1960, 746; in Conant, ed., Mechanisms of Intelligence, pp. 179–180 for the rule and the trial. "Can a Mechanical Chess-Player Outplay its Designer?", British Journal for the Philosophy of Science 3(9), 1952, 44–57 — the random-number table and "chaotic information is by no means useless" are pp. 290–291 of the Conant reprint; Bigelow's objection is in the discussion of "Mechanical Chess Player", in H. von Foerster, ed., Cybernetics: Transactions of the Ninth Conference, Macy Foundation, 1952, 151–154, as reported in Peter Asaro, "From Mechanisms of Adaptation to Intelligence Amplifiers: The Philosophy of W. Ross Ashby", in Husbands, Holland and Wheeler, eds., The Mechanical Mind in History, MIT Press, 2008; "Design for an Intelligence-Amplifier", in C. E. Shannon and J. McCarthy, eds., Automata Studies, Princeton, 1956, 215–234; "What Is an Intelligent Machine?" (1961) — the definition is its abstract, p. 295 of the Conant reprint, and "intelligent is as intelligent does" p. 297; the kaleidoscope and the winnowing are "Design for an Intelligence-Amplifier", pp. 263–264; and "Cybernetics Today and Its Future Contribution to the Engineering Sciences", address to the Foundation for Instrumentation Education and Research, New York, 1961 — all collected in Roger Conant, ed., Mechanisms of Intelligence: Ross Ashby's Writings on Cybernetics, Intersystems, 1981, whose pages are given here: the electric fluid, 1952, pp. 282–283; the medieval engineers and the stoker, 1956, p. 264; the gusher, 1961 address, p. 328, the perpetual-motion passage pp. 330–331, "We work now" p. 331; "If we accept the limitation", What Is an Intelligent Machine?, p. 305; the Aladdin's cave line is the classroom handout "Ashby Says", pp. 425–427. An Introduction to Cybernetics, Chapman & Hall, 1956, is at archive.org; footnote 5 quotes §7/7 (p. 126) and §§9/6, 9/11–9/13 (pp. 166, 173–177).
The result. Yang Yue et al., "Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?", arXiv:2504.13837, 2025.
The era of experience. David Silver and Richard Sutton, "Welcome to the Era of Experience", 2025, PDF; the trial-time accounting is "Design for an Intelligence-Amplifier", 1956, Section II, §§7–11 (Conant reprint pp. 270–278), listed under Ashby above.
The bound in the field's own histories. Harry Law, Learning From Examples, 2025: "Unnatural Selection" (AI Histories #16, Fisher), "Bell Labs' last trick" (#20, Vapnik), "Sorcerer and cell" (#18, Lighthill), learningfromexamples.com. R. A. Fisher, The Genetical Theory of Natural Selection, Oxford, 1930; V. N. Vapnik and A. Ya. Chervonenkis, "On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities", Theory of Probability and its Applications 16(2), 1971, 264–280; James Lighthill, "Artificial Intelligence: A General Survey", Science Research Council, 1973.
Einstein. "What Is the Theory of Relativity?", The Times, 28 November 1919; reprinted in Ideas and Opinions, Crown, 1954, 227–232.
The author's working papers, all at rsorourke.com/publications: The LLM as Variety Transducer (Kybernetes, published); Rebuilding Ashby's Amplifier; 70 Years Ago Ross Ashby Designed an 'Intelligence-Amplifier'. Have We Finally Built It?; The Machine That Outplayed Its Designer; The Status of Requisite Variety, Revisited; and, in preparation, a paper on how the machine's half of the loop, which the alignment field calls corrigibility, and the sponsor's half meet in Ashby's apparatus.
Footnotes
- ↩
The annotations are from my paper in Kybernetes, "The LLM as Variety Transducer: Ashby's Intelligence-Amplifier Architecture and Beer's Viable System Model as a Diagnostic and Predictive Framework," 2026, doi 10.1108/K-03-2026-0482. Its argument is the one this series assumes: the language model realises Beer's variety transducer; Ashby's amplifier is the coupled system in which the human and the model together regulate the problem, not the model alone; and the human supplies every function the model cannot, which in Beer's letters are Systems 3, 3*, 4 and 5. The chart's red region marks where that coupled system is meant to operate: to the right of the human distribution, which is the point of building an amplifier at all.
- ↩
The paper's route into McCarthy's volume is in Ronald Kline's account of the Automata Studies correspondence. Shannon and McCarthy had invited the Macy and Ratio Club people in 1953, Ashby among them, and Ashby "interpreted the invitation from Shannon and McCarthy as being an automatic acceptance for his paper." When McCarthy proposed to send it to referees, Ashby wrote back on 16 June 1954: "I am afraid I cannot consent to having it 'refereed,' as it is on a subject to which I have devoted a great deal of work, and on which I can now speak with some authority. Frankly, I know of no referee to whose opinion I would defer (though I would, of course, be prepared to answer specific objections)." MacKay, by contrast, welcomed referee reports. McCarthy and Shannon accepted both papers unrefereed. McCarthy's verdict on the volume that summer, to Shannon, was that the Ratio Club papers "contribute heuristic ideas of some value but none lead directly to a solution of the problem of thinking automata." When the Dartmouth invitations were drawn up the following year, "despite their difficulties working with Ross Ashby during the Automata Studies venture, McCarthy and Shannon put forward Ashby as 'the most attractive foreign name.'" Ronald R. Kline, "Cybernetics, Automata Studies, and the Dartmouth Conference on Artificial Intelligence", IEEE Annals of the History of Computing 33(4), 2011, pp. 7, 8 and 9; the letter is Ashby to McCarthy, 16 June 1954, Shannon Papers, Library of Congress, box 1.
- ↩
The search problem, in McCarthy's sense, is how a machine finds a solution without trying everything. Solomonoff's answer, from 1960 onwards, was algorithmic probability: weight every candidate hypothesis by the length of the shortest program that produces it, so that the simple ones are tried first ("A Formal Theory of Inductive Inference", Information and Control 7, 1964, 1–22 and 224–254). Leonid Levin showed in 1973 that a search which allots trial time in proportion to those weights is, up to a constant, as fast as any search can be. Solomonoff worked that out as a practical method in "Optimum Sequential Search" (Oxbridge Research, 1984) and spent the rest of his life on the constant. His 2009 description of the system is the note from Dartmouth answered in his own terms. Give the machine a simple problem and let Levin search find a solution. Then let the machine update itself "by modifying the reference machine so that the solutions found will have higher a priori probabilities", and continue "with a sequence of problems of increasing difficulty, updating after each solution is found" ("Algorithmic Probability: Theory and Applications", in Emmert-Streib and Dehmer, eds., Information Theory and Statistical Learning, Springer, 2009, p. 20). Start slowly and work up, which is what his notes say Ashby wanted, with the memory of previous solutions that his notes say the homeostat lacked.
- ↩
Bigelow's objection and Ashby's answer are the hinge of a working paper of mine, The Machine That Outplayed Its Designer. Ashby's 1952 question was whether a machine could outplay, not merely beat, its designer, and he set the condition in Shannon's bits: Descartes' dictum that a machine cannot put out more design than went in holds for a noiseless transducer and is broken by two information sources, the rules and "the inpouring stream of mutations." The designer must therefore write into the specification "admit other information," and aim "not at a machine that will play chess, but at a machine that can make trials, and select." At the Macy conference he drew the line himself: "I don't say 'beat' its designer; I say 'outplay.'" That rules out Deep Blue, brute force with a hand-built evaluation, and rules in AlphaZero, which won on the game he named from a reservoir the designer admitted but did not specify. The paper argues that this is a quantified prediction that waited sixty-five years for its instance, and that the instance came from Ashby's own lineage rather than the programme that dismissed him at Dartmouth. rsorourke.com/publications/2607-chess. The Macy wording is as reported in Peter Asaro's 2008 chapter, cited above.
- ↩
Where the information comes from is worth spelling out, because Ashby's own chain has four links and each adds something the last did not contain. A set of distinguishable elements has variety: the number of them, or its logarithm to base 2 (An Introduction to Cybernetics, §7/7). Probability arrives only with a way of drawing from the set: "If a set has variety, and we take a sample of one item from the set, by some defined sampling process, then the various possible results of the drawing will be associated with various, corresponding probabilities" (§9/11). His example is a set of traffic lights on for 25, 5, 25 and 5 seconds, which a motorist arriving at irregular times finds in the four states about 42, 8, 42 and 8 per cent of the time, "if this particular method of sampling be used." Shannon's entropy is that variety weighted by those probabilities; when they are equal, "H is then equal to log n, precisely the measure of variety defined in S.7/7" (§9/12). For a source that runs on, the weights are the proportions of time it spends in each state once it has settled to equilibrium, so the sampler is time itself (§§9/6, 9/12). Information is what such a source transmits, and it "cannot be transmitted in larger quantity than the quantity of variety allows" (§9/11). Ashby's conditions on the arithmetic are conditions on the process: a complete set of probabilities, a Markovian source, and a system run long enough to reach its equilibrial densities, all "by no means so commonly fulfilled in biological work" (§9/13). He reserves the word entropy for Shannon's measure and warns against playing "loosely, and on a merely verbal level, with the two entropies of Shannon and of statistical mechanics" (§§9/11, 9/13). So the distinctions are the observer's, the sampling procedure is the observer's, and the equilibrium is the process's. Everything after those choices is arithmetic, which is why the case that Ashby's law is a principle theory in Einstein's sense, made in The Status of Requisite Variety, Revisited, locates the law's content in the partition and not in the inequality.