Why we never farmed lions for food, and the energy arithmetic of the AI race

14 September 2026 · Law, privacy and artificial intelligence

Why we never farmed lions for food, and the energy arithmetic of the AI race

Every mammal domesticated for meat lives on plants. Cattle, sheep, pigs and goats. No predator ever made the list, not the lion, the tiger or the wolf, and the reason is not the danger of raising one. The reason is a simple piece of energy arithmetic, and the same arithmetic is back today in the data centres of artificial intelligence.

Fourteen out of 148

Jared Diamond devoted chapter nine of Guns, Germs, and Steel to the question, under the title Zebras, Unhappy Marriages, and the Anna Karenina Principle. The title points to the opening sentence of Anna Karenina, that all happy families are alike and each unhappy family is unhappy in its own way. A happy marriage requires success in a whole series of things, and failure in any one of them defeats the rest. Diamond applies that structure to domestication.

He lists six conditions: diet, growth rate, breeding in captivity, disposition, tendency to panic and social structure. Each has its own list of failures. An elephant reaches maturity at twelve, so anyone raising one for food waits twelve years for a single yield. The cheetah was tamed for hunting by Egyptian kings, Assyrians and Indian princes over thousands of years, and none of them managed to breed it in captivity, because courtship requires a prolonged chase in open country. The zebra looks like an African horse ready made, and it bites and does not let go. A deer bolts and injures itself when penned, and a horse stops.

Of 148 species of big wild terrestrial herbivorous mammals that were candidates for domestication, fourteen passed every condition.

Continent Candidates Domesticated
Eurasia 72 13
Sub-Saharan Africa 51 0
The Americas 24 1 (llama and alpaca)
Australia 1 0

Eurasia had both the most candidates and the most that passed, for two reasons. It is the largest landmass and the one with the widest range of habitats, so more species of big mammals lived there to begin with. In the Americas and in Australia most big mammals died out at the end of the ice age, around the time people arrived, and many of the possible candidates were no longer there. Even so, in Eurasia thirteen of seventy two passed, fewer than one in five, and in sub Saharan Africa none of the fifty one passed. The number of candidates alone does not explain the result.

Carnivores do not appear among the 148 candidates at all, and the reason is the first condition on the list. For them the question stops at diet, before disposition or social structure is ever reached. Of the six conditions, the one about energy is the strongest filter.

Conversion efficiency is what decides

When the energy cost of a way of doing things rises, the outcome is decided by conversion efficiency, by how much use is extracted from each unit of energy. Size stops being the advantage. In the animal world the winner is the cow, which converts cheap grass into protein, and not the lion, whose prey has to be raised for it first.

In artificial intelligence the corresponding role belongs to the frontier models. A frontier model is the newest and largest model a company runs at a given time. Its training took months of computation across whole data centres, and it is also the most expensive to operate, because every question sent to it uses more electricity than the same question sent to a small model. Each family has one such model, the highest version of GPT, of Claude and of Gemini, and alongside it the same family offers smaller and faster versions. Every few months a new model replaces the previous one at the top. The frontier model is what sets what can be done at all at any moment, and these models will keep growing. Most of the requests sent to a model on an ordinary working day do not call for the highest capability, and every such request draws electricity in proportion to the size of the model that answers it. Most of the actual work therefore moves to smaller models, which give a good enough answer for a fraction of the energy, and the frontier models are kept for what only they can do.

How much energy is lost at each step of the food chain

The number behind all of this is known and quantified. Diamond puts it this way:

Each time one animal eats a plant or another animal, the conversion of food biomass into the consumer’s biomass is far less than one hundred percent efficient, typically around ten percent.

A cow eating grass converts about a tenth of the plant’s energy into meat. To raise a lion for food you would have to raise a herd of cattle for it, and whole fields of feed for that herd, and each step in the chain loses most of the energy again. A global synthesis published in Science Advances in July 2026 found efficiency in practice lower still than the ten percent rule: about six percent on average, and under two percent in terrestrial systems. In modern agriculture the ratio is expressed as the feed conversion ratio, and according to Our World in Data about twenty five kilograms of feed are required for one kilogram of edible beef. An animal one step higher in the chain costs an order of magnitude more.

The same arithmetic in the data centres

In artificial intelligence the arithmetic has the same shape. To improve a model by a given amount, the compute required grows faster than the rate of improvement, so each further step of improvement costs more than the one before it.

That price is not peculiar to machines. A human brain is about two percent of body weight and consumes about a fifth of the energy the body burns at rest. Intelligence is expensive in whatever substrate it runs, and the price rises the more is asked of it.

A practical conclusion follows. When each further degree of capability costs more than the last, running a frontier model on a routine question pays the full price for capability that is not needed, rather like sending a lorry to fetch milk from the corner shop. The wider the gap in cost, the more is saved by matching the model to the task than by any improvement in hardware. The result shows up in the electricity market: technology companies have signed long term purchase agreements for the output of entire nuclear reactors. Microsoft and Constellation for 835 MW from Unit 1 at Three Mile Island in September 2024, Amazon and Talen Energy for 1,920 MW from the Susquehanna plant in June 2025, and Meta and Vistra for 2,609 MW over twenty years in January 2026.

A contract is not electricity. Running a new data centre takes six things, and money and chips are no longer the constraints. The table shows which of them are ready, which are not, and how long the rest take to obtain.

What is needed Status Waiting time
Money for the contracts Ready The contracts are signed
Chips Ready Were short until 2024
Power from the reactors ordered Missing Commissioning expected 2030 to 2035
Large power transformers Missing Two to four years
Grid connection Missing Set by the queue, not the manufacturer
Water for cooling and planning consent Missing Set by the local proceeding

Transformer lead times come from a shortage of grain-oriented electrical steel and of skilled winding labour, and no amount of money speeds those up. Alongside the deals for existing reactors, commitments have been signed for small modular reactors (SMR) not yet built: Google with Kairos Power, Amazon with X-energy, Meta with TerraPower and with Oklo.

How long this price can go on being paid is an open argument. Ilya Sutskever, co-founder of OpenAI, said in November 2024 that the returns from expanding pre-training had reached saturation, and repeated it at the end of 2025, saying the field was moving from the age of scaling to the age of research. Sam Altman has said that scaling has met no limit. Dario Amodei, co-founder and chief executive of Anthropic and the holder of a doctorate in biophysics from Princeton, writes in his essay The Adolescence of Technology, from January 2026, that behind the public volatility there continues a steady rise in capability. A neutral research body such as Epoch AI presents the question as open and empirical. While the ceiling cannot be located, betting on size is expensive, and investing in efficiency pays either way.

The models that actually get chosen

When it became clear that lions could not be farmed, people did not stop eating meat. They concentrated on animals that convert cheap feed into protein at high efficiency. The same move is under way now in artificial intelligence, in three directions.

The size of a model is measured in parameters, the internal numbers in which what it learned during training is held, and every use of it consumes computation and electricity in proportion to that size. An architecture is the way a model is built internally. In a Mixture of Experts architecture the model is divided into separate parts, and only a few of them are activated for each piece of text it processes. A model holding 671 billion parameters therefore works each time with a small part of itself. The division is not by subject, and it is decided while the text is processed, piece by piece. DeepSeek-V3 holds 671 billion parameters and activates 37 billion at a time, about five and a half percent, which is the knowledge of a large model at the running cost of a far smaller one.

The second direction is running on the device. In June 2026 Apple presented a model of 20 billion parameters in a sparse architecture, activating between one and four billion at a time and running on the device itself without sending anything to the cloud. Local execution saves the data centre’s electricity and shortens response time, and it also changes the status of the material: text that never leaves the device is not transferred to a third party, is not stored on a provider’s server, and is not subject to that provider’s terms of use.

The third direction is a division of labour by type of task. Research by NVIDIA from June 2025 argues that small language models are capable enough and more economical for most of the calls made in software where the model carries out a sequence of actions and is called dozens of times within a single task, and proposes systems that combine different sizes.

These complement the frontier models, and do not replace them. A small model on the device for routine tasks, a large model in the cloud for complex ones. What decides which model a task is sent to is what the task requires and what it costs to run the model that answers it.

Where it pays to be clever

The parallel here concerns the energy arithmetic alone. A data centre is not an organism, and what holds for a food chain is not evidence of what a market will do. It would be enough for the running cost of frontier models to fall fast enough, and efficiency would stop being a consideration.

But Diamond’s argument is not about animals. It is that what a continent had available decided what happened there afterwards. Sub-Saharan Africa had fifty one candidates and not one of them passed, so it had no draught animals, no carts and no agriculture resting on manure and pulling power. The consequences ran for thousands of years. The difference was not in the people. It was in what lay to hand.

In artificial intelligence the equivalent input is cheap electricity and a grid able to carry it, and it opens two different ways to compete. Whoever has both can buy more chips, run a longer training job and build a larger model, and their advantage is the scale of the investment. Whoever does not has to get comparable capability out of less hardware, and their advantage is the quality of the design: which parts of the model are activated at a time, how it is trained, and what can be made to run on the hardware that already exists.

The difference between the two is not only the size of the budget. Where compute is abundant, an idea that makes a model twenty percent more efficient saves money that would have been spent anyway. Where compute is limited, that same idea is the difference between a model that can compete and a model that does not exist. Take two teams that arrive at the same idea, one that makes running the model twenty percent cheaper. For the team with chips to spare, the idea reduces a bill it is paying anyway. For the team without them, that same idea is what makes it possible to run a model that otherwise would not run at all. That is why the constrained team puts more time and more people into efficiency research, and why the solutions appear there first.

DeepSeek is the example. The company built its model under restrictions on the export of advanced chips, and produced an architecture that holds 671 billion parameters and activates 37 billion at a time. The result is a model on the scale of the frontier at the running cost of a far smaller one. The constraint dictated the route to the result.

And there is the other side of it. Ideas born of a constraint do not stay with whoever was constrained. Mixture of Experts architectures are now used by companies with electricity to spare, because they are cheaper for them too. There are two ways to grow, to add electricity or to get more capability out of each unit of it. The first depends on transformers, grid queues and planning permits, and its timetables are measured in years. The second depends on an idea, and whoever improves conversion efficiency grows without adding a single megawatt.

Further reading

Sources

Jared Diamond, Guns, Germs, and Steel, chapter 9, including Table 9.2 · Science Advances, 8 July 2026, energy transfer efficiency in food webs · Our World in Data, Meat and Dairy Production · Dario Amodei, The Adolescence of Technology (January 2026) · Ilya Sutskever interview, November 2025 · Epoch AI, Scaling · announcements of Constellation (September 2024), Talen Energy (June 2025), Vistra (9 January 2026) · Carnegie Endowment, Beyond the Hype: Assessing Hyperscaler Nuclear Commitments (June 2026) · PwC via Reuters Events, transformer lead times (May 2026) · DeepSeek-V3 Technical Report, arXiv 2412.19437 · Apple, Introducing the Third Generation of Apple Foundation Models (June 2026) · NVIDIA Research, Small Language Models are the Future of Agentic AI, arXiv 2506.02153

A general overview. The data and sources are current as at the date of writing, 14 September 2026.

x
סייען נגישות
הגדלת גופן
הקטנת גופן
גופן קריא
גווני אפור
גווני מונוכרום
איפוס צבעים
הקטנת תצוגה
הגדלת תצוגה
איפוס תצוגה

אתר מונגש

אנו רואים חשיבות עליונה בהנגשת אתר האינטרנט שלנו לאנשים עם מוגבלויות, וכך לאפשר לכלל האוכלוסיה להשתמש באתרנו בקלות ובנוחות. באתר זה בוצעו מגוון פעולות להנגשת האתר, הכוללות בין השאר התקנת רכיב נגישות ייעודי.

סייגי נגישות

למרות מאמצנו להנגיש את כלל הדפים באתר באופן מלא, יתכן ויתגלו חלקים באתר שאינם נגישים. במידה ואינם מסוגלים לגלוש באתר באופן אופטימלי, אנה צרו איתנו קשר

רכיב נגישות

באתר זה הותקן רכיב נגישות מתקדם, מבית all internet - בניית אתרים.רכיב זה מסייע בהנגשת האתר עבור אנשים בעלי מוגבלויות.