RunTheSim
What Is a Power Law? The Giant Case Is the Rule

What Is a Power Law? The Giant Case Is the Rule

A power law is the reason averages keep failing you. In a power law the giant case is not the exception. It is the shape. The wider the category you look at, the more one rare case owns it, and the middle you were planning around never existed.

School statistics trains the other reflex. Heights, shoe sizes, exam scores: pick one at random and it sits near the middle, and that middle is a fact you can build on. Apply the same reflex to earthquakes, wildfires, city sizes or attention online and it breaks quietly. Not because the maths gets hard, but because in those systems the average case and the case that matters are two different animals.

Switzerland has a thousand earthquakes a year and no typical one

The Swiss Seismological Service at ETH Zurich records between 1,000 and 2,000 earthquakes a year in Switzerland and the regions next door. People feel 20 to 30 of them.

So what does the average Swiss earthquake look like? Too small to feel. That is a true statement about the catalogue and a useless one for anyone planning a hospital. The events that decide how you build are the ones the average hides: magnitude 5 every 8 to 15 years, magnitude 6 every 50 to 150 years, and Basel in 1356. Each step up the magnitude scale makes events roughly ten times rarer, the ratio known as the Gutenberg-Richter relation. A constant ratio like that has no preferred size, so the data never settles into a middle. Draw the box wider (all of Switzerland, all magnitudes, all century) and you do not get a smoother average. You get more room for the one quake that owns the whole damage total.

25

quakes of magnitude 2.5 or above per year

Out of the 1,000 to 2,000 the network records in and around Switzerland.

Quelle: Swiss Seismological Service
Long-term average, Switzerland and neighbouring regions.

That gap is the power law at work: nearly all the events sit where nobody notices, and nearly all the consequence sits in the handful that get through.

Nobody designs this shape, simple rules produce it

Once the shape is real, the next question is where it comes from. Not from a designer, and not from anything exotic. Take the standard forest fire model: trees grow on a grid at some rate, lightning hits at random, fire spreads to touching trees. Let it run and the burn sizes come out as a power law. Mostly nothing, sometimes a patch, rarely the whole board. Nobody wrote the big fire into the rules. It falls out of the density the small fires leave behind, and on a log-log plot the whole distribution lies down as a straight line.

Simple local rules making a global shape is the same idea behind a reaction diffusion simulation, though there the output is a stripe pattern and here it is a distribution. The lopsidedness shows up far from physics too, in traffic, in sales, in the outcomes of a business simulation browser game.

QuestionBell curvePower law
Where is the typical case?near the averagefar below the average
How big can the biggest get?a few steps from the middleno fixed ceiling
What happens as data comes in?the average settlesthe average keeps climbing
What should you plan for?the middlethe tail

Read the right column as one instruction: stop asking what is typical, start asking how big it gets.

Most things called a power law are not power laws

Here is the honest objection. The term gets stuck on any chart with a long tail, usually by eye, usually without a test. Clauset, Shalizi and Newman tested twenty-four real data sets that had all been claimed as power laws, and found that in some cases the claim holds and in others the power law is ruled out. So a lot of confident power-law talk is decoration.

What survives the objection is the part worth keeping. The practical difference was never the exact exponent, it is the heavy tail. Whether your data is a true power law or a log-normal with a fat end, the average is still unstable and the largest case still runs the outcome. The label is negotiable. The behaviour is not.

This stops holding where a quantity has a physical ceiling. Human height, body temperature, marathon times: there the average is the story, and no amount of extra data will produce someone four metres tall.

So before you average anything, sort it and look at the top of the list first. Then plot it on log-log and see whether it walks down as a straight line. And when you pick a category to stand in, remember the broad one is not just bigger, it is steeper: the wider the field, the more of it one name already holds.

Sorting a list is one thing, watching the shape assemble itself is another.

What RunTheSim is