What a machine does with it

How many features a scale can carry

Töpfer's radical law is quoted everywhere as a rule of thumb. It is not one: it is a theorem about a size distribution with a Pareto exponent of exactly one half, exact to 1.8 per cent for that population and out by 99.4 per cent for a lognormal one.

Five essays in this ladder ask what happens to a feature that is kept: how long it turns out to be, whether the tolerance survives the projection, what the tolerance promises, what a shared boundary does, and how the area drains away over a population. Which features are kept at all is a different problem, and it has a law.

nF=nASASFn_F = n_A\sqrt{\frac{S_A}{S_F}}

Töpfer’s radical law: the number of features that should survive to a smaller scale is the number at the source scale times the square root of the ratio of the scale denominators. It appears in every generalisation textbook, it is quoted as an empirical rule of thumb, and it is not one.

Töpfer's square root is one line of a family. The fraction of features surviving to a smaller scale, for four stated populations whose size distributions differ only in their exponent. Every one is a straight line on these axes, and the slope of each is its own exponent: 0.3, 0.5, 0.8, 1.2. Töpfer's radical law is the line at 0.5 — the square root — and it is exact for that population and for no other. The law is not a rule of thumb with exceptions; it is a theorem with a hypothesis nobody states.
Fig. 1 The fraction of features surviving to a smaller scale, for four stated populations whose size distributions differ only in their exponent. Every one is a straight line on these axes and the slope of each is its own exponent. Töpfer’s radical law is the line at 0.5 — the square root — and it is exact for that population and for no other.

Where the square root comes from

The derivation is three lines and it is the reason the law is a theorem.

A feature survives to a scale if it is still legible there — and the smallest legible mark has a ground size that grows with the denominator, and legibility is a fixed size on the paper — half a millimetre, say. At a scale denominator SS that is a ground size of 0.0005S0.0005\,S metres, so the cut-off grows in proportion to the denominator.

Now suppose the number of features whose ground size exceeds xx is N(x)=CxαN(x) = Cx^{-\alpha}. Then the fraction surviving from SAS_A to SFS_F is

N(cSF)N(cSA)=(SASF)α,\frac{N(cS_F)}{N(cS_A)} = \left(\frac{S_A}{S_F}\right)^{\alpha},

exactly. Töpfer’s square root is the case α=12\alpha = \tfrac12.

There is nothing empirical about the exponent. It is the exponent of the size distribution, and it can be measured.

Four populations, four exponents

The exponent the selection recovers is the population's own. The exponent fitted to the surviving counts, for four populations of stated exponent and one with none. Each of the four comes back within 0.024 of the value it was built with, at a coefficient of determination above 0.9997. The lognormal fits an exponent too — 1.67 — and fits it badly, at 0.9934, which is the only warning the arithmetic gives that the number means nothing.
Fig. 2 The exponent fitted to the surviving counts, for four populations of stated exponent and one with none. Each of the four comes back within 0.024 of the value it was built with, at a coefficient of determination above 0.9996. The lognormal fits an exponent too — 1.68 — and fits it badly, at 0.9934, which is the only warning the arithmetic gives that the number means nothing.

Build a population with α=0.3\alpha = 0.3 and the selection recovers 0.298. Build one with 0.80.8 and it recovers 0.790. Build one with 1.21.2 and it recovers 1.176.

The small shortfalls at the high end are the finite population running out of large features, not a defect in the argument: at α=1.2\alpha = 1.2 the number surviving to the smallest scale is in the hundreds, and a count in the hundreds is a noisy estimate of a fraction.

For α=0.5\alpha = 0.5 — Töpfer’s own case — the law’s prediction matches the count to within 1.8 per cent at every scale in the sweep, and the residual is that same finite-sample effect. On the population it is a theorem about, the law is not approximately right; it is right.

And on a population it is not a theorem about

The same law on a population it was not a theorem about. A population whose sizes are lognormal rather than power-law — no exponent anywhere in it — put through the same selection. Töpfer's law over-predicts at every scale and the error grows without bound: 42 per cent at one halving and 100.0 per cent at 1:5.0 million, where it predicts 1057 features and 0 survive.
Fig. 3 A population whose sizes are lognormal rather than power-law — no exponent anywhere in it — put through the same selection. Töpfer’s law over-predicts at every scale and the error grows without bound: 42 per cent at one halving of the scale and 99.4 per cent at 1:1,000,000, where it predicts 2,364 features and fifteen survive.

The lognormal is not a pathological choice. It is what a multiplicative process produces, it is the standard model for the sizes of many natural and human things, and it has no exponent at all.

Put through the same selection it gives a curve that bends. Töpfer’s law over-predicts everywhere and the over-prediction grows: 42 per cent at one halving of the scale, 71 per cent at two, 99.4 per cent at a factor of forty.

Predicting 2,364 features and getting fifteen is not a rule of thumb failing gracefully. It is a rule applied outside the hypothesis it needs.

The warning the arithmetic gives, and how weak it is

The lognormal population still fits an exponent. Run a least-squares line through its log-log selection curve and it comes back at 1.68, with a coefficient of determination of 0.9934.

That is a good-looking number. On the four power-law populations the same statistic is above 0.9996, so the two are distinguishable — but only if somebody computes it, and only if they know what value to expect. A fitted exponent of 1.68 with an R2R^2 of 0.993 looks like a measurement.

What it actually is is a straight line through a curve, over a range too short to see the bend. Extending the range is what exposes it, and the range a real generalisation exercise covers is usually a single step.

Why a straight line is the whole test

The log-log plot is doing more work here than it looks, and it is worth saying what it tests.

A power law is the only family for which the fraction surviving depends on the ratio of the two scales and not on either of them separately. That is the scale-invariance property, and it is what makes a single exponent enough: going from 1:25,000 to 1:50,000 keeps the same fraction as going from 1:250,000 to 1:500,000.

Any other distribution breaks that. For the lognormal, the fraction kept in the first halving is 41 per cent and in the fourth is 25 per cent, so there is no single number that answers what fraction survives a halving. The question has an answer only for a power law, and the law’s form presupposes that the question has one.

So the straight line is not a goodness-of-fit convenience. It is the test of whether the law’s central assumption — that the answer depends on the ratio alone — holds for this population at all.

What this changes about using the law

It has a hypothesis and the hypothesis is checkable. Take the source layer, count features above a sequence of size thresholds, and plot the counts against the thresholds on log-log axes. A straight line means the law applies and its slope is the exponent to use; a curve means it does not.

The exponent is not always a half. Populations with α\alpha well away from a half are ordinary. Settlements, lakes and islands have been reported at exponents from 0.4 to well above 1 in different regions, and using 0.5 for a population at 1.2 under-selects by a factor that grows with every step of scale.

Using it in reverse is worse. The law is sometimes inverted to answer what scale can carry this many features, and inverting a wrong exponent compounds the error into the scale itself.

What the law is really a statement about

It is worth being clear that Töpfer’s law is not about cartography at all.

It says: if a population’s counts above a threshold follow a power law, then the count above a threshold that is scaled by a factor is the original count times that factor to a fixed power. That is a property of power laws, and the cartographic content is only the identification of the threshold with the scale denominator.

So every criticism of it as an empirical rule misses the point twice. It is not empirical, and it does not fail because cartographic selection is more subtle than counting — it fails when the population is not a power law, which is a fact about the data and not about the drawing.

Töpfer's square root is one line of a family. The fraction of features surviving to a smaller scale, for four stated populations whose size distributions differ only in their exponent. Every one is a straight line on these axes, and the slope of each is its own exponent: 0.5, 0.7, 1. Töpfer's radical law is the line at 0.7 — the square root — and it is exact for that population and for no other. The law is not a rule of thumb with exceptions; it is a theorem with a hypothesis nobody states.
Fig. 4 The same experiment on three populations closer together in exponent. The lines are still straight and still separated, so the recovery is not a coarse discrimination between wildly different cases — the method resolves the exponent to a few hundredths, which is well inside the range real feature populations differ by.

The exponent is a fact about the ground

There is a positive result buried in this and it is worth extracting, because it turns a caveat into a measurement.

If the selection curve of a real layer is a straight line, its slope is a property of the landscape rather than of the cartography. It says how the sizes of that layer’s features are distributed, over the whole range the selection sweeps, and it says it from counts alone — no size measurements, no fitting to a distribution, no assumption about a functional form beyond the straightness the plot itself checks.

That is a cheap and unusual instrument. Measuring a size distribution directly needs the sizes; measuring its exponent this way needs only the counts above a sequence of thresholds, which is what a generalisation pipeline computes anyway on its way to deciding what to keep.

So the honest use of Töpfer’s law is inverted from the usual one. Rather than assuming an exponent and predicting a count, measure the counts and read off the exponent — and then the prediction is a prediction rather than an assumption, and the layer has told the pipeline about itself.

What "10 metres" means when the metres are Mercator's. A simplification tolerance is nearly always typed as a number of the projection's own units, because that is what the geometry is in when the algorithm runs. On Mercator those units are ground metres on the equator and nowhere else: 10 m of map is 10.00 m of ground at 0°, 5.00 m at 60° and 0.87 m at 85°. The same line in the same file is generalised 11.5 times more finely at the top of the map than at the bottom.
Fig. 5 The other half of a generalisation specification, and the one this ladder priced two rungs ago: a tolerance quoted in metres is a tolerance in whatever units the geometry was in when the algorithm ran. Selection and simplification are two decisions with two thresholds, and a single “generalisation level” is neither.

Selection is not the only thing that decides the count

The law predicts how many features a scale can carry, and a real generalisation decides how many it does, which is a different quantity.

A cartographer keeps a small feature because it is a landmark, which is a purpose rather than a property, drops a large one because it is redundant, and merges two into one. Every one of those is a judgement the law does not model, and the law’s own literature has always said so: it is a starting point for the count, to be adjusted.

What the derivation adds is that the starting point has a hypothesis. Adjusting a number that came from a wrong exponent is adjusting from the wrong place, and the adjustment a cartographer makes is small compared with a factor of sixty.

Vertices kept, for the same ground tolerance. The ground answer keeps 129 vertices at every latitude, because the tolerance is a ground distance and the region is the same region. A pipeline that simplifies in degrees keeps more and more of them the further north the region is — 367 at 80°, a factor of 2.84 — because a degree of longitude is a shorter distance up there and the algorithm is comparing degrees. It is not simplifying badly; it is simplifying something else.
Fig. 6 The other half of the same operation, which this ladder priced four rungs ago: for the features that are kept, how many of their vertices survive a stated ground tolerance. The two decisions are made by different machinery and are usually reported as one number — a “generalisation level” — which is neither.

The legibility threshold is the other assumption

The derivation puts a fixed size on the paper and everything follows from it. That is close to true and not exactly true.

The smallest legible mark depends on the medium — a printed sheet, a screen at some pixel density, a screen at another — and on the symbol. A road drawn as a line has a minimum width rather than a minimum length, which is the failure the screen ladder measures from the other end. A point symbol has a minimum spacing rather than a minimum size, which is the pixel’s own footprint on a screen map.

So the cut-off is a family of thresholds rather than one, and a layer with several symbol types has several selections running at once. The law applies to each of them separately and to their union not at all, because a sum of power laws with different exponents is not a power law.

The refusal, and what it catches

The check this rung ships has two halves and the second is the one that matters.

The first requires that four populations built with stated exponents come back at those exponents to better than 0.06 and at a coefficient of determination above 0.999. That establishes the machinery: the selection counts what it is supposed to count and the fit reads what it is supposed to read.

The second requires that a population with no exponent must fail to fit one at that quality. It does, at 0.9934 rather than 0.9996, and the margin between those two numbers is the whole detection.

Without the second half the check would pass on a fitting routine that reported 0.999 for anything, which is a real failure mode for a two-parameter fit on eight points. With it, the routine is known to be capable of rejecting, and the four passes above are four measurements rather than four instances of an instrument that only ever says yes.

Where the model stops

The populations here are stated and drawn from with a fixed seed, so the counts are reproducible and the exponents are exact by construction. That is the point — it makes the recovery a check on the machinery rather than a claim about any real dataset.

What it does not do is tell anybody what exponent their own layer has. Measuring that needs the layer, and the measurement is the one this rung recommends and cannot perform.

Nor is anything here about which features survive. The law counts and does not choose, and two selections that keep the same number can keep entirely different sets — one by size, one by importance — with the same count and different maps.

What a wrong exponent costs, in features

The size of the mistake is worth one paragraph in the units a cartographer uses.

Going from 1:25,000 to 1:250,000 is a factor of ten in denominator. At an exponent of 0.5 the law keeps 31.6 per cent of the features; at 0.3 it keeps 50.1; at 0.8 it keeps 15.8; at 1.2 it keeps 6.3.

So a layer whose true exponent is 1.2, selected with Töpfer’s half, keeps five times too many features. The sheet is then over-full by a factor of five, which is not a subtle degradation — it is the difference between a legible map and a grey one, and the cartographer fixing it by hand is fixing an arithmetic error rather than exercising judgement.

In the other direction a layer at 0.3 selected with a half keeps 63 per cent of what it should, and the sheet is under-populated in a way that reads as missing data.

The one-step problem

The reason the failure survives is that the exercise is almost always a single step.

Generalising one series to the next is a factor of two, or four, or occasionally ten. Over a single step every distribution looks like a power law, because every smooth positive function looks like a straight line over a short enough interval on a log-log plot — and one step is short enough.

So the law works, locally, on almost any population, with an exponent fitted to that step. What it does not do is transfer: the exponent that describes 1:25,000 to 1:50,000 is not the one that describes 1:250,000 to 1:500,000 unless the population really is a power law, and nothing about the first step reveals whether it is.

That is the sharpest practical statement available. Töpfer’s law with a fitted exponent is safe for one step and unsafe for a series, and a national mapping programme is a series.

Who found it, and when

Fritz Töpfer and Wolfgang Pillewizer published the radical law in 1966, as an empirical regularity observed across map series, and the derivation from a size distribution came later and is usually attributed to the generalisation literature of the 1980s.

The order matters for how the law is taught. Presented as an observation it invites the reading that it is approximately true of maps in general; presented as a theorem it invites the reading it deserves, which is that it is exactly true of one family of populations and says nothing about the rest.

It is worth adding that a selection is not a simplification: the vertices of what survives are thinned separately, by a different threshold, in a different pass. The square root has been so durable partly because a great many real populations do sit near an exponent of a half. That is a fact about the world worth knowing and it is not a licence, and the difference between the two is a log-log plot anybody can make in a minute.

One more thing follows from the framework reading, and it is about which way to be wrong.

An exponent set too high keeps too many features, and the map arrives at the plotter overcrowded. That failure announces itself: somebody looks at the sheet, sees labels colliding and symbols merging, and sends it back. An exponent set too low drops features that would have fitted, and the map arrives clean, legible and short of content. Nobody sends that one back, because there is nothing on the sheet to point at — the missing towns are missing.

So the two errors are not symmetric in how they are caught, only in how they are made. A programme that tunes its exponent by looking at its output will drift downwards, towards emptier maps, one series at a time.

The law’s own defence

Töpfer and Pillewizer were not naive about any of this, and it is worth quoting the shape of their own caveat.

They presented the radical law as one of a family, with a base exponent and a set of multiplicative corrections for symbol type, importance and the map’s purpose — a framework rather than a formula. The square root is the framework’s default, and defaults become the thing everybody remembers.

What has been lost in the retelling is not accuracy but status. A default in a framework invites the question of when to depart from it; a rule of thumb invites the question of how far it is off. The first question has an answer — measure the exponent — and the second does not.

Where the ladder goes next

Six rungs of this anchor treat a feature and a population of features. What none of them asks is what happens to the attribute attached to a feature when its geometry changes — which is a rung on a different ladder and belongs to the essays about stored geometry.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Feature populationGeneralisationLegibilityLognormalPareto exponentPower lawRadical lawRule of thumbScale denominatorSelectionSimplificationTopfer