What a machine does with it

Which features survive is not a sample

The rung below answers how many features a scale can carry and treats the population as a number. Which ones survive is a different question: keeping one feature in ten carries 99.99 per cent of the total length and inflates the median feature by a factor of 95, and the shape of the size distribution survives both exactly.

Assumes How many features a scale can carry.

The rung below this one answers how many. Töpfer’s radical law turns out to be a theorem about a size distribution rather than a rule of thumb, exact to 1.8 per cent for a population with a Pareto exponent of a half and out by 99.4 per cent for one without.

Everything in that rung treats the population as a count. This one asks which members of it are still on the sheet, and what a reader can conclude from them.

Three selection rules, and what each one keeps. Keeping one feature in ten from a stated population whose size distribution has a Pareto exponent of a half — the exponent Töpfer's law is a theorem about. Keeping the largest carries 99.99 per cent of the total size and inflates the median feature by a factor of 95. A random sample keeps the median to 1.068 and carries 5.0 per cent of the total. The two rules are right about different things and there is no rule that is right about both, because the total lives in the tail and the median does not.
Fig. 1 Keeping one feature in ten from a stated population, under three rules. Keeping the largest carries 99.99 per cent of the total size and inflates the median feature by a factor of 95. A random sample keeps the median and carries 5 per cent of the total. The two are right about different things.

The rule nobody chose

A generalisation drops features because they are too small to draw — which is the constraint how many features a scale can carry turns into a law, and which a real edge has a width prices from the other end. That is not a policy; it is the legibility limit doing arithmetic. A feature smaller than about 0.4 millimetres on the paper cannot be shown, so at 1:1,000,000 anything under 400 metres goes, and what remains is everything above a threshold.

So the surviving population is not a sample of the original in any statistical sense. It is the upper tail, cut at a place that moves with the scale.

This is worth stating because the word generalisation invites the opposite reading. A generalised map is routinely described as showing the important features, or a representative selection, or the ones that matter — all of which suggest a choice about relevance. What is happening is a threshold on size, applied to whatever is there.

What was computed, and how

A stated population of twenty thousand features whose sizes follow a Pareto law with exponent α = ½ — the exponent Töpfer’s law is exactly a theorem about. No dataset, and the whole distribution is a closed form a reader can regenerate.

Three selection rules are applied, each keeping the same number of features:

  • keep the largest, which is what a legibility limit forces;
  • keep a random sample, which is unbiased in the statistical sense and unusable as a map;
  • keep the same share of every size band, which is the compromise between them.

And three statistics are measured against the truth over the whole population: the total size the survivors carry, the median survivor’s size, and the fitted tail exponent of the survivors’ own size distribution.

The mean is deliberately absent, and its absence is part of the rung. A Pareto population with α below one has no finite mean: the sample mean is a measurement of the largest draw and not of the population. Töpfer’s own case is exactly the case in which “the average feature” is not a quantity, which is a fact about the subject rather than about the arithmetic.

Two of the three statistics survive

keep the largest a random sample the same share of each band
median / true median 95.4× 1.07× 3.08×
share of the total carried 99.99% 4.98% 5.42%
fitted tail exponent 0.478 0.543 0.522
(the population’s own) 0.497 0.497 0.497

Reading the columns rather than the rows is what makes this a finding rather than a complaint.

The total survives, almost perfectly. One feature in ten carries 99.99 per cent of the total size. A reader estimating how much river, coastline or road there is from a small-scale sheet gets very nearly the right answer, because in a distribution with this exponent the total lives in the tail and the tail is exactly what was kept.

What the survivors carry. The share of the population's total size carried by the features that survive to each scale, under the rule every generalisation uses. Keeping one feature in a hundred still carries 99.93 per cent of the total. The figures beside the bars are the fitted tail exponent of the survivors, against the population's own 0.497: truncating a Pareto distribution from below leaves a Pareto distribution with the same exponent, so the shape survives the selection exactly while the location does not.
Fig. 2 The share of the population’s total size carried by the survivors at each scale, with the fitted tail exponent of the survivors beside it. Keeping one feature in a hundred still carries 99.93 per cent of the total, and the exponent stays where it was.

The shape survives, exactly. Truncating a Pareto distribution from below leaves a Pareto distribution with the same exponent — that is the defining property of the family — so a reader who fits an exponent to what is on the sheet gets the population’s own. That is a strong positive result and it is why the fractal-dimension literature can work off published maps at all.

The location does not survive, by two orders of magnitude. The median surviving feature is 95 times the median feature, and the smaller the scale the worse it is.

The bias is worst where it is least visible. The median surviving feature against the median of the whole population, as the scale gets smaller and the selection tighter. At a half the survivors' median is 4.0 times the truth; at one in a hundred it is 12795. A reader forming an impression of the typical feature from a small-scale sheet is looking at a population selected for being large, and the smaller the scale the more severely selected it is.
Fig. 3 The median survivor against the median of the whole population, as the selection tightens. At a half it is 4×; at one in a hundred it is 12,795×. The bias is worst at exactly the scales a reader is least able to check it.
share kept median inflation total carried
1 in 2 4.0× 100.00%
1 in 4 16.8× 100.00%
1 in 10 95.4× 99.99%
1 in 20 411× 99.99%
1 in 100 12,795× 99.93%

Read the last two columns against each other and the whole of this rung is in them: the total is intact at every scale in the table, and the median is wrong by four orders of magnitude at the bottom of it. Both statistics are computed from the same survivors.

Töpfer's square root is one line of a family. The fraction of features surviving to a smaller scale, for four stated populations whose size distributions differ only in their exponent. Every one is a straight line on these axes, and the slope of each is its own exponent: 0.3, 0.5, 0.8, 1.2. Töpfer's radical law is the line at 0.5 — the square root — and it is exact for that population and for no other. The law is not a rule of thumb with exceptions; it is a theorem with a hypothesis nobody states.
Fig. 4 The rung below’s own picture, for context: how many features survive to each scale, counted against Töpfer’s prediction. That curve is the count, and everything in this rung is about which features are underneath it.
The exponent the selection recovers is the population's own. The exponent fitted to the surviving counts, for four populations of stated exponent and one with none. Each of the four comes back within 0.024 of the value it was built with, at a coefficient of determination above 0.9997. The lognormal fits an exponent too — 1.67 — and fits it badly, at 0.9934, which is the only warning the arithmetic gives that the number means nothing.
Fig. 5 And the exponent fitted from the selection at four stated populations, which is the rung below’s finding: the square root in Töpfer’s law is the exponent of the size distribution, recovered to three decimals where there is one. That fitted exponent is the third of the three statistics measured here, and it is the one selection leaves alone.

Which questions the map can answer

The three statistics divide the questions a reader might ask into two piles, and the division is sharp.

Answerable off the sheet: how much is there in total; what the size distribution’s shape is; how the count scales between two sheets. All three are properties of the tail, and the tail is what survived.

Not answerable off the sheet: how large a typical feature is; how many small ones there are; what fraction of features are of any particular size. All three are properties of the part that was removed, and no amount of care in reading the map recovers them.

The uncomfortable part is that nothing distinguishes the two piles visually. A reader looking at a 1:1,000,000 sheet sees features of a range of sizes, forms an impression of what is typical, and that impression is a measurement of the threshold rather than of the ground — which is a thousand features wrong in the same direction with a different mechanism and the same shape: a systematic error that looks like data because every feature on the page is individually correct.

The one place the bias is visible

There is a single circumstance in which a reader can see the selection happening, and it is worth naming because it is the exception that shows the rule.

Comparing two sheets of the same ground at two scales. The larger-scale sheet’s population is the smaller-scale sheet’s plus everything between the two thresholds, so the difference between them is the removed band. A reader with both sheets can recover the size distribution over that band exactly, which is what makes a multi-scale series a better instrument than any one of its sheets.

That is not usually how a map is read. A reader has the sheet that is appropriate to the question, uses it alone, and has no access to the band it dropped. And the practice a tolerance is a promise about the picture describes — publishing the tolerance a sheet was drawn to — states the displacement of what survived and says nothing at all about what did not.

The missing declaration is the threshold. A sheet that stated the size below which nothing appears would let a reader convert every one of the unanswerable questions above into an answerable one, because the population is a truncated distribution and a truncated distribution with a known truncation point is a full one. It is a single number per sheet.

Why the alternatives are worse

The natural objection is that generalisation should use a better rule, and the middle column of the table is the answer.

A random sample is statistically unbiased and it is not a map. It drops the largest feature — the one landmark a reader navigates by — with probability nine in ten, and it keeps a scatter of features too small to draw, which is the constraint a tolerance is a promise about the picture starts from. It answers the questions the size-threshold rule cannot and fails the ones it can.

A stratified sample is the compromise and it inherits both problems in proportion: it still inflates the median threefold, still carries only 5 per cent of the total, and still drops most of the large features. It is what one would build to preserve the shape, and the shape was already surviving.

There is no rule that keeps the total and the median at once, because the total lives in the tail and the median lives in the body, and a fixed number of features cannot be drawn from both. That is not a limitation of any algorithm; it is a property of a distribution whose tail carries essentially all of its mass.

The same law on a population it was not a theorem about. A population whose sizes are lognormal rather than power-law — no exponent anywhere in it — put through the same selection. Töpfer's law over-predicts at every scale and the error grows without bound: 42 per cent at one halving and 100.0 per cent at 1:5.0 million, where it predicts 1057 features and 0 survive.
Fig. 6 The counts themselves at each scale, which the selection is drawn from. Every bar is a population, and the three statistics in the table above are computed on the population each bar represents rather than on the bar’s height.

What a practitioner would do with this

Two consequences, and they point in different directions.

For anybody computing off a published map, the division above is a checklist. A total, a density, a fitted exponent and a count ratio between two scales are all safe; a typical size, a size histogram, a proportion of features under some size and anything computed from the number of small features are not. The line between the two lists is whether the statistic is a functional of the tail, which is decidable before the number is computed.

For anybody making a map, the finding is a mild defence of the crude rule. A size threshold looks like the least intelligent possible selection and it preserves exactly the statistics a quantitative reader wants, because it is aligned with where the mass of the distribution is. A more thoughtful rule — one that keeps a representative spread, or that keeps features by importance rather than size — moves the map closer to what a reader expects and further from what a reader can measure.

Which of those two the map should serve is a purpose question rather than a measurement, and this collection’s standing rule — no essay may say a projection is best without naming the purpose — applies to a selection rule exactly as it does to a projection.

Where the model stops

One distribution. Everything above is measured on a Pareto population with α = ½. The rung below establishes that real feature populations are sometimes that and sometimes not — a lognormal has no exponent at all, and its apparent one drifts with the scale it is measured at. On a lognormal population the total is not concentrated in the tail, so the first of the three findings would go, and the direction of the argument would change with it.

Size is one number. A real feature has a length, an area, a width and an importance, and a legibility limit acts on the smallest of them. Treating size as a scalar is what makes the arithmetic clean, and it is why the numbers here are about a model population rather than about any country’s rivers.

And selection is not the only operation. A generalisation also simplifies what it keeps — which is how a line has a length only at a scale — and merges what is nearly touching, displaces what would collide and exaggerates what would vanish. This rung prices the first alone, holding the others fixed — which is possible here because the population is a list of sizes and not a set of shapes.

The generalisation

A threshold is a sampling rule, and a sampling rule decides which statistics survive.

The collection has met this with other objects. The ellipses are a sample, drawn at a size somebody chose is the same statement about an indicatrix field: the sampling scheme decides what a reader takes away, and it is stated nowhere. The answer depends on the cells it was counted in is the same statement about a partition rather than a threshold.

What is distinctive here is the split. Usually the finding is that a statistic is wrong; here two of three are exactly right, and which two is decidable in advance from the distribution’s own shape. A threshold at x preserves every statistic that is a functional of the tail above x and destroys every one that is not, and that is a checkable property of a statistic rather than a judgement about a map.

The check that the machinery can see the difference

The three rules are compared on one population with one keep fraction, and a comparison of three numbers is only worth having if the machinery could have returned three equal ones.

Two controls run alongside. The random sample must keep the median — it does, to seven per cent, which is sampling noise on a two-thousand-feature draw — and it must not keep the largest feature, which it does not; if it had, the comparison would be between two rules that happened to agree. And the largest-first rule must carry more than three times the total the random sample does, which is the statement that the tail is where the mass is rather than an artefact of the keep fraction.

Both are asserted rather than inspected, because the interesting failure here is a silent one: a bug that made all three rules keep the same features would produce a table of three identical columns and look like a strong null result.

Who found it, and when

Töpfer and Pillewizer’s radical law is from 1966 and is about counts. The literature on cartographic selection since then — Langran and Weber, Regnauld, the whole model-generalisation programme — is largely about doing better than a size threshold, using importance, topology and context, and it is motivated by exactly the observation that a threshold is not a selection.

What appears not to be stated in that literature is which statistics a size threshold nevertheless preserves. The answer turns out to be the ones a quantitative reader most often wants — the total and the distributional shape — which is a defence of the crude rule that the sophisticated alternatives do not obviously improve on.

A crude rule with an unexpected defence

The finding inverts the usual reading of the selection literature, and the inversion is worth stating plainly because it is easy to mistake for a defence of doing nothing.

The literature’s complaint about size thresholds is correct. A threshold discards small features regardless of importance, breaks topology, drops the only lake in a region and keeps three in another — and every one of those is a real defect that the model-generalisation programme exists to repair. Nothing measured here contradicts any of it.

But the complaint is about the map, and the finding is about the statistics. A reader looking at the sheet is served by the sophisticated methods; a reader computing a total, a mean or a distribution shape from the surviving features is served by the crude one, because a size threshold’s survivors preserve exactly those quantities and an importance-weighted selection need not.

Which means the two kinds of consumer want different generalisations of the same data, and a producer supplying one is not serving the other. That is not an argument against the sophisticated methods; it is an argument that a dataset generalised for legibility carries statistics that were never meant to survive, and nothing in the file says so.

And it explains why the defence is missing from the literature. The programme was motivated by what a map looks like, so it measured what a map looks like; nobody asked what a selection does to a total, because a total is not a cartographic quantity. The answer turns out to favour the method the programme was written to replace, on a criterion the programme was not addressing.

The practical form is a labelling problem rather than an algorithmic one. A producer shipping a generalised layer knows which method made it, and a consumer computing a total from it needs to know — but nothing in a vector format records how a selection was made, so the two kinds of layer are indistinguishable once written. That is the same omission this collection keeps meeting: the information exists at the moment of production, has nowhere to be recorded, and is exactly what a downstream user needs in order to know which questions their copy can answer.

A layer that has been generalised for legibility and a layer that has been thresholded by size look identical in a file and answer different questions.

The two are produced for different readers and there is no field in which to say which reader a given file was made for.

Nothing in a vector format has a place to record which of the two a file is, so the question cannot be asked of the data.

Where the ladder goes next

Seven rungs price what a generalisation does to a line, to a shared boundary, to a population and now to the statistics that population carries. The next question is about the chain: a national series is not generalised once but repeatedly, each scale derived from the one above it, and whether that is the same as generalising the source directly turns out to have a surprising answer and an exact condition attached.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

BiasClosed formExponentGeneralisationLegibilityQuadratic lawRepresentative fractionSamplingScaleSelectionSimplificationTest field