Concept

Generalisation — where it appears

Reducing the detail of a geometry so that it can be drawn or stored at a coarser scale. Its tolerance is a promise about the drawn displacement and about nothing else — not the enclosed area, not which side of the boundary a point falls, and not whether the curve still fails to cross itself.

Named by 18 essays across 4 fields — each of them below, with the objects they name alongside it.

Everywhere within 200 km of the route from London to Tokyo. The shortest route between the two places, and the set of places within 200 kilometres of it, with both edges computed on the sphere and then projected. On the ground the set holds 3.949 million square kilometres — the band 3.823 and the two end caps 0.126, which between them make one disc of the corridor's own width. Drawn in Mollweide the corridor is visibly wider at one end than the other, and the ground it stands for is not.

A corridor has a width the page cannot keep

Buffering a line is the most-run operation in spatial analysis and the corridor it produces is a reach set with two failures nobody separates: stroked on the page, one width covers 190 to 471 kilometres of ground along a single route; computed in closed form, the formula stops being the area at a width the route's own length fixes, and eventually claims more ground than the sphere has.

paths · Reach
Two rules, three populations. The taught rule keys on latitude; the replacement keys on shape. On the thirty regions the replacement was read from it scores 83 per cent against the taught rule's 63. On forty-five different regions built the same way it scores 91 — higher, not lower, so the generalisation the shortfall doubted is real. On the seven named regions this collection actually uses, both rules score 29 per cent, and following either costs a mean factor of 14.6. The third bar of each group puts the taught rule's polar clause back into the shape rule, which is what the out-of-sample failures ask for: it repairs four of them, breaks two that were right, and reaches 43 of 45.

The rule scored out of sample

A replacement rule was read off thirty regions and scored on the same thirty, and this collection recorded that as not being evidence about any other thirty. It is: on forty-five different regions the rule scores 91 per cent against the 83 it managed at home. What it cannot do is the seven regions the collection actually uses, where both it and the rule it replaced name the winner twice out of seven and cost a mean factor of 14.6.

wrong · Audit
How far each page reorders the shapes. The number of pairs of shapes whose order on the page differs from their order on the ground, out of 36, for ten projections. Four of them put a shape other than the geodesic disc at the top — Equirectangular, Lambert azimuthal equal-area, Robinson, Miller cylindrical — which means the shape that attains the isoperimetric bound on the sphere is not the most compact thing on those sheets. The projection with none is not the equal-area one; it is whichever one's stretching happens to leave this particular set of shapes alone.

The most compact shape depends on the paper

A geodesic disc attains the isoperimetric bound on a sphere — its score is one, exactly, at any radius. Score the same nine regions from their images on ten projections and four of them put something else on top, an oval and its own 45° rotation come out 4.7 per cent apart, and the projection that preserves the order is not the equal-area one.

paths · Reach
The score does not settle at any resolution. The compactness of one stated boundary — a circle with cosine ripples at eight geometrically spaced wavenumbers, so it has structure at every scale — read at sixteen vertices up to two thousand and forty-eight. The ground score falls from 0.980 to 0.834, and it keeps falling: the boundary's length grows without bound as it is resolved while the area it encloses converges, so the quotient has no limit. The four page curves sit within a fraction of a per cent of the ground curve and of each other, which is the comparison this rung exists to make.

The score is not stable at any scale

One boundary, read at eight resolutions from sixteen points to two thousand and forty-eight: the compactness score falls from 0.980 to 0.834 and is still falling. Changing the projection instead moves it by 0.69 per cent. The two decisions are made by the same person on the same afternoon and only one of them is ever reported.

paths · Reach
One unit of a vector tile, in metres of ground. A vector tile's coordinates are integers on a lattice 4096 units across the tile, and the tile halves at every level, so one unit is a distance that halves too: 5.48 m at z10 and 0.086 m at z16, at 55°. It is also a different distance at every latitude, by cos φ, because the tile is in Web Mercator — the same factor that makes a grid metre a different quantity of ground at every latitude, arriving in the file format rather than in the projection.

A vector tile has an integer grid

Six essays on this ladder treat a vector tile as the thing a raster tile is not: geometry, resolution-free, styled at draw time. Its coordinates are integers on a lattice 4,096 units across a tile, the tile halves at every level, and at 55° north one unit is 88 metres at zoom 6 and 21 millimetres at zoom 18.

applied · Screen
A curve built to have dimension 1.2619. The generator replaces every segment with four of equal length at headings 0, +60.0°, −60.0° and 0. Closing the displacement fixes the length ratio at 0.33335, and four copies at that ratio give a dimension of exactly log 4 / log(2 + 2 cos θ) = 1.2619. Nothing here is measured yet: this is the construction the measurement will be checked against. Drawn at depth 5, which is 1024 segments, with the second-level shape shown faint beneath it.

A line has a length only at a scale

Every measurement on this site so far has been of a curve given by a formula, sampled as finely as the picture needed. A map is not that: the geometry that reaches the page has been through an algorithm whose job is to throw most of it away. The first thing that goes is the idea that the line had a length.

applied · Generalise
The same tolerance, applied in two orders, at 65°. The faint line is the region's boundary as built, 1025 vertices across 400 km of ground. Both pipelines were given the same tolerance of 2000 m on the ground. Simplifying in degrees and then projecting keeps 311 of them; projecting into Mercator and then simplifying keeps 129; doing it on the ground itself, which no pipeline does, keeps 129. The two drawn lines separate by 1883 m, which is 94 per cent of the tolerance that was supposed to bound the whole operation.

Simplification does not commute with the projection

A pipeline either simplifies the geometry and then projects it, or projects it and then simplifies. Both orders are in use, neither is recorded, and given the same tolerance in ground metres they keep different vertices — 129 of them on the ground, 367 in degree space at 80°, and 459 on an equal-area page.

applied · Generalise
The picture is kept, at four tolerances. One closed curve of 3001 vertices, simplified at four tolerances. Douglas–Peucker's promise holds in every panel: no discarded vertex is further than ε from the line drawn in its place, measured at 0.1158 against 0.128 in the last. The picture survives. The enclosed area does not: it falls by 5.43 per cent, and it falls rather than wandering, because cutting a corner takes area off and never puts it back.

A tolerance is a promise about the picture

Douglas–Peucker guarantees exactly one thing: no vertex it discarded is further than ε from the line drawn in its place. It says nothing about the enclosed area, nothing about which side of the boundary a point ends up on, and nothing about whether the curve still fails to cross itself — and all three are what the geometry is usually being asked.

applied · Generalise
The whole of what an unlabelled map gives you. The outline of Japan as drawn on Conformal conic, delivered as an ordered list of page positions with nothing attached to any of them. No latitude, no longitude, no scale, no north. The rung's question is whether a projection can be recovered from that, and it can: the correspondence between the ink and the ground is found by sweeping the starting point round the curve and both directions, and the true candidate comes back with a residual of 2.88e-14 against the runner-up's 4.57e-4.

A map with no graticule

Ten rungs are handed control points, and a great many maps have none. Handed an outline with no labels on it at all, the method still works — and works better: the correspondence between ink and ground is recoverable exactly, because a similarity preserves ratios of arc length, and the margin on clean observations is 1.6 × 10¹⁰ against a graticule's 9.9 × 10⁶. What breaks it is noise, at three parts in a thousand.

wrong · Identify
What simplifying a boundary does to the number stored beside it. One region, simplified at five tolerances, with the error in the two quantities a consumer computes from the pair. If the density was stored, the total it implies moves by exactly the area's error — -1.55 per cent at the loosest tolerance. If the total was stored, the density it implies moves the other way by the same amount. Nothing in the file says which of the two was measured and which is being derived, and the simplification is normally done by a tool that never opens the attribute table.

The attribute is a claim about the geometry

Fourteen essays price what a stored coordinate means and not one asks what the number stored beside it means. A rate is a quantity divided by an area, the area belongs to the geometry, and no format records which area — so a simplification that moves the outline by nothing visible moves the implied total by 1.55 per cent, an unweighted average of densities is 4.09 per cent out, and a choropleth gives a polar square kilometre fifteen times the ink of an equatorial one.

applied · Dataset
The same 2-pixel road at three latitudes, zoom 5. The dark bar is the mark as drawn — 2 pixels, identical in all three panels, because that is what the stylesheet says. The pale band behind it is the ground that mark covers, drawn to one common ground scale: 9.78 kilometres at the equator, 6.92 at 45° and 1.70 at 80°. The reader sees the dark bar and is being told about the pale one.

The road is drawn two pixels wide

Seven rungs measure what a screen map does to position. Nothing on a map is a point: every mark has a width, the width is chosen in pixels, and a two-pixel road covers 9.78 kilometres of ground at the equator and 1.70 at 80° north. That is a generalisation applied at a strength varying by a factor of six across one sheet, by a stylesheet with no latitude in it.

applied · Screen
Two features, one shared boundary, simplified apart. Two neighbouring areas whose common boundary is a curve with structure at every scale — a river or a ridge, in effect — each stored with its own copy of that boundary and each simplified on its own at a tolerance of 0.01. The faint outlines are the originals and the solid ones what came back. The two copies of the shared boundary were within 0.01 of each other before the simplification and are not afterwards: 144 probe cells of 40000 now lie inside both features and 0 inside neither.

A boundary that two features share

Three rungs simplify one curve and price what a tolerance covers. Almost no boundary in a real dataset belongs to one feature: a county's edge is the next county's edge, it is stored twice, and it is simplified twice. What opens between the two answers is a region belonging to both features or to neither, and its area is not bounded by the tolerance.

applied · Generalise
Twenty-four versions of one shape, and not one of them gains area. The same closed boundary rotated twenty-four times and simplified at the same tolerance. If the area error were noise the values would straddle zero and their mean would fall towards it; they do not. Every one is negative, the mean is -0.4644 per cent, and the mean is 71 standard errors from zero. A bias of that size cannot be removed by averaging over more boundaries, which is the only defence anybody has against a rounding error.

A thousand features are wrong in the same direction

The area a simplification costs is unpredictable in sign for one feature. Over a population it is not: twenty-four presentations of one shape all lose area, the mean is seventy standard errors below zero, and no amount of aggregation removes it.

applied · Generalise
Töpfer's square root is one line of a family. The fraction of features surviving to a smaller scale, for four stated populations whose size distributions differ only in their exponent. Every one is a straight line on these axes, and the slope of each is its own exponent: 0.3, 0.5, 0.8, 1.2. Töpfer's radical law is the line at 0.5 — the square root — and it is exact for that population and for no other. The law is not a rule of thumb with exceptions; it is a theorem with a hypothesis nobody states.

How many features a scale can carry

Töpfer's radical law is quoted everywhere as a rule of thumb. It is not one: it is a theorem about a size distribution with a Pareto exponent of exactly one half, exact to 1.8 per cent for that population and out by 99.4 per cent for a lognormal one.

applied · Generalise
Three selection rules, and what each one keeps. Keeping one feature in ten from a stated population whose size distribution has a Pareto exponent of a half — the exponent Töpfer's law is a theorem about. Keeping the largest carries 99.99 per cent of the total size and inflates the median feature by a factor of 95. A random sample keeps the median to 1.068 and carries 5.0 per cent of the total. The two rules are right about different things and there is no rule that is right about both, because the total lives in the tail and the median does not.

Which features survive is not a sample

The rung below answers how many features a scale can carry and treats the population as a number. Which ones survive is a different question: keeping one feature in ten carries 99.99 per cent of the total length and inflates the median feature by a factor of 95, and the shape of the size distribution survives both exactly.

applied · Generalise
How far the two routes end up apart. The distance between a line simplified directly at the final tolerance and the same line simplified through two intermediate products, as a multiple of the final tolerance, with a stated extra step applied to each intermediate. With nothing in between the two are the same line to the last bit, because Douglas–Peucker's outputs are nested. Rounding the intermediate to half the tolerance, smoothing it for legibility, or running it through a moving average each break that, and the last of them puts the final product 0.142 away from where a direct route would have put it — nine times the tolerance the product is published under.

Two routes to one scale

A national series is cascaded — the million is derived from the quarter-million, which was derived from the fifty — and the folklore is that the errors accumulate. They do not: Douglas–Peucker and Visvalingam both cascade to the same line the direct route produces, bit for bit, because both output a sublevel set of a per-vertex number. What breaks it is anything else in the chain, and a moving average puts the product nine tolerances away.

applied · Generalise

A tolerance in map units is not a tolerance

A snapping tolerance is a number, and the number is in whatever units the file is in. Five map units on Web Mercator is 4.97 metres of ground at the equator and 0.87 at eighty degrees — so a rule that merges two features three metres apart merges them everywhere below 52.8° north and refuses everywhere above it, in one pass, over one dataset, with nothing recording where the boundary is.

applied · Dataset

The area is unbiased and the perimeter is not

A boundary measured from noisy vertices comes out long, always, by σ²/d on every leg. The area enclosed by the same vertices comes out exactly right, because a shoelace is bilinear and the cross terms vanish. So densifying a boundary makes its area five times more precise and its perimeter three thousand times more wrong, and every compactness score computed from it falls short.

distortion · Precision

Named alongside it

The objects these essays reach for when they reach for this one.

ToleranceSimplificationClosed formVerificationAggregationAreaPurposeResolutionScaleAnisotropyBiasBoundary

All concepts