Ladder

Generalise — the ladder

8 distinct arguments against one idea, from the one that introduces it to the one that assumes the rest.
  1. A curve built to have dimension 1.2619. The generator replaces every segment with four of equal length at headings 0, +60.0°, −60.0° and 0. Closing the displacement fixes the length ratio at 0.33335, and four copies at that ratio give a dimension of exactly log 4 / log(2 + 2 cos θ) = 1.2619. Nothing here is measured yet: this is the construction the measurement will be checked against. Drawn at depth 5, which is 1024 segments, with the second-level shape shown faint beneath it.

    A line has a length only at a scale

    Every measurement on this site so far has been of a curve given by a formula, sampled as finely as the picture needed. A map is not that: the geometry that reaches the page has been through an algorithm whose job is to throw most of it away. The first thing that goes is the idea that the line had a length.

    rung 1 · applied
  2. The same tolerance, applied in two orders, at 65°. The faint line is the region's boundary as built, 1025 vertices across 400 km of ground. Both pipelines were given the same tolerance of 2000 m on the ground. Simplifying in degrees and then projecting keeps 311 of them; projecting into Mercator and then simplifying keeps 129; doing it on the ground itself, which no pipeline does, keeps 129. The two drawn lines separate by 1883 m, which is 94 per cent of the tolerance that was supposed to bound the whole operation.

    Simplification does not commute with the projection

    A pipeline either simplifies the geometry and then projects it, or projects it and then simplifies. Both orders are in use, neither is recorded, and given the same tolerance in ground metres they keep different vertices — 129 of them on the ground, 367 in degree space at 80°, and 459 on an equal-area page.

    rung 2 · applied
  3. The picture is kept, at four tolerances. One closed curve of 3001 vertices, simplified at four tolerances. Douglas–Peucker's promise holds in every panel: no discarded vertex is further than ε from the line drawn in its place, measured at 0.1158 against 0.128 in the last. The picture survives. The enclosed area does not: it falls by 5.43 per cent, and it falls rather than wandering, because cutting a corner takes area off and never puts it back.

    A tolerance is a promise about the picture

    Douglas–Peucker guarantees exactly one thing: no vertex it discarded is further than ε from the line drawn in its place. It says nothing about the enclosed area, nothing about which side of the boundary a point ends up on, and nothing about whether the curve still fails to cross itself — and all three are what the geometry is usually being asked.

    rung 3 · applied
  4. Two features, one shared boundary, simplified apart. Two neighbouring areas whose common boundary is a curve with structure at every scale — a river or a ridge, in effect — each stored with its own copy of that boundary and each simplified on its own at a tolerance of 0.01. The faint outlines are the originals and the solid ones what came back. The two copies of the shared boundary were within 0.01 of each other before the simplification and are not afterwards: 144 probe cells of 40000 now lie inside both features and 0 inside neither.

    A boundary that two features share

    Three rungs simplify one curve and price what a tolerance covers. Almost no boundary in a real dataset belongs to one feature: a county's edge is the next county's edge, it is stored twice, and it is simplified twice. What opens between the two answers is a region belonging to both features or to neither, and its area is not bounded by the tolerance.

    rung 4 · applied
  5. Twenty-four versions of one shape, and not one of them gains area. The same closed boundary rotated twenty-four times and simplified at the same tolerance. If the area error were noise the values would straddle zero and their mean would fall towards it; they do not. Every one is negative, the mean is -0.4644 per cent, and the mean is 71 standard errors from zero. A bias of that size cannot be removed by averaging over more boundaries, which is the only defence anybody has against a rounding error.

    A thousand features are wrong in the same direction

    The area a simplification costs is unpredictable in sign for one feature. Over a population it is not: twenty-four presentations of one shape all lose area, the mean is seventy standard errors below zero, and no amount of aggregation removes it.

    rung 5 · applied
  6. Töpfer's square root is one line of a family. The fraction of features surviving to a smaller scale, for four stated populations whose size distributions differ only in their exponent. Every one is a straight line on these axes, and the slope of each is its own exponent: 0.3, 0.5, 0.8, 1.2. Töpfer's radical law is the line at 0.5 — the square root — and it is exact for that population and for no other. The law is not a rule of thumb with exceptions; it is a theorem with a hypothesis nobody states.

    How many features a scale can carry

    Töpfer's radical law is quoted everywhere as a rule of thumb. It is not one: it is a theorem about a size distribution with a Pareto exponent of exactly one half, exact to 1.8 per cent for that population and out by 99.4 per cent for a lognormal one.

    rung 6 · applied
  7. Three selection rules, and what each one keeps. Keeping one feature in ten from a stated population whose size distribution has a Pareto exponent of a half — the exponent Töpfer's law is a theorem about. Keeping the largest carries 99.99 per cent of the total size and inflates the median feature by a factor of 95. A random sample keeps the median to 1.068 and carries 5.0 per cent of the total. The two rules are right about different things and there is no rule that is right about both, because the total lives in the tail and the median does not.

    Which features survive is not a sample

    The rung below answers how many features a scale can carry and treats the population as a number. Which ones survive is a different question: keeping one feature in ten carries 99.99 per cent of the total length and inflates the median feature by a factor of 95, and the shape of the size distribution survives both exactly.

    rung 7 · applied
  8. How far the two routes end up apart. The distance between a line simplified directly at the final tolerance and the same line simplified through two intermediate products, as a multiple of the final tolerance, with a stated extra step applied to each intermediate. With nothing in between the two are the same line to the last bit, because Douglas–Peucker's outputs are nested. Rounding the intermediate to half the tolerance, smoothing it for legibility, or running it through a moving average each break that, and the last of them puts the final product 0.142 away from where a direct route would have put it — nine times the tolerance the product is published under.

    Two routes to one scale

    A national series is cascaded — the million is derived from the quarter-million, which was derived from the fifty — and the folklore is that the errors accumulate. They do not: Douglas–Peucker and Visvalingam both cascade to the same line the direct route produces, bit for bit, because both output a sublevel set of a per-vertex number. What breaks it is anything else in the chain, and a moving average puts the product nine tolerances away.

    rung 8 · applied

All ladders