Paths and directions

The score is not stable at any scale

One boundary, read at eight resolutions from sixteen points to two thousand and forty-eight: the compactness score falls from 0.980 to 0.834 and is still falling. Changing the projection instead moves it by 0.69 per cent. The two decisions are made by the same person on the same afternoon and only one of them is ever reported.

Assumes The most compact shape depends on the paper.

The previous rung scored nine shapes on ten projections and found that the page decides which of them is most compact. Every one of those shapes was smooth, and the boundaries were read at 720 points because that was plenty.

A real boundary is not smooth, and a line has a length only at a scale is the whole of the reason: a coastline’s measured length grows without bound as the ruler shrinks, so a perimeter is not a property of the feature but of the feature together with the resolution it was captured at.

A compactness score has that perimeter in its denominator, squared.

The score does not settle at any resolution. The compactness of one stated boundary — a circle with cosine ripples at eight geometrically spaced wavenumbers, so it has structure at every scale — read at sixteen vertices up to two thousand and forty-eight. The ground score falls from 0.980 to 0.834, and it keeps falling: the boundary's length grows without bound as it is resolved while the area it encloses converges, so the quotient has no limit. The four page curves sit within a fraction of a per cent of the ground curve and of each other, which is the comparison this rung exists to make.
Fig. 1 One boundary read at eight resolutions. The ground score falls from 0.980 at thirty-two vertices to 0.834 at two thousand and forty-eight, and keeps falling. The four page curves sit within a fraction of a per cent of it and of each other.

The boundary, written down

The shape is a stated closed form, not a coastline, for the reason this site’s coastline decision gives: a boundary read from a dataset has somebody’s generalisation in it, and the whole question here is what a generalisation does, so importing one would beg it.

So the boundary is a circle with cosine ripples at eight geometrically spaced wavenumbers, each with an amplitude falling as a stated power of its wavenumber:

r(θ)=R(1+k=18a3kHcos ⁣(3kθ+ϕk)),r(\theta) = R\left(1 + \sum_{k=1}^{8}\frac{a}{3^{kH}}\cos\!\left(3^{k}\theta + \phi_k\right)\right),

with a=0.09a = 0.09 and H=0.75H = 0.75. That is a radial Weierstrass series, its box dimension is 2H=1.252 - H = 1.25, and it looks like a coastline for exactly the reason coastlines look like this — structure at every scale with a stated relation between amplitude and wavelength. Every number in this essay comes out of that formula and nothing comes out of a file.

One boundary, read three ways. The stated boundary — a circle with cosine ripples at eight geometrically spaced wavenumbers — read at sixteen, a hundred and twenty-eight and two thousand and forty-eight vertices, drawn as Lambert azimuthal equal-area draws it. The enclosed ground is the same to seven figures in all three. The way round grows by a fifth, and the compactness score falls from 0.978 to 0.834, because each refinement admits an octave of ripples that simply were not there before.
Fig. 2 The boundary read at sixteen, a hundred and twenty-eight and two thousand and forty-eight vertices. The ground it encloses is the same to seven figures in all three; the way round grows by a fifth; the compactness score falls from 0.978 to 0.834.

What the ladder does

vertices perimeter ground score on Mercator
16 8,311 km 0.978 0.977
32 8,382 0.980 0.980
64 8,448 0.970 0.969
128 8,507 0.958 0.957
256 8,590 0.939 0.939
512 8,749 0.906 0.906
1,024 8,882 0.879 0.879
2,048 9,120 0.834 0.833

The perimeter grows by ten per cent over seven doublings and shows no sign of stopping, which is what a boundary of dimension 1.25 does. The area converges almost immediately — the ripples at high wavenumber contribute area proportional to their amplitude squared, which is negligible by the third term.

One of these converges. The enclosed area and the boundary length of the same stated boundary, each divided by its own value, against the number of vertices it is read at. The area is settled to seven figures by a thousand vertices — 5578362 km² — and does not move again. The perimeter grows from 8311 km to 10638, a rise of 28 per cent, and is still climbing at thirty-two thousand. The compactness score is the first divided by the square of the second, so it inherits the one that never stops.
Fig. 3 The enclosed area and the boundary length of the same boundary, each relative to where it starts, against the resolution it is read at. The area is settled to seven figures by a thousand vertices and does not move again. The perimeter rises by 28 per cent over the ladder and is still climbing at thirty-two thousand.

That figure is the control the score cannot supply on its own. If both quantities grew, the fall would be a statement about the sampler; if neither did, there would be nothing to report. The area converges to 5,578,362 km² and stays there through five further doublings, while the perimeter goes from 8,311 km to 10,638 and shows no sign of a limit.

So the numerator settles and the denominator does not, and the quotient inherits the denominator’s behaviour: it falls, without limit, and there is no resolution at which it is right.

The first two rows are the exception and they are worth explaining rather than hiding. At sixteen and thirty-two vertices the polygon is too coarse to resolve even the first ripple, so it is reading a slightly deformed circle and scoring it accordingly. The monotone fall begins once the resolution catches the largest wavelength, which is where the ladder starts saying anything.

The comparison this rung exists to make

Which of the two decisions moves the score more. The spread the score takes on, as a percentage of itself, under the two decisions a producer makes without thinking of either as a decision. Changing the resolution the boundary is read at moves it by 15.7 per cent. Changing the projection, at a fixed resolution, moves it by 0.69. The ratio is 23 to one, so a published compactness score that states its projection and not its capture scale has named the smaller of its two problems.
Fig. 4 The spread the score takes under the two decisions, as a percentage of itself. Reading the boundary at sixteen points rather than two thousand moves it by 15.7 per cent. Changing the projection at fixed resolution moves it by 0.69.

Twenty-three to one.

That number is the rung, and the direction of it is the uncomfortable part. The previous rung spent an essay establishing that a compactness score computed from projected geometry is a property of the projection, which is true and is worth knowing. This one says the same score is a property of the capture resolution by a factor of twenty-three more, and the capture resolution is never stated at all.

A published score names its formula. It sometimes names its projection. It essentially never names the scale of the boundary data, because the boundary data arrives as a file and a file does not look like a measurement — it looks like the shape itself.

What the score is being asked to do

It is worth being clear about why a producer wants one number here at all, because it explains why the failure survives.

A reach set, a district, a catchment and a habitat patch are all objects with complicated boundaries and a question attached: is this shape a normal one, or is something odd about it? The isoperimetric quotient is the cheapest possible answer — one number, bounded, with a known extremal case, computable from geometry that is already in hand. How many features a scale can carry prices the same instinct from the cartographic side: a single number describing a whole map’s content is enormously convenient and is a statement about the capture as much as about the ground.

The convenience is real. What it buys is a number whose two ingredients behave completely differently under the one operation everybody performs on boundary data, which is resampling it.

Why the two effects are so unequal

The asymmetry is structural rather than accidental, and it comes from the two quantities having different orders of contact with the thing being varied.

A projection is a smooth map. Over a region of twenty degrees at mid-latitude it stretches by a bounded factor with bounded variation, so both the area and the perimeter change by a bounded fraction, and because the quotient is a ratio much of that cancels. The residual is what the previous rung measured: a fraction of a per cent for a smooth boundary, larger for an elongated one, and always bounded.

A resolution is not a map at all. It changes which of the boundary’s features exist. Halving the ruler admits a whole new octave of ripples that were simply absent before, each contributing its own length and almost no area. There is nothing to cancel, because the numerator does not change.

So one effect is a distortion and the other is an omission, and this collection’s recurring lesson is that omissions are the larger and the quieter of the two. Which features survive is not a sample of the ones that do not makes the same observation about a population of features rather than about one boundary.

What was computed, and how

The boundary is evaluated at 4,096 points once, and each resolution takes every 4096/n4096/n-th of them, so the coarse polygons are genuinely subsets of the fine one rather than separately generated curves. That matters: it makes the ladder a statement about reading one boundary at different resolutions rather than about eight different boundaries.

The ground area is the spherical excess summed edge by edge and the ground perimeter is a sum of great-circle distances, both as in the previous rung. The page quantities are a shoelace and a Euclidean perimeter over the projected vertices.

The assertion has two parts and both could fail. The score must fall monotonically from the third rung of the ladder onward — a boundary that was not multi-scale would settle, and then there would be nothing to report — and the resolution spread must exceed the projection spread by at least a factor of two. If the second failed the honest conclusion would be that the projection is the larger problem and this essay would say so. It does not fail; it clears the bar by a factor of eleven.

The same boundary on four sheets

The projection effect is small here, and saying so precisely is what earns the comparison.

The score does not settle at any resolution. The compactness of one stated boundary — a circle with cosine ripples at eight geometrically spaced wavenumbers, so it has structure at every scale — read at sixteen vertices up to two thousand and forty-eight. The ground score falls from 0.980 to 0.834, and it keeps falling: the boundary's length grows without bound as it is resolved while the area it encloses converges, so the quotient has no limit. The four page curves sit within a fraction of a per cent of the ground curve and of each other, which is the comparison this rung exists to make.
Fig. 5 The same ladder with the four page curves drawn against the ground curve. They track it so closely that at this scale the five are nearly one line — the spread between the best and worst projection at any fixed resolution is 0.69 per cent, against a 15.7 per cent fall down the ladder.

The reason the projection matters so little on this shape is the previous rung’s mechanism running in reverse. A projection’s effect on the quotient comes from anisotropy — stretching a shape along one axis — and it is largest for an elongated shape whose long axis lines up with the stretching. This boundary is a perturbed circle: it has no long axis, so the stretching has nothing to grip, and what is left is the second-order residual.

That is a genuine limitation on the comparison and it is stated rather than buried. On the elongated shapes of the previous rung the projection moves the score by up to 21.8 per cent, which is the same order as the resolution effect here. The honest statement is that both effects are large and only one of them is ever reported, and on a compact rough boundary the unreported one is twenty-three times the reported one.

What a producer can do

The repairs are unequal in cost and in effect, which is worth stating in order.

Report the capture scale beside the score. “Polsby–Popper 0.42, from 1:250,000 boundaries” is a complete statement and “Polsby–Popper 0.42” is not. This costs nothing, it is the only repair available to somebody comparing published figures, and it is not standard practice anywhere.

Compare only within one scale. Two districts scored from the same national dataset at the same generalisation level are comparable, because the effect is common to both and much of it cancels in the comparison. Two districts from different sources are not, and the difference between a state’s own boundary file and a national one is easily a factor this size.

Resample to a stated resolution before scoring. Densify or simplify every boundary to a common vertex spacing — a stated number of metres — and the scores become comparable across sources. This is a few lines and it makes the score a well-defined function of the geometry and the stated spacing, which is the best a quantity with no limit can be given.

And compute on the ellipsoid, which is the previous rung’s repair and removes the smaller of the two problems. It is still worth doing; it is simply not the one to do first.

The order matters because it is the reverse of the attention the two get. The projection question is discussed in the literature; the scale question is not discussed at all, and it is twenty-three times larger.

Where the model stops

The series is truncated, so it does have a limit. Eight terms means the smallest ripple has wavenumber 3⁸ = 6,561, and a polygon fine enough to resolve it — somewhere past thirteen thousand vertices — would start to see the perimeter converge. A real coastline has no such term; this one does, and the claim being made is about the range of resolutions anybody actually captures data at, which is entirely inside the range where it behaves like a boundary with no limit. At thirty-two thousand vertices the perimeter is still rising.

One boundary, one exponent. The ratio between the two effects depends on how rough the boundary is. A smooth boundary — an administrative unit defined by straight lines between monuments, say — has no ripples to admit, its perimeter converges, and the projection becomes the dominant effect again. So the factor of twenty-three is a property of a boundary with H=0.75H = 0.75, and the general claim is about which effect is unbounded rather than about the number.

Simplification is not the same as subsampling. Taking every kk-th vertex is the crudest possible generalisation and is not what a cartographer does; a tolerance is a promise about the picture describes the algorithms that are actually used, and they preserve a boundary’s extremes rather than its parametrisation. Subsampling is used here because it is exactly reversible and states the resolution as a number, and a real simplification would give a similar ladder with different constants. A boundary that two features share is where the difference between the two starts to matter for its own reasons.

The projection effect measured here is small because the shape is round. The previous rung’s elongated shapes move by up to 21.8 per cent on the plate carrée, so the factor of twenty-three is not a general exchange rate between the two effects; it is the ratio on a boundary with no preferred direction.

And nothing here says the score is meaningless. A fractal boundary’s compactness at a stated resolution is a perfectly well-defined number, computed from a well-defined procedure, comparable with any other number computed the same way. What it is not is a property of the region, and the whole practice of publishing it as one is what this rung objects to.

The generalisation

Two rungs have taken one number apart and the pair of them makes a rule.

A quantity that has no limit cannot be a property, and the ways it can fail to have one are not equally visible. A perimeter has no limit under refinement of the boundary; that is Richardson’s coastline observation and it is famous. A compactness score inherits it, invisibly, because the score is dimensionless and looks like the kind of number that would be scale-free. Dimensionlessness is not scale-invariance, and confusing the two is what lets a ratio of two lengths carry a resolution dependence that neither of its parts hides.

The general test is one question: does the quantity converge as the description of the object is refined? An area does. A length does not. A ratio does whatever its worst part does, and a ratio’s appearance of being a pure number is not evidence either way.

Who found it, and when

Richardson measured the coastline effect in the 1950s while looking for a relation between the length of a border and the probability of a war between the countries either side of it, and found instead that the length was not a number. Mandelbrot named it in 1967. Neither of them was talking about compactness.

The compactness literature knows about it, in the sense that papers occasionally note that scores should be computed at a consistent scale. What is missing is the size: the effect has not, as far as this collection can find, been measured against the projection effect on the same boundary, and the two are discussed as though they were comparable. They are not, and the one that gets the attention is the smaller.

Which quantities pass the convergence test

The test at the end of the argument — does the quantity converge as the description of the object is refined — sorts the things this collection computes about a region, and the sorting is worth having in one place because the answer is not obvious from what the quantities look like.

These converge. An area, because it is a measure and refining the boundary adds and removes slivers that vanish. A centroid, for the same reason. A diameter, since it is decided by two extreme points that stabilise. A bounding box. An inscribed or circumscribed circle. Anything that is an integral over the interior, or a distance between features of the interior.

These do not. A perimeter, which grows without limit on a real boundary. A vertex count, which is a property of the file. A count of features above a size threshold, which changes as the threshold’s meaning changes with the representation. Anything that integrates along the boundary rather than over the region.

A ratio inherits the worse half. Compactness scores built from area and perimeter do not converge, because the perimeter does not, and no amount of care about the area repairs it. Scores built from two areas do converge. The appearance of being a dimensionless number says nothing either way — dimensionlessness is about units and convergence is about limits.

Which gives a usable rule for reading any published statistic about a region: find the boundary in it. If the boundary enters only as the edge of an integration, the number is a property of the region. If it enters as something to be traversed and totalled, the number is a property of the drawing, and comparing two of them compares two capture resolutions.

And the rule is cheap to apply from the definition alone, without computing anything — which is the point of stating it. The convergence question is settled by what the formula does, not by how the numbers behave on any particular dataset.

Where this anchor goes next

Seven rungs have measured a set’s position, area, width, membership, asymmetry, shape and now the stability of its shape. Every one of them has treated the set’s boundary as known — sharp, stated, and correct. What a reach set actually has is a boundary that is a contour of an estimate: travel times come from a model, the model has uncertainty, and the set within three hours is really the set the model believes is within three hours. That uncertainty has a shape of its own, and it is the one thing this anchor’s machinery has never carried.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

BoundaryClosed formConvergenceEstimatorFractal dimensionGeneralisationPurposeReachRegional distortionSamplingToleranceVerification