Concept

Sampling — where it appears

Recording a continuous field at a discrete set of places, which fixes the finest structure any later calculation can see. The interval is a choice with two opposite costs — a coarse grid loses the field and a fine one amplifies its noise — so there is a best one and it is not the finest.

Named by 14 essays across 5 fields — each of them below, with the objects they name alongside it.

The score does not settle at any resolution. The compactness of one stated boundary — a circle with cosine ripples at eight geometrically spaced wavenumbers, so it has structure at every scale — read at sixteen vertices up to two thousand and forty-eight. The ground score falls from 0.980 to 0.834, and it keeps falling: the boundary's length grows without bound as it is resolved while the area it encloses converges, so the quotient has no limit. The four page curves sit within a fraction of a per cent of the ground curve and of each other, which is the comparison this rung exists to make.

The score is not stable at any scale

One boundary, read at eight resolutions from sixteen points to two thousand and forty-eight: the compactness score falls from 0.980 to 0.834 and is still falling. Changing the projection instead moves it by 0.69 per cent. The two decisions are made by the same person on the same afternoon and only one of them is ever reported.

paths · Reach
A finer grid makes a measured slope worse. The error in the direction of steepest ascent, against the spacing the field was sampled at, for four noise levels. With exact values the curve falls at a fitted slope of 2.00 — second order, which is what a central difference is. Add noise and the same curve turns over: a finite difference divides the noise by the spacing, so halving the grid doubles the noise in the slope while quartering an error that was already negligible. The minimum is where the two meet, and it is not at the fine end.

The slope of a field that was measured

Three rungs differentiate a formula, which is what makes the projection the only thing under test. A real field is a grid of numbers with an error on each of them, and differencing such a thing divides the noise by the spacing — so a finer grid gives a worse slope, there is a best spacing, and it is the cube root of the noise.

distortion · Gradient
One of these settles. The largest departure of the fitted map's boundary scale from constant — the quantity the previous rung showed the solver cannot see — against the number of collocation nodes, for the two placements. The clustered fit reaches 7.030e-5 at forty-eight nodes and returns exactly that at every count above it. The evenly spaced fit does not settle at all: it wanders by a factor of 1.43 across the same range, going up as often as down. Refining an evenly collocated fit is not convergence, and the previous rung's finding that more samples improve the report and not the map is this seen from one side.

The nodes were evenly spaced

The previous rung showed that refining an evenly collocated fit improves the solver's report and leaves the map alone. Moving the same number of nodes to the Chebyshev positions — crowded toward the corners, where a conformal map of a polygon is singular — makes the fit settle: 7.030 × 10⁻⁵ at forty-eight nodes and exactly that at every count above it, against an even fit that wanders by 43 per cent and never converges at all.

choosing · Condition
What a seven-parameter fit's residual is made of. A published transformation accuracy is the root-mean-square residual at the common points, and here it is 1.51 metres. Almost all of it — 1.51 — is the network's own distortion, which a rigid motion and a scale cannot follow and which is present at every point of the country whether it was used in the fit or not. The transformation's own error, measured as the disagreement between the fitted parameters and the true ones over a clean grid, is 0.091 metres: 6 per cent of the quoted figure. The last bar is the control — the same fit with the distortion switched off, at 2.9e-4 metres.

What another common point buys

Rung three finds that a seven-parameter datum fit leaves a pattern rather than noise. Six per cent of the residual it reports is the transformation's own error and the other ninety-four is distortion no seven parameters can follow — so adding common points improves a term that was already small and cannot touch the one that is quoted.

datums · Datum
A 900 km circular accuracy at 55° north, projected. Six thousand ground positions drawn from a circular error of 900 kilometres about one place, each projected in Mercator and plotted as a displacement from the projected place. The curve is the nominal 95 per cent ellipse, computed the standard way — the ground covariance sandwiched between the projection's own derivatives. It holds 93.83 per cent of the points, the cloud is measurably longer than it along its own long axis by 5.52 per cent, and it is not symmetric: the third moment along the page's second axis is 0.727 rather than zero.

The error ellipse is not an ellipse

Rung two pushed a covariance through a projection with the same matrix sandwich that draws an indicatrix. That is a first-order operation on a map with a second derivative, so the propagated distribution is not the ellipse the sandwich draws — and a nominal 95 per cent ellipse holds 93.06 per cent on one projection and 95.63 on another, in opposite directions, from the same input.

distortion · Precision
Three selection rules, and what each one keeps. Keeping one feature in ten from a stated population whose size distribution has a Pareto exponent of a half — the exponent Töpfer's law is a theorem about. Keeping the largest carries 99.99 per cent of the total size and inflates the median feature by a factor of 95. A random sample keeps the median to 1.068 and carries 5.0 per cent of the total. The two rules are right about different things and there is no rule that is right about both, because the total lives in the tail and the median does not.

Which features survive is not a sample

The rung below answers how many features a scale can carry and treats the population as a number. Which ones survive is a different question: keeping one feature in ten carries 99.99 per cent of the total length and inflates the median feature by a factor of 95, and the shape of the size distribution survives both exactly.

applied · Generalise
Two conventions for one indicatrix field on Mercator. The left panel is the convention nearly every published field uses: every ellipse drawn at the same area on the paper, so the picture carries the shape and throws the size away. The right panel draws the image of the same ground circle at every place, so the drawn area IS the areal scale factor. On Mercator the axis ratio spans 1.0000 and the areal factor spans 14.93, so the left panel shows nothing the other one shows. Neither caption states which is which, on any published map this collection has found.

The ellipses are a sample, drawn at a size somebody chose

Thirteen essays measure with the indicatrix and none audits it as an instrument. A published field has a gauge nobody states and a placement nobody states: on Mercator the standard convention draws twenty-five identical circles while the areal factor runs over a factor of 14.9, and the average a reader takes off any of these fields is between 22 and 64 per cent too high.

distortion · Tissot
The same count, scattered two ways. 22 dots in every cell of a 30-cell covering, drawn on Mercator. The upper panel places them uniformly on the GROUND — uniform in longitude and in the sine of latitude, which is what uniform on a sphere means — and the lower places them uniformly on the PAGE, which is what a drawing routine handed a polygon does. Both panels carry exactly the same number of dots in exactly the same regions, so both are honest as totals. They are different pictures, and a reader reads a dot map by density.

A dot map's density is partly the projection's

A dot map carries the right number of dots in every region whichever way it is drawn, so it is honest as a total under both placements. It cannot be honest as a density under both: ground on a uniform field reads 0.099 of its equatorial density at 72° north on Mercator, and scattering inside the polygon on the page moves 64.3 per cent of a cell's dots into its northern half without one of them leaving the cell.

distortion · Thematic
The areal factor over 0° to 60° north, and the points it is measured at. Mercator's areal factor shaded over 0° to 60° north, from 1.0001 to 3.8473, with the 12 × 12 grid of cell centres a regional measurement uses drawn on top of it. The cross marks where the quantity is actually largest, found by a search that is allowed to leave the sample; the ring marks the largest value the sample contains. The grid's answer is 3.4639 and the real one is 4.0000, short by 13.40% — and the reason is visible in the picture, because no cell centre is ever on an edge.

The worst point is not on the grid

Every maximum distortion this collection has printed is a maximum over a sample, and a maximum over a sample is a lower bound. On Mercator over a sixty-degree band a twelve-by-twelve grid reports 3.464 where the answer is exactly 4, and the shortfall does not go away with refinement so much as decay at a rate that says where the extreme is hiding.

distortion · Sampling
One of these averages exists. Mercator's area-weighted mean areal factor, and its Kavrayskiy number, against how close the sampled band comes to the pole. The first is artanh(sin Φ)/sin Φ in closed form and has no limit: it passes 7.04 at a tenth of a degree from the pole and keeps going, gaining a fixed amount every time the remaining gap is halved. The second settles by 85° and does not move again. Both are published as summary distortion figures for the same map.

A mean that does not exist can still be printed

Mercator's area-weighted mean areal factor is artanh(sin Φ)/sin Φ, and it has no limit. A sampler asked for it returns the logarithm of its own sample count plus 1.512 — measured slope 1.001 against ln n — so the number is a property of the person who computed it. The Kavrayskiy number for the same map over the same sphere settles at 0.52124 and is a number.

distortion · Sampling
Three ways to cover a sphere with 900 points. The three samplers this collection's numbers are computed with, each with about 900 points, drawn on the Mollweide projection so that equal areas on the sphere are equal areas on the page and the crowding is the samplers' rather than the map's. The graticule piles points at the poles; the equal-area rings space them evenly by area and unevenly by distance; the Fibonacci lattice trades a little of each. None of them is equal-area, and no finite set is.

There is no equal-area lattice on a sphere

Every number in this collection is an average over a point set, and the three point sets available all fail to be equal-area in different ways. The equal-area ring sampler this site has used since its early essays is the one whose outermost ring sits half a step inside the rim — which is how a measured scale spread once came in 3.42 parts in a thousand below a proved bound.

distortion · Sampling
The doubling ladder is the one sequence that cannot see it. The largest departure of a small circle from its own indicatrix that a grid of n latitudes over 10° to 70° north finds on the Robinson projection. The filled marks are 4, 8, 16, 32, 64 and 128 — the doubling ladder every convergence study runs — and they rise smoothly to about 4.66e-4 with the increments halving, which is what a convergent first-order sequence looks like. The open marks are grids whose samples land on the projection's five-degree table entries. They report 1.28e-2, twenty-seven times higher, and whether a grid does that is decided by whether n is a multiple of four.

A refinement that stops moving

Doubling the sample and watching the answer settle is how every quadrature in every field is checked. On the Robinson projection the doubling ladder — 4, 8, 16, 32, 64, 128 — converges beautifully, with its increments halving at every step, on a limit that is wrong by a factor of twenty-seven. Whether a grid finds the answer is decided by whether n is a multiple of four.

distortion · Sampling
Which of this collection's numbers are the sampler's. How much each of two published quantities moves when the sampler behind it goes from twenty samples a side to sixty, for eight projections over the whole sphere. The Kavrayskiy numbers — the summary means the rankings are built from — move by at most 1.15 per cent. The worst-point angular deformations move by up to 19.0 per cent, all in the same direction, because they are maxima over a sample and a maximum over a sample is a lower bound. The split is clean, and it says which numbers here need refining and which do not.

Which of these numbers are the sampler's

Four rungs have shown that a sampled maximum understates, a sampled mean can be a report on the sampler, no arrangement of points is neutral, and refining until the answer settles proves nothing. So the collection re-measured itself. The means move by at most 1.1 per cent, the worst points by up to 19, and the rankings — which is what the essays actually argue with — do not move at all.

distortion · Sampling
One map, two samples, two answers. The same Miller cylindrical projection carrying two sets of sample points of the same size. On the left the points are uniform over the SPHERE — equal ground area between them, which is what every mean in this collection integrates against. On the right they are uniform over the PAGE, which is what a raster, a pixel loop or any figure that walks its own canvas produces. The right-hand set crowds where the map stretches, and the mean angular deformation it returns is 18.0° against the left-hand set's 7.2°.

The sample was drawn on the page

Every mean in this collection integrates over the sphere, because that is where the ground is. A raster, a pixel loop and any figure that walks its own canvas integrate over the page instead, and the difference is exactly the covariance between the quantity being measured and the map's own area distortion — 7.2° of mean angular deformation on Miller becoming 18.0°.

distortion · Sampling

Named alongside it

The objects these essays reach for when they reach for this one.

VerificationClosed formEstimatorAreal factorEqual-areaConvergenceQuadratureToleranceAngular deformationMercatorRegional distortionBias

All concepts