What a machine does with it

A centroid belongs to a plane

Every renderer labels a region at its centroid, and every centroid is a shoelace over coordinates as stored — which is a statement about the plane they are in. Six planes put the middle of one 20° × 20° region up to 273 kilometres apart, the equal-area member is 66 kilometres out, and the disagreement falls as the square of the region's size.

Every map with labels on it has answered the question where is the middle of this region, usually without being asked. A renderer puts a label at a polygon’s centroid, and the centroid is a shoelace over the coordinates as they are stored.

A shoelace is a plane formula. Applied to coordinates in one plane it returns one point; applied to the same region’s coordinates in another plane it returns another point; and the two are not the same place on the ground.

Where the middle of a 20° × 20° region is, in five planes. The region is drawn in longitude and latitude — which is itself a projection, and one of the ones being compared. Each filled mark is the shoelace centroid computed in one projected plane and inverted back to the ground; the hollow mark is the centre of area on the sphere, by integration. They spread over 273 kilometres. The equal-area member is 66 kilometres out, because a centroid is a first moment and preserving area says nothing about where the area sits.
Fig. 1 The middle of one 20° × 20° region, computed in five projected planes and inverted back to a longitude and a latitude, against the centre of area on the sphere obtained by integration. The five answers spread over 273 kilometres, and the hollow mark belongs to no plane at all.

The five answers

plane centroid latitude distance from the surface centroid
Mercator 61.568° 272.6 km
equirectangular 60.000° 98.2 km
Gall–Peters 58.525° 65.8 km
Mollweide 58.818° 33.2 km
Lambert azimuthal 58.965° 16.8 km
the sphere 59.117°

Mercator pulls the answer north because it stretches the high latitudes and therefore gives them more weight in the average; the plate carrée returns exactly 60°, the arithmetic mean of the latitude bounds, which is what a shoelace on latitude and longitude is; and the equal-area members pull it slightly south of the truth.

That last row is the surprising one and it is the point of the essay.

The spread is worth reading as a whole before the mechanism. Five projections, all of them in daily use, all of them applied to the same twenty degrees of the world, and the answers they give for its middle occupy a north–south interval three degrees long. A reader shown any one of the five numbers on its own would have no way to tell it from the others, because each is the output of an exact calculation on exact inputs — and the calculation each performed is a different calculation wearing the same name.

Preserving area does not preserve the centre of area

An equal-area projection multiplies every region’s area by one. A centroid is not an area; it is a first moment divided by an area, and the moment involves where the area is, not merely how much of it there is.

Gall–Peters preserves the amount of area in every part of the region and rearranges where that area sits — a cell at 65° north is the same size as one at 55° but is drawn at a different distance from the reference — and the first moment changes accordingly. The result is 65.8 kilometres from the surface centroid, which is worse than Mollweide’s and four times worse than Lambert’s.

The projection that gets a centroid nearest right is not the one that preserves area but the one that is least distorted over the region, which here is the azimuthal equal-area projection centred nearby. That is a different criterion, and it is the criterion choosing for a line, not a region and its neighbours are about.

The law, and it is the one this site keeps meeting

The disagreement falls quadratically with the region’s size:

region span worst disagreement
0.67 km
2.66 km
16.7 km
10° 66.9 km
20° 272.6 km
40° 1,191 km

Fitted exponent 2.02, which is the second-order term of two maps that agree to first order at a point. Relative to the region’s own size the error is therefore first order: a region half the size has half the relative error, not a quarter.

That is the identical law the indicatrix is a limit finds for the departure of a finite circle from Tissot’s ellipse, and for the same reason. Two descriptions that agree to first order differ at second order, and dividing by anything proportional to the size loses one power.

The practical form: a county-sized polygon has a centroid wrong by a few kilometres and a continent-sized one by hundreds. Small regions are safe and the safety is only linear in how small they are.

The area of one 20° × 20° cell at 50–70° north, seven ways. The cell has an exact area — R²Δλ(sin φ₂ − sin φ₁), 2,460,334 square kilometres — so every other row is a measurement of the method rather than of the ground. The equal-area projection returns it to 1.000000 and the spherical polygon formula to 1.000000, which is three routes agreeing — and the same cell integrated on the ELLIPSOID comes out 0.53 per cent away from all three, because the sphere is a model. Taking the shoelace in Mercator gives 4.17 times too much, and treating degrees as a length gives 2.01 times — about sec φ at the cell's middle, which is where that error comes from.
Fig. 2 The zeroth moment for comparison: the same region’s area computed in the same planes. Here the equal-area member is exactly right and the others are wildly wrong, which is the opposite ordering to the centroid table above. Preserving the zeroth moment and preserving the first are different requirements, and only the first has a projection named after it.

The same region, seen in the planes that disagree

Putting the five planes side by side says where the disagreement comes from without any arithmetic: each one draws the region a different shape, and the middle of a different shape is in a different place.

The region in question spans 50° to 70° north, which is where these planes disagree most about how tall a band of latitude should be drawn: Mercator stretches it, the plate carree leaves it alone, Gall-Peters compresses it. The centroid follows the stretching, because a shoelace formula weights by drawn area and by nothing else.

Mercator’s northward pull is the clearest case and it has a closed form. Its northing derivative is secφ\sec\varphi, so the band from 60° to 70° is drawn 1.37 times as tall as the band from 50° to 60° while covering less ground — the shoelace weights the northern half more, and the centroid moves up. The 272-kilometre displacement is that ratio integrated.

Five identical cells on Mercator. Five patches, each 20° of longitude by 10° of latitude. On the sphere the higher ones are genuinely smaller, because the meridians converge. On Mercator the cell at 70° comes out 15.4 times larger than the equatorial one relative to its true size.
Fig. 3 The same weighting shown as cells: equal-angle cells drawn at wildly different sizes, each labelled with its inflation. A centroid computed in this plane is an average weighted by these numbers rather than by the ground, which is the whole mechanism in one picture.

What the true answer even is

There is a prior question the table above skates over: which point is the “right” middle of a region on a sphere?

The definition used here is the centre of area — integrate the position vector over the region with the area element, normalise, and project the result out to the surface. It is what a physicist would call the centre of mass of a uniform lamina, radially projected, and it has two properties that make it the right reference: it is invariant under rotations of the sphere, and it reduces to the planar centroid for a small region.

It is not inside the region in general, and neither is a planar centroid. A crescent-shaped country’s centroid is in the sea; a ring-shaped region’s is in the hole. That is a defect of the centroid as a label position rather than of any particular way of computing it, and no choice of plane fixes it.

So the essay’s claim is narrow and worth stating narrowly: given that a centroid has been chosen as the answer, which plane it is computed in moves it by hundreds of kilometres, and that is a separate failure from the centroid being the wrong kind of answer.

The right answer has a closed form, for the shape the table is about

The reference above is computed by integration on a grid, refined until it stops moving, which is the honest way to get a number that has to be trusted. For the particular region this essay measures — a cell bounded by two meridians and two parallels — the integral is elementary and the answer can be written down, which is worth doing because it removes the last soft number from the comparison.

The centre of area is the mean position vector over the region. Its vertical component integrates sinφ\sin\varphi against the area element cosφ\cos\varphi, so it is proportional to the difference of sin2φ\sin^2\varphi at the two bounds; its horizontal magnitude integrates cosφ\cos\varphi against the same element, giving a half-angle plus a quarter of the difference of sin2φ\sin 2\varphi, times the chord factor 2sin(Δλ/2)2\sin(\Delta\lambda/2) that the longitude sweep contributes. The latitude of the centroid is the arctangent of the first over the second.

Evaluated on the 20° × 20° cell that the table is about, that expression gives 59.1168°, against the 59.117° the integration reports. The two agree to every digit the table prints, which is the check the reference needed: a numerical integral compared against nothing is a number, and a numerical integral that reproduces a closed form is a measurement.

The closed form also says what is wrong with the two intuitive answers in one line each. The arithmetic mean of the latitude bounds is 60°, and it is what the plate carrée returns because a shoelace on longitude and latitude is exactly that mean; it ignores the cosφ\cos\varphi in the area element entirely, so it over-weights the northern half of the cell by the ratio of the two bounding cosines. The other guess — the latitude whose sine is the mean of the two bounding sines, which is the latitude that halves the cell’s area — comes out at 58.525°, and it is not the centroid either, because halving the area is a statement about the zeroth moment and the centroid is the first. Those two guesses bracket the truth, one 0.88° north of it and the other 0.59° south, and neither is the answer.

No plane’s shoelace returns the closed form, and it is worth being clear about why, since it is the sharpest form of the essay’s claim. A shoelace in any plane returns the centroid of the drawn figure. For that to be the centre of area of the region on the ground, the map would have to carry every region’s first moment as well as every region’s zeroth — and a map that preserved the centroid of every sub-region would preserve the centroid of every infinitesimal one, which forces the areal factor to be constant, and then forces the map to be affine over the whole region. An affine map of a curved surface onto a plane does not exist, which is the impossibility this collection was founded on, arriving as an argument about labels.

That is the precise sense in which there is no equal-centroid projection. It is not that nobody has constructed one; it is that the condition collapses into a stronger one that Gauss’s theorem forbids, and an equal-area projection satisfies the first half of it and not the second.

The three moments, ranked differently by each plane

Reading the site’s three area-adjacent measurements together gives a small table nobody publishes, and it is the useful summary of the whole field.

quantity what preserves it worst plane here
area (zeroth moment) an equal-area projection, exactly Mercator, 3.06×
centroid (first moment) a locally undistorted projection Mercator, 273 km
shape (second moment) a conformal projection, at a point Gall–Peters, 38.9°

No projection is best in more than one row, which is the trade-off the site was founded on arriving in the form a data operation meets it. And the middle row is the one with no named projection attached: there is no “equal-centroid” projection, because preserving first moments over every region is a stronger condition than any map satisfies.

four planes a dataset might be stored in, scored on three operations. Each candidate measured over -10° to 30° east and 35° to 60° north: the worst areal error, the worst angular deformation, and the spread of the scale factor, which are what an area query, a shape and a distance respectively depend on. The best plane for areas is Gall–Peters, for shapes Web Mercator, and for distances Lambert azimuthal equal-area — three different answers, and no fourth candidate would collapse them, because a projection exact in two of these columns has a = b = 1 everywhere and is the isometry Gauss's theorem forbids. area of a polygon costs 3.06× too large in the wrong plane; drawing a line between two points costs 194 km from the ground it claims.
Fig. 4 The field’s own summary table: five candidate planes scored on the three things a spatial operation depends on. No candidate is exact in two columns and none ever will be, which is the operation decides the coordinate system. The centroid is a fourth column this essay adds, and it ranks the candidates in an order none of the other three produces.

One more consequence follows from the same collapse, and it is the reason the error in the table is not a constant that could be tabulated and subtracted. Because the condition that fails is a condition on the areal factor, the displacement a plane introduces depends on how that factor varies across the particular region — so it is a function of where the region is and how big it is, and two regions of identical shape at different latitudes get different corrections. There is nothing to publish.

Where this shows up

Labels. A renderer placing a country’s name at the centroid of its polygon places it in whatever plane the renderer works in. Two renderers of the same data, one in Web Mercator and one in an equal-area projection, put the same label in visibly different places.

Joins by nearest centroid. Assigning points to regions, or regions to regions, by centroid proximity inherits the whole error, and nearest is a question about the metric measures what that does to an identity rather than to a position.

Published “centre of” figures. The geographic centre of a country is a published number in several places and is a different number in each, partly for this reason and partly because the definitions differ. The disagreements are of the order this table predicts.

Anything weighted. A population-weighted centroid is the same integral with a density in it, and the plane’s distortion multiplies the weights, so the error is the projection’s areal factor correlated with the population distribution — which can be larger or smaller than the unweighted case and is not predictable without the data.

Anything decomposed before it is labelled. Clipping the same region to the tiles it crosses and marking each piece’s centroid moves the label by up to 1,863 kilometres — a tile is drawn without its neighbours — which is a far larger error than this essay’s and arises from cutting rather than from projecting. The two compound, and only one of them is visible in the output.

The control that makes it a measurement

Three, and the second is what stops the whole thing being a measurement of the inverse projection’s own error.

Each centroid is inverted back to the ground. The comparison is between places, not between page coordinates, so a projection that draws the region twice as large does not thereby get a centroid twice as far out.

The disagreement must vanish for a small region, quadratically. If the numbers stayed large as the region shrank, the measurement would be reporting the numerical inverse’s convergence rather than the effect.

The surface centroid is computed by integration, on a grid fine enough that halving the step does not move it, and it is not one of the projected answers. Using one of them as the reference would make the essay’s central number a comparison of two projections rather than of a projection with the ground.

What to do instead

The operation that is correct at any size is the same integral the reference uses: sum the position vectors over the region with the area element on the ellipsoid, normalise, and project out. It costs an integration rather than a shoelace and it belongs to no plane.

For a polygon rather than a cell there is a closed form — the surface integral over a spherical polygon reduces to a sum over its edges — and it is the same construction as the spherical area formula this site already uses for computing an area needs a surface. So the exact answer is available for the same cost as the exact area, which is a small multiple of the shoelace’s.

The reason it is not used is the reason nothing on this site is used: the shoelace is already written, it returns a plausible number, and nothing anywhere reports that the number is 273 kilometres from the place it names.

For a region too small for any of this to matter, the shoelace in a locally sensible plane is fine and the arithmetic above says how small that is: at a stated tolerance τ\tau kilometres, the region may span roughly 20τ/27320\sqrt{\tau/273} degrees. A ten-kilometre tolerance buys 3.8°, and a hundred-metre one buys 0.38° — which is the shape of every tolerance-to-size inversion on this site, and the same square root how small is flat enough gets for a plane.

What a data format could say and does not

The whole difficulty is that a file of coordinates does not record which plane its geometry is meant to be interpreted in for the purposes of an operation. It records a coordinate reference system, which says what the numbers mean as positions, and that is a different declaration.

A polygon stored in longitude and latitude is a set of positions on the ellipsoid. Whether its interior is the set of points enclosed by great-circle edges, by straight-in-plate-carrée edges, or by anything else is not stated, and neither is the plane its centroid should be computed in. Standards leave both to the reader.

That is one more entry on the list what a grid is made of started for coordinates and a coordinate without its system is not a location continued: the declarations a file carries are about the numbers, and the declarations an operation needs are about the geometry.

A smaller case, at the size real polygons come in

The 20° region is chosen because it makes the effect visible. Most polygons in a real dataset are administrative units a degree or two across, so the useful question is what the numbers are there.

At 2° the worst disagreement over this pair of planes is 1.9 kilometres, and over the wider set in the table above it is 2.66; at 1° it is 0.67. Those are small on a map and are not small for every purpose: a label two or three kilometres from where it should be is inside the wrong district in a dense country, and a centroid-based join at that scale assigns points to the wrong unit near every boundary.

The two-degree figure also has an uncomfortable property. It is a per cent or so of the width of many of the units it would be computed for, so the error is a fraction of the polygon rather than a fraction of its position — which is the regime where a reader looking at the label would notice something is wrong and would be unable to say what.

Where the middle of a 2° × 2° region is, in four planes. The region is drawn in longitude and latitude — which is itself a projection, and one of the ones being compared. Each filled mark is the shoelace centroid computed in one projected plane and inverted back to the ground; the hollow mark is the centre of area on the sphere, by integration. They spread over 2 kilometres. The equal-area member is 1 kilometres out, because a centroid is a first moment and preserving area says nothing about where the area sits.
Fig. 5 The same measurement on a 2° × 2° region — the size of a large county. The planes now agree to within 1.9 kilometres, which is a hundredth of the 20° case and is exactly what a quadratic law predicts for a tenth of the size.

Where this ladder goes

The dataset ladder has established that a coordinate without its system is not a location, that a degree is not a unit of length, that computing an area needs a surface, that a straight segment is a claim about a plane, that the antimeridian is a cut in the numbers, and that nearest is a question about the metric.

This adds the first moment to the zeroth: an area computed in the wrong plane is wrong by a factor, and a centroid computed in the wrong plane is wrong by a distance, and the two failures rank the projections in opposite orders.

What follows is the operation that asks about the boundary rather than about the interior — whether a point is inside — where the answer depends not on the plane the area was measured in but on the plane the edges were understood to be straight in, and where the disagreement is a band rather than a point.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 10 that link here.

The objects this essay names

Each one links to every other essay that touches it.

Areal scaleCentroidCoordinate reference systemEqual-areaLocalityMetricNumerical integrationQuadratic lawShoelaceThematic mappingToleranceVerification