A cartogram keeps the shapes it inflates
Seven rungs of this anchor build cartograms. They price the deformation, prove the target is met, find where the construction folds, show that the cheapest map is not the one anybody draws and that changing one number moves the whole page. Every one of them is about making the map.
A cartogram exists to be read, and reading one is a different operation with its own failure modes. A reader has to do two things: recognise which region is which, and turn a drawn area back into a value. This rung measures both, and the construction turns out to be working against both — not by accident of the algorithm, but for reasons that follow from what a cartogram is.
Measuring what a reader has to do
Recognising a region on a map is not a matter of its position, its size or which way round it has been turned. A cartogram is entitled to change all three — moving and resizing regions is what it is for — and a reader knows to allow for it.
What is left after position, size and rotation have been allowed for is the shape, and there is a standard measurement of how much of it survives: the residual after the best similarity transform, reported as a fraction of the shape’s own size. Zero means a region was moved and resized and otherwise left alone, so it is exactly as recognisable as it was. Approaching one means the residual is as large as the shape.
The choice of that measure is not arbitrary and it is worth a sentence, because two obvious alternatives are wrong for this question. Comparing outlines without removing size would report the growth twice, once as growth and once as shape change, and would guarantee the correlation the rung is trying to test. Comparing them without removing rotation would count a region that has been turned as unrecognisable, which is not how anybody reads a map. What is left after both are removed is the part that genuinely cannot be undone by a reader, and it is the only part worth calling a loss.
The measurement needs a control, and the control is exact. Under a uniform density request, the cartogram is the identity: every region’s growth comes out 1.000000 and every region’s shape loss comes out 0.000000. A construction that deformed anything under a flat request would be reporting its own solver rather than the data, and everything below would be about the solver.
The correlation, and its sign
Run it across the four cities’ request and the pattern is not subtle.
| region | grew by | shape lost |
|---|---|---|
| London | 8.28× | 0.16 |
| Delhi | 7.76× | 0.21 |
| Tokyo | 7.13× | 0.23 |
| São Paulo | 5.85× | 0.23 |
| Quito | 0.41× | 0.54 |
| Lagos | 0.38× | 0.55 |
The correlation between the logarithm of the growth and the shape loss is −0.996. Negative: the more a region is enlarged, the better its shape survives.
It holds on every request tested: −0.996 on four cities, −0.795 on two, −0.719 on the belt, −0.694 on one sharp city. A cartogram protects the identity of exactly the regions it is inflating, and destroys the identity of the ones it is deflating.
Reading down the columns rather than across the rows says something the correlation does not. Lagos loses 0.55, 0.58, 0.64 and 0.33 under the four non-uniform requests — it is being shrunk in all four, because it is in none of the concentrations. London loses 0.16, 0.17, 0.06 and 0.27, and is being enlarged in all four. So a region’s legibility on a cartogram is not a property of the map; it is a property of the region’s own value, which means the picture’s readability is correlated with the data it is showing. That is the sort of coupling a chart is normally designed to avoid.
Why it has to be that way
The mechanism is not in the algorithm and would not be changed by choosing a different one.
A region being enlarged sits near the centre of a concentration of the density it was drawn for. Near the centre of a bump the deformation is close to a pure dilation — it pushes outwards symmetrically, so a circle goes to a circle and only the radius changes. That is the same fact one number changed and the whole map moved records from the other side: the place whose number changed moves least, because it is at the bottom of a symmetric bowl.
The regions being shrunk are wherever is left. They are being squeezed by several bumps at once, from directions that do not agree, and the deformation there is strongly anisotropic — which is exactly what everything else on the page pays for the areas measures as angular deformation. Shape loss is the finite-region version of the same quantity.
So the two halves are one fact. Anisotropy and shrinkage are the same thing seen twice, because the total area is conserved: a region that grows is being pushed apart evenly and a region that shrinks is absorbing every direction’s worth of displacement at once.
Five of six are mistakable
Shape loss is a number; what matters is whether it is enough to break identification. It is.
Four or five of six regions are confused on every non-uniform request, and none is confused on the uniform one. The test is deliberately hard on the reader — shape alone, no position, no size, no labels — and that is the point of stating it that way: it isolates the channel. Real cartograms are labelled and their regions stay roughly where they were, so nobody is genuinely lost. What the measurement establishes is that the shape channel is carrying no information after the transformation, so everything a reader recovers is coming from the label and from memory of where things are.
That matters because it decides what a cartogram can be used for. A picture in which position is approximate, size is the data and shape is destroyed is a bar chart with a memorable arrangement — which is a perfectly good thing to be, and is not what a cartogram is usually claimed to be.
The second half: turning an area back into a value
The value a cartogram encodes is , up to one constant. Both areas are needed, and only one of them is on the page.
A reader who assumes all regions started the same size recovers values wrong by up to 82 per cent on these six, whose base areas differ by a factor of four. Real regions differ by far more than four: the ratio between the largest and smallest of a country’s administrative divisions is routinely a hundred.
The consequence is a statement about what the reader must already have. A cartogram is not a self-contained picture. It is a picture that has to be read against a map the reader is expected to be carrying in their head, and the accuracy of the reading is the accuracy of that memory. A choropleth is read by area records the complementary failure: there the base map is on the page and the reader integrates the wrong thing anyway.
What it does to a comparison
The two failures compound in a specific direction, and the direction is the awkward one.
Suppose a reader wants to check whether one small region really is smaller than another. Both are being shrunk, so both have lost most of their shape — 0.54 and 0.55 here — and both are candidates for confusion with something else. The value each carries has to be recovered by dividing by a base area the reader does not have, and the two base areas are exactly what the reader is least likely to remember, because small values on a cartogram are usually carried by regions that were not prominent on the base map either.
Now suppose the reader wants to check a large value. The region is recognisable, its shape is nearly intact, and its base area is likely to be one they know. The cartogram is easiest to read precisely where the answer is most obvious, and the reading gets harder in step with how much work the reader is doing.
That is not a fatal criticism and it is not an argument against cartograms. It is a statement about which claims a cartogram supports. It supports this one is much bigger than that one very well, since that is the comparison it is built to make legible and the one whose regions it protects. It supports these two are nearly equal at the small end hardly at all, and no amount of care in the construction changes that, because the construction is not what is causing it.
What was computed, and how
The regions are spherical caps of stated radii, from eight to sixteen degrees, so that their base areas differ by a factor of four and the recovery question is not vacuous. Their outlines are sampled at ninety-six points, and their areas are computed by the shoelace formula in the equal-area rectangle where longitude and the sine of latitude are the coordinates — so an area on that page is a ground area exactly, with no quadrature and no approximation.
The cartogram is the diffusion construction this anchor has used since its fourth rung. Its forward takes a latitude and returns the rectangle’s own coordinate, and mixing the two is a mistake this file has now made once: the displacement field was being sampled at , which never exceeds 48.2°, so that measurement was made on the middle half of the world. It is corrected, and the essay it belongs to is corrected with it — the place whose density changed is now last of seven rather than sixth.
Shape loss is Procrustes distance: both outlines are centred, scaled to unit norm, and the residual minimised over rotation in closed form. It needs no units and no threshold.
The assertions require five things separately: that every drawn area equal its base area times the density asked for; that shape loss fall as growth rises, with a correlation past −0.6; that a uniform request produce growth of exactly one and shape loss of exactly zero; that some region be confusable on a high-contrast request and none on a flat one; and that reading the value without the base areas be wrong by a stated, measurable amount.
The one thing that would fix it
There is a repair, it is not a better solver, and stating it is the useful half of the rung.
Draw the base area on the page. Every published cartogram has room for it: a faint outline of each region at its original size and shape, behind or beside the deformed one, turns the two unknowns a reader is missing into one thing they can see. The value is then read as a ratio of two areas both present on the paper, and the identification is made against the original outline rather than against memory.
That is not a novel suggestion — a map drawn to a density it was handed shows the pair side by side, and so do most careful presentations — but the reasoning for it is usually aesthetic or pedagogical. What the measurements here supply is the size of the thing it fixes: 82 per cent of value error at a factor-of-four spread in base areas, and five confusions in six regions. Those are not decoration.
The alternative repair, which is what non-contiguous constructions do, is to keep every region’s shape exactly and give up adjacency instead. That trades one channel for another and the trade has never been priced against this one. It is the next rung.
Where the model stops
These are discs, not countries. A real region’s outline has its own shape, and a shape with a long axis will lose more or less than a circle depending on how that axis sits against the deformation. Circles are the neutral choice — they have no orientation to be caught out by — so the measurement is a lower bound on what a real boundary would suffer.
Procrustes distance is one similarity measure. A reader does not compute residuals; they recognise. A turning-function distance, a Hausdorff distance or a human trial would each give a different number, and the claim here is about the sign and the ordering rather than about the units.
And the confusion test is harder than reading a real map. Position and label are both available in practice, and both carry more information than shape does. What the test isolates is what the shape channel is worth after the transformation, which is close to nothing — and that is a fact about the picture rather than about any reader.
The generalisation
The rule is that an encoding that uses one visual channel usually damages another, and the damage is not distributed evenly over the data.
A cartogram encodes value in area. Area cannot be changed without changing shape unless the change is a pure dilation, and a pure dilation is available only where the deformation is locally isotropic — which is exactly where the value is high. So the damage falls on the low values, systematically, and a reader trying to check a small number is working with the least legible part of the picture.
The same shape recurs across thematic cartography. A symbol has a size on the page and an area on the ground is the same trade with the symbol standing in for the region; the answer depends on the cells it was counted in is the same statement about aggregation. In each case the encoding is defensible and the unevenness of its cost is what nobody states.
The habit that follows is cheap. When a picture encodes a quantity in a geometric property, ask which other properties had to move, and then ask whether they moved most where the quantity is large or where it is small. If they moved most where it is small, the picture is best exactly where it is least needed.
This collection has been making the same move about projections since measuring instead of naming: a property is not a label, it is a measurement, and the measurement has to be made where the reader will use it rather than averaged over the sheet. What survives a change of coordinates settles which quantities are worth measuring at all. Shape loss under a similarity fit is such a quantity — it does not depend on where the region ended up, how large it became or which way round it is drawn, which is precisely why it is the right thing to ask a reader’s question with.
Who found it, and when
Cartograms have been drawn since the 1860s and the literature about them is almost entirely about construction: how to meet the areas, how fast, with what continuity, with what preservation of adjacency. Tobler’s diffusion method of 2004 is the standard, and the papers that followed it argue about the trade between area accuracy and shape preservation as a global quantity — one number for the map.
That global framing is what hides this. A cartogram’s mean shape distortion is a real number and it is the number everybody reports; splitting it by which regions bore it is one line of extra arithmetic that nobody had cause to write, because the question being asked was always whether the construction was good rather than what the reader could recover.
The one part of the subject that has looked at reading rather than making is the perceptual work on how badly people estimate areas at all, which finds that judged area grows as roughly the 0.7 power of true area — so the encoded quantity is systematically compressed before any of this begins. That is a second, independent problem with reading a cartogram, and the two compound: the values are compressed, and the regions carrying the small ones are the unrecognisable ones.
Where the ladder goes next
Eight rungs have treated a cartogram as a continuous deformation of the sphere, which is what makes the areas exact and the shapes pay. The constructions people actually publish often are not continuous — they cut regions apart, or replace them with circles, and give up adjacency deliberately in exchange for keeping something else. What each of those gives up, and what it buys, is a comparison this anchor has not made.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- The orientation is a policy area · design · purpose · shape distortion · verification
- A current drawn on a page has sources anisotropy · equal-area · thematic mapping · verification
- A partition under a directed cost has two versions anisotropy · area · purpose · verification
- Every density can be met and none is free cartogram · density · equal-area · shape distortion
- The most compact shape depends on the paper anisotropy · equal-area · purpose · verification
- A centroid belongs to a plane equal-area · thematic mapping · verification
The objects this essay names
Each one links to every other essay that touches it.
AnisotropyAreaCartogramDensityDesignEqual-areaIdentificationInverse problemLegibilityPurposeShape distortionThematic mappingVerification