Measuring distortion

A cartogram keeps the shapes it inflates

Seven rungs build cartograms and none reads one back. Reading means recognising a region and dividing its drawn area by its base area, and the construction is against the reader twice: the correlation between how much a region grows and how much of its shape it keeps is −0.996, and the base area a reader has to divide by is the map the cartogram replaced.

Assumes One number changed and the whole map moved.

Seven rungs of this anchor build cartograms. They price the deformation, prove the target is met, find where the construction folds, show that the cheapest map is not the one anybody draws and that changing one number moves the whole page. Every one of them is about making the map.

A cartogram exists to be read, and reading one is a different operation with its own failure modes. A reader has to do two things: recognise which region is which, and turn a drawn area back into a value. This rung measures both, and the construction turns out to be working against both — not by accident of the algorithm, but for reasons that follow from what a cartogram is.

Six regions, before and after. Six circular regions of unequal size on the equal-area rectangle, and the same six after the cartogram of four cities has been solved. Each drawn area is exactly its base area times the density it was asked for. The four that grow stay recognisably round; the two that shrink are drawn out into shapes that no longer resemble what they were, and nothing in the construction chose to treat them differently.
Fig. 1 Six circular regions of deliberately unequal size on the equal-area rectangle, and the same six after the cartogram of four cities has been solved. Each drawn area is exactly its base area times the density it was asked for. The four being enlarged stay recognisably round; the two being shrunk are drawn out into shapes that no longer resemble what they were. Nothing in the construction chose to treat them differently.

Measuring what a reader has to do

Recognising a region on a map is not a matter of its position, its size or which way round it has been turned. A cartogram is entitled to change all three — moving and resizing regions is what it is for — and a reader knows to allow for it.

What is left after position, size and rotation have been allowed for is the shape, and there is a standard measurement of how much of it survives: the residual after the best similarity transform, reported as a fraction of the shape’s own size. Zero means a region was moved and resized and otherwise left alone, so it is exactly as recognisable as it was. Approaching one means the residual is as large as the shape.

What is left after size, position and turn are removed. Each region's original outline, shaded, with its drawn outline over it after the best similarity fit — so the difference shown is the part a cartogram is not entitled to change and a reader cannot undo. The number is the residual as a fraction of the shape's own size. The two regions being shrunk are drawn into crescents; the four being enlarged remain circles.
Fig. 2 Each region’s original outline, shaded, with its drawn outline laid over it after the best similarity fit — so what is shown is the part a cartogram is not entitled to change and a reader cannot undo. The numbers are the residuals. The four regions being enlarged remain circles to within a fifth of their own size; the two being shrunk are drawn into crescents.

The choice of that measure is not arbitrary and it is worth a sentence, because two obvious alternatives are wrong for this question. Comparing outlines without removing size would report the growth twice, once as growth and once as shape change, and would guarantee the correlation the rung is trying to test. Comparing them without removing rotation would count a region that has been turned as unrecognisable, which is not how anybody reads a map. What is left after both are removed is the part that genuinely cannot be undone by a reader, and it is the only part worth calling a loss.

The measurement needs a control, and the control is exact. Under a uniform density request, the cartogram is the identity: every region’s growth comes out 1.000000 and every region’s shape loss comes out 0.000000. A construction that deformed anything under a flat request would be reporting its own solver rather than the data, and everything below would be about the solver.

The correlation, and its sign

Run it across the four cities’ request and the pattern is not subtle.

region grew by shape lost
London 8.28× 0.16
Delhi 7.76× 0.21
Tokyo 7.13× 0.23
São Paulo 5.85× 0.23
Quito 0.41× 0.54
Lagos 0.38× 0.55

The correlation between the logarithm of the growth and the shape loss is −0.996. Negative: the more a region is enlarged, the better its shape survives.

The regions it shrinks are the ones it makes unrecognisable. Every region under every density, plotted by how much it grew against how much of its shape it lost. The correlation is -0.465 — negative. A cartogram protects the identity of exactly the regions it is inflating and destroys the identity of the ones it is deflating, which is the opposite of what a reader trying to check a large value would want. The uniform density sits at (1, 0) and is the control: no growth, no loss, exactly.
Fig. 3 Every region under every density request, plotted by how much it grew against how much of its shape it lost. The uniform request sits at exactly (1, 0) and is the control. Within each request the correlation runs from −0.69 to −0.996; pooled across requests it is −0.47, weaker only because each request has its own scale of growth. Every one of them is negative.

It holds on every request tested: −0.996 on four cities, −0.795 on two, −0.719 on the belt, −0.694 on one sharp city. A cartogram protects the identity of exactly the regions it is inflating, and destroys the identity of the ones it is deflating.

Shape lost, region by region, request by request. How much of its identity each region loses under each density request. The uniform row is exactly zero everywhere and is the control: a construction that deformed anything under a flat request would be reporting its own solver rather than the data. Every other row has the same shape — the two regions being shrunk lose most, whichever request is being met.
Fig. 4 Shape lost, region by region, under each of five requests. The uniform row is exactly zero across the board and is the control. Every other row has the same profile whatever is being asked for: Lagos and Quito lose most under four cities, Lagos and Delhi under two, Lagos under one sharp city — the region losing most is always one being shrunk, and which regions those are changes with the request while the pattern does not.

Reading down the columns rather than across the rows says something the correlation does not. Lagos loses 0.55, 0.58, 0.64 and 0.33 under the four non-uniform requests — it is being shrunk in all four, because it is in none of the concentrations. London loses 0.16, 0.17, 0.06 and 0.27, and is being enlarged in all four. So a region’s legibility on a cartogram is not a property of the map; it is a property of the region’s own value, which means the picture’s readability is correlated with the data it is showing. That is the sort of coupling a chart is normally designed to avoid.

Why it has to be that way

The mechanism is not in the algorithm and would not be changed by choosing a different one.

A region being enlarged sits near the centre of a concentration of the density it was drawn for. Near the centre of a bump the deformation is close to a pure dilation — it pushes outwards symmetrically, so a circle goes to a circle and only the radius changes. That is the same fact one number changed and the whole map moved records from the other side: the place whose number changed moves least, because it is at the bottom of a symmetric bowl.

The regions being shrunk are wherever is left. They are being squeezed by several bumps at once, from directions that do not agree, and the deformation there is strongly anisotropic — which is exactly what everything else on the page pays for the areas measures as angular deformation. Shape loss is the finite-region version of the same quantity.

So the two halves are one fact. Anisotropy and shrinkage are the same thing seen twice, because the total area is conserved: a region that grows is being pushed apart evenly and a region that shrinks is absorbing every direction’s worth of displacement at once.

Five of six are mistakable

Shape loss is a number; what matters is whether it is enough to break identification. It is.

Which region a reader matching by shape would call it. Every drawn outline against every original outline, by the same residual-after-best-fit measure, with the closest match in each row ringed. On the cartogram of four cities, 5 of the 6 regions look more like some other region's original than like their own. A reader with the cartogram and the base map, and no labels, would misidentify them.
Fig. 5 Every drawn outline compared with every original outline by the same residual measure, with each row’s closest match ringed. On the cartogram of four cities, five of the six regions look more like some other region’s original than like their own — Tokyo’s drawn shape resembles Lagos’s original, and São Paulo’s resembles Tokyo’s. A reader holding the cartogram and the base map, with the labels removed, would get five of six wrong.

Four or five of six regions are confused on every non-uniform request, and none is confused on the uniform one. The test is deliberately hard on the reader — shape alone, no position, no size, no labels — and that is the point of stating it that way: it isolates the channel. Real cartograms are labelled and their regions stay roughly where they were, so nobody is genuinely lost. What the measurement establishes is that the shape channel is carrying no information after the transformation, so everything a reader recovers is coming from the label and from memory of where things are.

That matters because it decides what a cartogram can be used for. A picture in which position is approximate, size is the data and shape is destroyed is a bar chart with a memorable arrangement — which is a perfectly good thing to be, and is not what a cartogram is usually claimed to be.

The second half: turning an area back into a value

The value a cartogram encodes is V=Adrawn/AbaseV = A_{\text{drawn}} / A_{\text{base}}, up to one constant. Both areas are needed, and only one of them is on the page.

Reading the value off the page without the base map. A cartogram encodes a value as an area, so recovering the value means dividing the drawn area by the base area — and the base map is the thing the cartogram replaced. A reader who assumes the regions were all the same size to begin with is wrong by up to 82 per cent here, on regions whose true areas differ by a factor of four. The error is not noise: it is the base map, read back as data.
Fig. 6 The error in the recovered value for a reader who assumes every region started the same size. On regions whose true base areas differ by a factor of four — which is a mild spread by the standards of real administrative units — the recovered value is wrong by up to 82 per cent. The error is not noise; it is the base map, read back as data.

A reader who assumes all regions started the same size recovers values wrong by up to 82 per cent on these six, whose base areas differ by a factor of four. Real regions differ by far more than four: the ratio between the largest and smallest of a country’s administrative divisions is routinely a hundred.

The consequence is a statement about what the reader must already have. A cartogram is not a self-contained picture. It is a picture that has to be read against a map the reader is expected to be carrying in their head, and the accuracy of the reading is the accuracy of that memory. A choropleth is read by area records the complementary failure: there the base map is on the page and the reader integrates the wrong thing anyway.

What it does to a comparison

The two failures compound in a specific direction, and the direction is the awkward one.

Suppose a reader wants to check whether one small region really is smaller than another. Both are being shrunk, so both have lost most of their shape — 0.54 and 0.55 here — and both are candidates for confusion with something else. The value each carries has to be recovered by dividing by a base area the reader does not have, and the two base areas are exactly what the reader is least likely to remember, because small values on a cartogram are usually carried by regions that were not prominent on the base map either.

Now suppose the reader wants to check a large value. The region is recognisable, its shape is nearly intact, and its base area is likely to be one they know. The cartogram is easiest to read precisely where the answer is most obvious, and the reading gets harder in step with how much work the reader is doing.

That is not a fatal criticism and it is not an argument against cartograms. It is a statement about which claims a cartogram supports. It supports this one is much bigger than that one very well, since that is the comparison it is built to make legible and the one whose regions it protects. It supports these two are nearly equal at the small end hardly at all, and no amount of care in the construction changes that, because the construction is not what is causing it.

What was computed, and how

The regions are spherical caps of stated radii, from eight to sixteen degrees, so that their base areas differ by a factor of four and the recovery question is not vacuous. Their outlines are sampled at ninety-six points, and their areas are computed by the shoelace formula in the equal-area rectangle where longitude and the sine of latitude are the coordinates — so an area on that page is a ground area exactly, with no quadrature and no approximation.

The cartogram is the diffusion construction this anchor has used since its fourth rung. Its forward takes a latitude and returns the rectangle’s own coordinate, and mixing the two is a mistake this file has now made once: the displacement field was being sampled at arcsin(sinφ)\arcsin(\sin\varphi), which never exceeds 48.2°, so that measurement was made on the middle half of the world. It is corrected, and the essay it belongs to is corrected with it — the place whose density changed is now last of seven rather than sixth.

Shape loss is Procrustes distance: both outlines are centred, scaled to unit norm, and the residual minimised over rotation in closed form. It needs no units and no threshold.

The assertions require five things separately: that every drawn area equal its base area times the density asked for; that shape loss fall as growth rises, with a correlation past −0.6; that a uniform request produce growth of exactly one and shape loss of exactly zero; that some region be confusable on a high-contrast request and none on a flat one; and that reading the value without the base areas be wrong by a stated, measurable amount.

The one thing that would fix it

There is a repair, it is not a better solver, and stating it is the useful half of the rung.

Draw the base area on the page. Every published cartogram has room for it: a faint outline of each region at its original size and shape, behind or beside the deformed one, turns the two unknowns a reader is missing into one thing they can see. The value is then read as a ratio of two areas both present on the paper, and the identification is made against the original outline rather than against memory.

That is not a novel suggestion — a map drawn to a density it was handed shows the pair side by side, and so do most careful presentations — but the reasoning for it is usually aesthetic or pedagogical. What the measurements here supply is the size of the thing it fixes: 82 per cent of value error at a factor-of-four spread in base areas, and five confusions in six regions. Those are not decoration.

The alternative repair, which is what non-contiguous constructions do, is to keep every region’s shape exactly and give up adjacency instead. That trades one channel for another and the trade has never been priced against this one. It is the next rung.

Where the model stops

These are discs, not countries. A real region’s outline has its own shape, and a shape with a long axis will lose more or less than a circle depending on how that axis sits against the deformation. Circles are the neutral choice — they have no orientation to be caught out by — so the measurement is a lower bound on what a real boundary would suffer.

Procrustes distance is one similarity measure. A reader does not compute residuals; they recognise. A turning-function distance, a Hausdorff distance or a human trial would each give a different number, and the claim here is about the sign and the ordering rather than about the units.

And the confusion test is harder than reading a real map. Position and label are both available in practice, and both carry more information than shape does. What the test isolates is what the shape channel is worth after the transformation, which is close to nothing — and that is a fact about the picture rather than about any reader.

The generalisation

The rule is that an encoding that uses one visual channel usually damages another, and the damage is not distributed evenly over the data.

A cartogram encodes value in area. Area cannot be changed without changing shape unless the change is a pure dilation, and a pure dilation is available only where the deformation is locally isotropic — which is exactly where the value is high. So the damage falls on the low values, systematically, and a reader trying to check a small number is working with the least legible part of the picture.

The same shape recurs across thematic cartography. A symbol has a size on the page and an area on the ground is the same trade with the symbol standing in for the region; the answer depends on the cells it was counted in is the same statement about aggregation. In each case the encoding is defensible and the unevenness of its cost is what nobody states.

The habit that follows is cheap. When a picture encodes a quantity in a geometric property, ask which other properties had to move, and then ask whether they moved most where the quantity is large or where it is small. If they moved most where it is small, the picture is best exactly where it is least needed.

This collection has been making the same move about projections since measuring instead of naming: a property is not a label, it is a measurement, and the measurement has to be made where the reader will use it rather than averaged over the sheet. What survives a change of coordinates settles which quantities are worth measuring at all. Shape loss under a similarity fit is such a quantity — it does not depend on where the region ended up, how large it became or which way round it is drawn, which is precisely why it is the right thing to ask a reader’s question with.

Who found it, and when

Cartograms have been drawn since the 1860s and the literature about them is almost entirely about construction: how to meet the areas, how fast, with what continuity, with what preservation of adjacency. Tobler’s diffusion method of 2004 is the standard, and the papers that followed it argue about the trade between area accuracy and shape preservation as a global quantity — one number for the map.

That global framing is what hides this. A cartogram’s mean shape distortion is a real number and it is the number everybody reports; splitting it by which regions bore it is one line of extra arithmetic that nobody had cause to write, because the question being asked was always whether the construction was good rather than what the reader could recover.

The one part of the subject that has looked at reading rather than making is the perceptual work on how badly people estimate areas at all, which finds that judged area grows as roughly the 0.7 power of true area — so the encoded quantity is systematically compressed before any of this begins. That is a second, independent problem with reading a cartogram, and the two compound: the values are compressed, and the regions carrying the small ones are the unrecognisable ones.

Where the ladder goes next

Eight rungs have treated a cartogram as a continuous deformation of the sphere, which is what makes the areas exact and the shapes pay. The constructions people actually publish often are not continuous — they cut regions apart, or replace them with circles, and give up adjacency deliberately in exchange for keeping something else. What each of those gives up, and what it buys, is a comparison this anchor has not made.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

AnisotropyAreaCartogramDensityDesignEqual-areaIdentificationInverse problemLegibilityPurposeShape distortionThematic mappingVerification