Measuring distortion

A dot map's density is partly the projection's

A dot map carries the right number of dots in every region whichever way it is drawn, so it is honest as a total under both placements. It cannot be honest as a density under both: ground on a uniform field reads 0.099 of its equatorial density at 72° north on Mercator, and scattering inside the polygon on the page moves 64.3 per cent of a cell's dots into its northern half without one of them leaving the cell.

Assumes A symbol has a size on the page and an area on the ground.

A dot map is the most honest thematic form there is, and it is usually introduced that way. One dot means one hundred of the thing; count the dots and the total comes back; nothing is aggregated into a class, nothing is smoothed, nothing is coloured by a scheme somebody chose.

Every word of that survives this essay. The count is right. What is not right is the only thing anybody actually reads off a dot map, which is how thick the dots look.

The same count, scattered two ways. 22 dots in every cell of a 30-cell covering, drawn on Mercator. The upper panel places them uniformly on the GROUND — uniform in longitude and in the sine of latitude, which is what uniform on a sphere means — and the lower places them uniformly on the PAGE, which is what a drawing routine handed a polygon does. Both panels carry exactly the same number of dots in exactly the same regions, so both are honest as totals. They are different pictures, and a reader reads a dot map by density.
Fig. 1 Twenty-two dots in each of thirty cells, drawn on Mercator. The upper panel scatters them uniformly on the ground — uniform in longitude and in the sine of latitude, which is what uniform on a sphere means. The lower scatters them uniformly on the page, which is what a routine handed a polygon and a random number generator does. Both panels carry exactly the same dots in exactly the same regions.

Two placements, one count

A dot map is made by two decisions and only the first is ever discussed. The first is how many dots a region gets, and it is the data: value divided by the dot value, rounded. The second is where inside the region they go, and it is almost always left to whatever the drawing routine does.

What the drawing routine does is rejection sampling in the polygon’s bounding box, in the coordinate system the polygon is stored in — which is the page. That is uniform on the page.

The alternative is to scatter uniformly on the sphere, which means uniform in longitude and uniform in the sine of latitude, the same closed form every area calculation on this site rests on. That is uniform on the ground.

Both produce the same number of dots in the same regions. Neither loses or invents a single one. And they are different maps of the same number, because a dot map is read by density and density is dots per unit of something.

The identity, twice

Let a region carry n dots over ground area A, drawn with page area As.

Scattering on the ground makes the ground density n/A uniform, and the page density is then n/(As) — it carries the whole areal factor. The dots thin out exactly where the map inflates.

Scattering on the page makes the page density n/(As) uniform inside the region, so the implied ground density is ns/A — the factor moves to the other side.

Either way the areal factor is in one of the two densities, and it cannot be in neither, because their ratio is the areal factor. The choice is which of the two readings to make wrong, and it is the same forced trade as a proportional symbol has to choose between its total and its density, one rung down, with the roles of the two quantities swapped.

What was computed, and how

The field is a constant ground density: the same number of dots per unit area of ground everywhere, which is the cleanest possible input because any variation in the picture is then entirely the map’s.

A uniform ground, drawn sparse. A field of constant ground density — the same number of dots per square kilometre everywhere — drawn with the dots placed on the ground and the visible page density measured. It falls away from the equator on every projection that inflates area, because the same dots are spread over more paper: on Mercator the page density at 72° is 0.10 of its equatorial value. The ground is uniform. The picture of it is not, and the flat line is the equal-area member.
Fig. 2 A perfectly uniform ground, drawn with the dots placed on the ground, and the visible page density measured against latitude. It falls away from the equator on every projection that inflates area, because the same dots are spread over more paper.
latitude Mercator Miller plate carrée Gall–Peters
1.000 1.000 1.000 1.000
28° 0.793 0.828 0.891 1.000
50° 0.424 0.504 0.652 1.000
72° 0.099 0.172 0.319 1.000

A ground that is uniform everywhere is drawn ten times sparser at seventy-two degrees north than at the equator. Nothing has been left out; the dots are all there, on a sheet that gave that latitude ten times the paper.

The reading a viewer forms — this phenomenon thins out towards the pole — is a reading of the projection, and it is the exact opposite of the impression the choropleth’s covariance produces, where the high latitudes are given more weight than they deserve. Both are the areal factor; the sign differs because one is a total spread over paper and the other is a colour filling it.

The two placements, and why one of them has to be wrong. Scattering on the ground puts the areal factor into the density a reader SEES; scattering on the page puts it into the density the map IMPLIES about the ground. On Mercator at 77° the first reads 0.051 of its equatorial value and the second 19.76 — reciprocals, because their ratio is the areal factor by construction. The product of the two curves is one at every latitude, which is the statement that there is no third placement.
Fig. 3 The two placements on one axis. Scattering on the ground puts the areal factor into the density a reader sees; scattering on the page puts it into the density the map implies about the ground. The curves are reciprocals, so their product is one at every latitude — which is the statement that there is no third option.

The product is one, so there is no third placement

The two curves multiply to one everywhere, and that is not a coincidence to be checked but the definition rearranged. The seen density and the implied density differ by the areal factor; making one of them uniform forces the other to carry it; and a placement rule that made both uniform would be a map with an areal factor of one, which is an equal-area projection.

So the trade is closed. There is no clever sampler, no weighting, no jitter that produces a dot map correct in both senses on a projection that is not equal-area. This is the same shape of argument as the trade-off is two lines — the site’s oldest one — arriving at a much smaller object: two conditions, one degree of freedom, and the only way to satisfy both is for the thing being traded to be absent.

What that buys is a clean statement of what a dot map on Mercator is. It is a correct map of totals, drawn with the dots deposited where the paper is rather than where the ground is, and its visible texture is the paper’s.

The skew inside a region, which no count can see

The between-region effect above is at least visible in principle: a reader who knows the projection can allow for it. The within-region effect is not, and it is the finding of this rung.

Take one cell running from 60° to 75° north and cut it at the latitude that halves its ground area — 66.34°, by the closed form. Scatter four thousand dots in it and count how many land in the northern half.

Where inside one region the dots land. Four thousand dots scattered inside a single cell running from 60° to 75° north, and the share of them that falls in the northern half of it by ground area. Placed on the ground the share is 50.0%, which is what an equal-area cut means. Placed on the page it is whatever the projection's north–south stretch makes it: 64.3 per cent on Mercator. The dots have moved without leaving the region, so no count anywhere has changed.
Fig. 4 Four thousand dots inside one cell, and the share of them that falls in its northern half by ground area. Ground placement gives 49.95 per cent, which is an equal split to sampling noise. Page placement gives whatever the projection’s north–south stretch makes it. No dot has left the cell, so no count anywhere on the map has changed.
projection share in the northern half
ground placement, any projection 49.95%
Mercator 64.30%
Miller 60.20%
plate carrée 56.95%
Robinson 53.20%
Mollweide 49.20%
Gall–Peters 48.85%

On Mercator, page placement puts nearly two thirds of a region’s dots into the half of it that has half the ground. The region’s total is untouched, the map’s totals are all untouched, and a reader looking at the picture sees a concentration that is a property of the sheet.

This is the failure that has no reader-side repair. A reader who knows the projection can mentally correct a between-region density comparison, in the same way a reader who knows Mercator can allow for Greenland. A concentration inside a region looks exactly like data, because within one region a reader has no reason to expect the map to be doing anything at all.

The skew has a closed form, and the sampling agrees with it

The table above is measured by scattering four thousand dots and counting, which is the right way to demonstrate the effect and is not the only way to know it. For a cylindrical projection the answer is a ratio of two lengths on the page and needs no dots at all.

Page-uniform placement inside a graticule cell is uniform in the projection’s own northing, so the share of dots falling north of a stated parallel is the share of the cell’s page height above it. The cut is at 66.34°, the latitude that halves the ground area between 60° and 75°, so the closed form is

y(75)y(66.34)y(75)y(60)\frac{y(75^\circ) - y(66.34^\circ)}{y(75^\circ) - y(60^\circ)}

with y the projection’s northing function. On Mercator, whose y is ln tan(45° + φ/2), the three values are 1.3170, 1.5657 and 2.0275, and the share is 0.6500. On the plate carrée, whose y is the latitude itself, it is (75 − 66.34)/15 = 0.5773.

Against the sampled 64.30 and 56.95 per cent those are agreements rather than checks: four thousand dots give a standard error of about 0.8 percentage points on a proportion near a half, and both closed forms are inside one of them.

Two things follow from having the expression rather than the number. The skew is a property of the cell and the projection alone — it does not depend on how many dots the region carries, so a sparsely populated region is skewed exactly as much as a dense one and the effect cannot be averaged away by drawing more dots. And it can be computed for any cell before anything is drawn, which makes it available as a warning rather than as a diagnosis: a producer can be told that this particular region on this particular sheet will put 65 per cent of its dots in the half that holds half its ground, and can decide what to do about it while there is still something to decide.

The expression also says where the effect goes to zero. It is a half exactly when y is a function of the sine of the latitude — which is the equal-area condition for a cylindrical projection, and is why Gall–Peters and Mollweide come back at a half in the table. The closed form and the measurement agree about that too, and they agree for the same reason the whole rung has: there is one degree of freedom and the equal-area condition spends it.

Why the equal-area map does not fix everything either

Gall–Peters comes out at 48.85 per cent, which is a half to within sampling noise, and Mollweide at 49.20. An equal-area projection makes the page-uniform and ground-uniform placements the same operation, so the whole of this rung collapses on one.

That is a stronger result than the previous rung got. Equal-area repaired the symbol’s density reading and left the crowding alone, because crowding depends on a principal scale rather than on the product. A dot map has no crowding constraint of that kind — a dot has no size that a legend has promised anything about — so here the equal-area condition is not merely sufficient, it is the whole answer.

The same count, scattered two ways. 22 dots in every cell of a 30-cell covering, drawn on Gall–Peters. The upper panel places them uniformly on the GROUND — uniform in longitude and in the sine of latitude, which is what uniform on a sphere means — and the lower places them uniformly on the PAGE, which is what a drawing routine handed a polygon does. Both panels carry exactly the same number of dots in exactly the same regions, so both are honest as totals. They are different pictures, and a reader reads a dot map by density.
Fig. 5 The same construction on Gall–Peters. The two panels are the same picture, because on an equal-area projection uniform-on-the-page and uniform-on-the-ground are the same operation — the sampler has nothing left to disagree about. This is the control for the hero figure, and the difference between the two figures is the whole of this rung.

Which makes dot maps the one thematic form where the standard advice is exactly the right advice, for exactly the reason usually given, with nothing left over.

That is worth saying plainly because the collection has spent two rungs finding residuals in the advice. A choropleth on an equal-area map is unbiased in its mean and can still be misread through its classification; a symbol map on an equal-area map has its density fixed and its crowding made worse. A dot map on an equal-area map has no residual of either kind, and the reason is structural: the only page quantity a dot map commits to is area, and area is the quantity the condition pins.

The other repair, and why it is not used

There is a second fix that works on any projection: scatter on the ground rather than on the page. It costs one inverse projection per dot, it produces the correct within-region distribution on every sheet, and it leaves the between-region density carrying the areal factor — which is at least the honest half, since it is the half a reader can be told about.

It is not what any drawing pipeline does, and the reason is architectural rather than considered. A polygon arrives in projected coordinates, a random point generator works in the plane it is handed, and rejecting points outside the polygon is a two-line routine that has been in every graphics library since there have been graphics libraries. Sampling on the sphere requires the polygon to be carried back through the inverse projection first, which is the operation deciding the coordinate system it is performed in — the general rule this site has been stating in one form or another for two hundred essays.

What a reader could be told, and what would not help

The three failures in this rung have different repairs and it is worth separating which are the map’s to make.

The between-region density is correctable by a reader, in principle, and the correction is the same one everybody already applies to Greenland. A note saying dots are placed on the ground, so the visible thinning towards the poles is the projection is one sentence and closes it, and it is the same class of repair as a scale bar right in one place admitting that it is.

The within-region skew is not correctable by a reader at all. There is nothing to correct: the reader cannot know that the concentration in the north of a region is the sheet, because within one region the sheet is not visibly doing anything. Only the producer can fix it, and only by changing the sampler.

Equal-area does not stop the symbols crowding. Two marks four degrees of ground apart, drawn on four projections, along a parallel (rising curves) and along a meridian (falling ones), relative to the same pair at the equator. Whether two symbols collide is decided by this quantity against the symbol's diameter. Gall–Peters holds its areal factor at exactly one everywhere and still compresses a north–south pair to 0.174 of its equatorial separation while stretching an east–west pair to 4.81. The two failures are separate and only one of them has an equal-area repair.
Fig. 6 Why the saturation limit is not shared out evenly either. Two marks four degrees of ground apart, drawn along a parallel and along a meridian, relative to the same pair at the equator. Where a curve falls, the dots are being pushed together and the map saturates sooner — and Gall–Peters, which fixes every density question in this rung, is the projection that compresses a north–south pair hardest.

The saturation threshold is nobody’s to correct. Where the map compresses, the dots crowd, and past a threshold the count stops being recoverable by eye. That is a legibility limit rather than a bias, and it is the one failure of the three where an equal-area projection genuinely does remove the whole problem rather than merely equalising it.

Ranking those three by how much they matter is a judgement rather than a measurement, but the middle one is the one this rung exists to name, because it is the only one that is invisible in the finished picture.

Where the model stops

Dots are not points. A drawn dot has a radius, dots overlap, and past some density a dot map saturates and stops being readable as a count at all. The saturation threshold is a page-density threshold, so it too inherits the areal factor: a map saturates first where it compresses. Nothing here prices that, and it interacts with the within-region skew in a way that would need the symbol footprint of the rung below.

Real dot maps are not random. Better cartographic practice places dots against ancillary data — population within the enumeration unit, land cover, exclusion of water — precisely because uniform placement inside an administrative unit is known to be a fiction. That practice repairs the within-region distribution for a reason unrelated to projection, and it repairs this one as a side effect if the ancillary data is handled on the ground.

The cells here are cells. The within-region skew grows with the north–south extent of the region: a fifteen-degree cell shows 64 per cent, and a one-degree cell would show a fraction of a per cent. So the effect is a large-region effect, and it is largest for exactly the enumeration units — Canadian territories, Russian federal subjects, Australian states — that are already the hardest cases for a dot map.

The generalisation

The pattern is one this collection has met with a different object each time: an operation that is not equivariant under the map, performed after the map instead of before it.

A centroid computed in a plane belongs to that plane. A buffer computed in degrees is an ellipse rather than a circle. A simplification carried out after projecting is a different simplification. A strain rate differenced from grid coordinates is partly a measurement of the grid.

Random sampling joins the list, and it is a slightly worse case than the others because the result carries no residual. A centroid in the wrong plane can be compared against one computed properly; a simplified line has a measurable deviation. A scatter of dots has nothing to compare against — it was random, it looks random, and a second run of the same routine produces a different picture that is wrong in the same way.

Who found it, and when

Dot maps are usually dated to Frère de Montizon’s 1830 map of the population of France, and the technique’s statistical properties have been studied since Wright’s work on dot placement in the 1930s, which is where the ancillary-data practice comes from. That whole literature is about placement within a unit relative to what is known about the ground, and it is the right question.

The projection question sits underneath it and is rarely separated out, for a reason that is honourable: national and regional dot maps are drawn on national grids, where the areal factor is within a fraction of a per cent of one across the sheet and none of this exists. It becomes a factor of ten on a world map, and a world dot map is precisely the case where the ancillary data that would otherwise have saved the placement is least likely to be available.

Where the ladder goes next

The three rungs so far price readings a viewer forms by eye — a mean, a ratio, a thickness. The last one prices something the software computes: the class boundaries a classifier picks. Those are computed from geometry the machine has, which is projected geometry, and moving them moves regions between colours. On a five-class quantile classification of one stated field, half the regions on the map change colour.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

Area weightingAreal factorBiasDensityDot mapEqual-areaMercatorProportional symbolQuadratureSamplingTest fieldThematic mapping