What a machine does with it

The answer depends on the cells it was counted in

Nine essays price the cell as a shape. The number reported out of it is priced nowhere: sliding a grid without changing its resolution moves the largest reported value by 12.6 per cent, which is more than halving the resolution costs.

Nine rungs of this ladder have priced the cell as a shape: whether hexagons can tile the sphere, what a cell system trades away, and what a disc query costs against it: whether it can tile the sphere, what it trades away, how a query interacts with it, what its address means, and what happens when the same data is put on two grids. All of them are about the geometry.

A cell system exists to produce a number. Somebody asks where the maximum is, how strongly two things go together, or how much of the world is above a threshold, and the answer comes out of the cells. That answer is a property of the cells as much as of the ground, and this rung measures how much.

The same resolution, the grid moved, and a different answer. The same field at a fixed 18 × 9 division, with the grid slid by fractions of a cell. Nothing is lost — every cell is the same size as before and there are exactly as many of them — and the largest reported value moves over a range of 12.59 per cent. That is more than a whole halving of the resolution costs, which is 10.15 per cent on the same field. The scale effect has an excuse and this one has none.
Fig. 1 The same continuous field at a fixed 18 × 9 division, with the grid slid by fractions of a cell. Nothing is lost — every cell is the same size as before and there are exactly as many of them — and the largest reported value moves over a range of 12.59 per cent. That is more than a whole halving of the resolution costs, which is 10.15 per cent on the same field.

Two effects, and only one of them has an excuse

The phenomenon splits cleanly into two halves.

The scale effect: the answer changes when the cells get bigger. Everybody expects this and everybody has a story for it — detail is lost, extremes are averaged away, resolution costs information.

The zoning effect: the answer changes when the cells are moved without changing size. Nothing is lost, the cell count is identical, the resolution is identical, and the number is different.

The second is the one with no excuse, and it is the larger of the two on the measurements here.

The field is stated and integrated exactly

Everything below is computed on a stated continuous field — a ridge and two peaks, written as a formula — binned into equal-area cells by exact quadrature over each cell.

That matters because it removes the usual explanation. No data is being resampled, no interpolation is happening, no points are being counted into boxes and no sampling noise is present anywhere. Every cell’s value is the field’s own area-weighted mean over exactly that cell, computed to six figures.

So whatever moves is not information loss in any ordinary sense. It is the arithmetic of averaging a varying quantity over a region, and the region is a choice.

The scale effect, measured

The largest value on the map is a property of the cells. One stated continuous field, integrated exactly over eight different cell systems. Nothing is resampled and no information is interpolated away — every value is the field's own area-weighted mean over the cell it is drawn in. The largest number anywhere falls from 3.179 at 10,368 cells to 1.407 at 32, a loss of 55.8 per cent, and the ground did not change.
Fig. 2 One stated continuous field, integrated exactly over eight different cell systems. The largest number anywhere falls from 3.179 at 10,368 cells to 1.407 at 32, a loss of 55.8 per cent, and the ground did not change. Nothing is resampled and no information is interpolated away.

The maximum falls monotonically as the cells grow, from 3.179 to 1.407 — a loss of 56 per cent of the reported peak.

That is the expected direction and the expected mechanism: a peak averaged over a larger region is diluted by the ground around it, and a cell large enough to contain the whole peak reports the peak’s mean rather than its height.

It is worth noticing how fast it moves. Halving the resolution once, from 18 × 9 to 12 × 6, costs 10 per cent of the reported maximum. A user comparing two datasets binned at different resolutions is comparing two numbers that differ by that much before anything about the world enters.

The zoning effect, measured

At a fixed 18 × 9 division, sliding the grid through one cell in longitude and half a cell in latitude moves the reported maximum by 12.59 per cent.

That is larger than the halving. Moving the grid without changing it costs more than throwing away three quarters of the cells.

One field, one resolution, three placements of the grid. The same continuous field binned into 18 × 9 equal-area cells three times, with the grid slid between the panels and nothing else changed. The largest cell reports 2.267, 2.402, 2.360 in the three, and a reader looking for where the field is highest is shown a different cell each time.
Fig. 3 The same field binned into 18 × 9 equal-area cells three times, with the grid slid between the panels and nothing else changed. The largest cell reports a different value in each, and a reader looking for where the field is highest is shown a different cell each time.

The mechanism is simple and is not usually stated. A peak that falls at the centre of a cell is averaged with the low ground on all sides of it; a peak that falls at a cell corner is split between four cells and each of them reports a quarter of it against three quarters of low ground. Where the boundaries fall relative to the feature decides the number, and where they fall is an arbitrary choice made when the grid was defined.

Tangent-warped cube cells, shaded by area. The tangent-warped cube at level 3, drawn on Mollweide with each cell shaded by its own measured area. The largest cell is 1.29 times the smallest. Every area is computed with the spherical polygon formula from the cell's own boundary, not from the scheme's intentions, and the shading is what the numbers say rather than what the mesh looks like.
Fig. 4 A cell system of the kind nine rungs of this ladder have priced, shaded by the area each cell actually covers. Everything measured in this rung is downstream of a choice like this one and of one more that nobody records: where its boundaries were put.

The two effects are entangled

There is a subtlety in the sweep worth being explicit about.

Changing the resolution also moves every cell boundary. So a resolution sweep carries a zoning change inside it, and the two cannot be separated by varying resolution alone — which is why the second measurement, at fixed resolution, is the one that isolates the effect with no excuse attached.

Run the resolution sweep with the grid aligned differently and the numbers move; run it at several alignments and the spread at each resolution is the zoning effect. What is reported as the scale effect in any single sweep is a scale effect plus one draw from a zoning distribution.

The correlation moves too, and further

How strongly two things go together is decided by the cells they were counted in. The correlation between two stated fields, computed from their cell means. The line is the resolution sweep and the vertical spreads are what sliding the grid does at a fixed resolution. The resolution moves the correlation by 0.070 across the whole range; sliding the grid at 12 × 6 moves it by 0.109, which is more. Neither field changed and no data was lost in either case.
Fig. 5 The correlation between two stated fields, computed from their cell means. The line is the resolution sweep and the vertical spreads are what sliding the grid does at a fixed resolution. The resolution moves the correlation by 0.035 across the whole range; sliding the grid at 12 × 6 moves it by 0.109, which is more. Neither field changed and no data was lost in either case.

The correlation between two fields rises as the cells grow — 0.502 at the finest resolution to 0.537 at the coarsest — which is the classic result: aggregating suppresses the fine-scale variation that the two fields do not share, and what is left is the coarse structure they do.

The zoning effect on the correlation is three times larger than the scale effect. Sliding the grid at a fixed 12 × 6 division moves the correlation by 0.109; the entire resolution sweep, over a factor of three hundred in cell count, moves it by 0.035.

That is the strongest form of the result. A correlation computed from cell means is decided more by where the cell boundaries fall than by how many cells there are, and the placement is not something anybody reports.

Why the correlation is the worst case

A correlation is not merely another statistic that moves. It moves for a reason that makes it uniquely unsafe, and the reason is worth stating.

The correlation of two fields over a set of cells is a ratio whose numerator is a covariance and whose denominators are two variances, and every one of the three is an average of squares. Averaging a quantity over a cell reduces its variance by the amount of variation that happens to fall inside the cell, and how much falls inside depends on where the boundaries are relative to the field’s own structure.

So the boundaries decide the denominators, and the two denominators are decided differently because the two fields have different structure. A grid placement that happens to split one field’s peaks and not the other’s suppresses one variance and not the other, and the ratio moves.

That is why the zoning effect is three times larger on the correlation than on the maximum here. The maximum is decided by one cell; the correlation is decided by every cell, through three sums that respond to the placement in different directions.

The same address length, a tenth of the area. Every cell of a lon/lat quadtree at level 3 carries an identifier of the same length. The heavy curve is each cell's area as a fraction of the largest, against its latitude: a polar cell is 5.0 times smaller than an equatorial one. The light curve is the inverse of the cell's aspect ratio, which falls from 0.98 near the equator to 0.20 at the top — the cells stop being anything like square long before they stop being usable.
Fig. 6 The same object from the ladder’s earlier work, for contrast: a cell’s address and the precision it implies, which is a property of the scheme and not of the data in it. The address is stable under everything measured in this rung — sliding the grid changes which cell a place is in and not what an address means — which is exactly why the addressing rungs could be written without any of this.

What a reader can and cannot take from a binned map

Two things survive and one does not.

The total survives — an area on a sphere has a closed form and the binning here integrates it exactly: the sum of the cell values weighted by cell areas is the integral of the field, at every resolution and every placement, because the binning is exact. Anything that is an integral over the whole domain is safe.

The rank ordering of large regions largely survives, because a feature much bigger than a cell is described by many cells and their arrangement matters less.

The extremes do not, and neither does anything derived from them: the maximum, the location of the maximum, the number of cells above a threshold, and the correlation, which is an extreme-sensitive statistic because it is dominated by the tails of both fields.

The area above a threshold

The third statistic is the one that appears most often in reporting and is the least stable of the three.

How much of the world is above this value is a question a binned map answers by adding up the areas of the cells whose means exceed the threshold, and a cell’s mean exceeds the threshold or does not — there is no partial credit. So a feature straddling four cells contributes nothing if none of the four averages above the line, and contributes a whole cell if one of them does.

Measured on this field, sliding the grid at 18 × 9 moves the area above a threshold by 0.62 percentage points against a value of about 1.4 per cent — which is a relative movement of nearly half. At 12 × 6 it moves by 1.39 points against 1.4, which is a relative movement of very nearly one hundred per cent.

That statistic is not merely uncertain. At coarse resolutions its value is roughly as large as its own variation under a choice nobody records.

The refusal

A uniform field is reported identically by every cell system at every placement — the same value in every cell, the same maximum, the same correlation with any other field, the same everything.

So the machinery is capable of reporting no effect, and it does whenever the ground has no structure at the scale of the cells. Every number above is therefore a measurement of an interaction between the field’s own structure and the grid’s, rather than an artefact of the binning code.

What to do about it

Nothing removes the effect, because it is not an error. Three things reduce the damage.

Report the placement, which is metadata of the same kind a cell address already carries. A grid is defined by an origin as well as a resolution, and the origin is metadata that costs nothing and is almost never recorded. An address is an area makes the same point about a cell identifier; the grid’s own origin is the same information one level up.

Sweep the placement rather than choosing one. Computing a statistic at several offsets and reporting the spread turns an unquantified arbitrary choice into a stated uncertainty, and it costs one loop.

Prefer statistics that are integrals. A total, a mean weighted by area, and anything that is a linear functional of the field are stable under both effects. A maximum, a threshold count and a correlation are not.

Where this sits against the rest of the field

The attribute is a claim about the geometry makes the general version of the argument one anchor over: a value stored against a geometry is only meaningful against that geometry, and every operation that changes the geometry invalidates it.

This rung is the case where the operation is binning and the geometry is a grid nobody thinks of as a choice. A projection is chosen; a cell system is chosen; a grid origin is defaulted, and the default is doing as much work as either of the other two.

Where the model stops

The field here is smooth and has two clear peaks, which makes the effects visible and probably understates them. A field with more structure at the cell scale — a point pattern, a coastline, an urban distribution — has more of its variance at the frequencies the binning interacts with, and the zoning effect is correspondingly larger.

What is also omitted is the shape of the zones. Everything here is a lat–lon banding slid in two directions, so the freedom explored is two numbers. Real aggregation units — districts, catchments, postal areas — differ in shape as well as in placement, and the freedom is then combinatorial rather than two-dimensional.

The comparison that quietly assumes a grid

The failure that does the most damage in practice is a comparison between two datasets that were binned separately.

Two agencies publish gridded products of the same quantity. Both use equal-area cells; both use a sensible resolution; neither states its origin, because an origin is not the kind of thing that goes in a product description. The two grids are offset by some fraction of a cell.

A user differencing them gets a field whose structure is partly the difference in the quantity and partly the difference in the zoning, and there is no way to separate the two from the products alone. The difference field even looks plausible — it has structure where the quantity has structure, because that is where the zoning effect lives.

That is why the recommendation to sweep the placement is not academic. A product that reported its statistic at four offsets, with the spread, would let a user tell a real difference from a zoning artefact in one step; a product that reports one number cannot.

What this does to the site’s own machinery

It is worth saying that this collection’s own binning is not exempt, and that the exemption it does have is narrow.

Every quantity this site computes from cells is an integral — an area, a total, a mean weighted by area — and integrals are the class that survives. That was not a defensive choice; it is what the questions happened to be. The moment a figure here reports a maximum over cells, or a count above a threshold, it inherits everything in this essay.

The one place it already did is worth naming. The same data on two grids measures what rebinning costs, and its round-trip error is a scale effect with a zoning effect inside it, exactly as described above — the source and target row edges coincide wherever their counts share a factor, so the measured advantage of one geometry over another was partly a measurement of arithmetic. Sliding the target is what separated them, and it is the same manipulation this rung is built on.

Who found it, and when

The modifiable areal unit problem was named by Stan Openshaw and Peter Taylor in 1979, and the underlying observation is older: Gehlke and Biehl reported in 1934 that correlations between census variables rose as tracts were grouped, and could not say why.

It is worth noting that the effect is not the projection’s: every scheme here is equal-area by construction, so nothing about it is the areal factor misbehaving. Openshaw’s demonstration was the sharp one. Given a set of zones and two variables, he searched over regroupings and produced correlations anywhere between strongly negative and strongly positive from the same underlying data — which is a stronger statement than anything measured here, because his freedom was combinatorial and this rung’s is two-dimensional.

What this rung adds is that the effect survives every excuse. No sampling, no interpolation, no resampling, no data loss, exact quadrature on a stated field: and the number still moves by more than a halving of the resolution when the grid is nudged.

The sensitivity test the effect supplies

The measurement is a warning, and it also supplies its own diagnostic, which is worth stating because it costs one extra run of an analysis that has already been written.

Nudge the grid and recompute. Shift the origin by a fraction of a cell, in a few directions, and look at the spread of whatever number the analysis produces. That spread is not an estimate of the zoning effect; it is the zoning effect, measured on this dataset with this field at this resolution, and it needs no theory, no null model and no assumption about the data.

The reading is the same one-sided reading this collection keeps arriving at. A large spread proves the number is an artefact of the grid; a small spread proves nothing — a jitter of the origin explores only translations, and Openshaw’s regroupings show how much more freedom a zoning has than that. What a small spread does establish is that the cheapest and commonest source of the effect is not present, which is worth knowing before spending anything on the expensive ones.

And the spread is the thing to publish beside the number. A correlation of 0.62 means something different when the same pipeline gives 0.61 to 0.63 under a nudge than when it gives 0.41 to 0.78, and no reader can distinguish those two situations from the published figure. It is the same move as reporting a residual with a fit: the number alone is a claim, and the number with its sensitivity is a claim somebody else can act on.

Where the ladder goes next

Ten rungs have priced the cell and now the number that comes out of it. The unasked question is temporal: a cell system that is redefined between two editions makes every comparison across them a comparison of two zonings, which is the same effect with a date attached.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

AggregationBinningCell systemCorrelationEcological fallacyEqual-areaGrid originMaupResolutionScale effectStatisticZoning effect