← All articles
technologydatameasurementscienceSeptember 17, 20263 min read

How Do You Know the Satellite Is Right? Go and Look

By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.

Any system that infers what is on the ground from a distance has to be checked against direct observation of the same place. That checking is unglamorous and decides whether the results mean anything.

What the term means

The phrase names information gathered by direct observation at the location itself, used to check or to train a system that infers the same information indirectly. A satellite records how much light of various wavelengths a patch of land reflects, and a model turns that into a claim that the patch is wheat, or flooded, or burnt. Somebody standing in that field, recording what is actually there and when, provides the measurement against which that claim is judged. The term originated in remote sensing and has spread to any field where a model's output must be compared with reality.

What the checking is used for

The same data serves several distinct purposes and confusing them causes trouble:

  • Training, where examples teach a model what each signal corresponds to
  • Validating, where held-back examples test how well it performs
  • Calibrating instruments against a known reference
  • Correcting for atmospheric effects between sensor and ground
  • Establishing what accuracy can be claimed for a published product
  • Investigating why a model fails in particular conditions

Why it is harder than it sounds

Collecting the reference data well is a serious methodological problem rather than a matter of going outside. Timing must match, since a field photographed in May and visited in August may have changed completely. Scale must match, since a satellite pixel may cover a hectare containing several different things while an observer records one point. Locations must be chosen without bias, and the accessible places beside roads are systematically different from the rest. The observer's own judgement introduces error, since two people classify the same vegetation differently. And using the same data for both training and testing produces a model that reports excellent accuracy and performs badly on anything new.

How much is needed

Collecting reference data is expensive, so how much is required is a practical question with a reasonably well understood answer. Accuracy assessment needs enough samples in every category, including rare ones, which means a stratified design rather than a simple random scatter, since a random sample of a landscape that is ninety per cent forest produces almost nothing for the other categories. Published guidance suggests a minimum of several dozen samples per class as a working rule. Rare classes therefore dominate the cost. Some of the demand can be met by high-resolution imagery interpreted by an expert instead of a site visit, though that substitutes one inference for another and has to be justified.

Where the phrase has spread

The concept now appears wherever automated inference needs checking, and recognising it clarifies what those systems actually rest on. Machine learning generally depends on labelled examples, and the quality of any model is bounded by the quality of those labels, which are produced by people making judgements and are frequently inconsistent. Medical imaging systems are judged against biopsy results or against a panel of specialists. Mapping companies drive streets to check what their automated systems inferred. Weather forecasts are scored against station observations. In every case the reference is itself imperfect, which means the accuracy figures reported are relative to a standard that has its own errors.

The takeaway

Direct observation at a location provides the reference against which a system inferring the same thing remotely is trained and judged. Timing, scale and location choice all have to match or the comparison is meaningless, and roadside accessibility biases the sample. Using the same examples to train and to test produces flattering accuracy and poor performance. The reference itself has errors, so reported accuracy is relative to an imperfect standard.

Practise this

Questions from Technology and Society

Reading about something is not the same as being able to recall it. These are real questions from the Technology and Society unit in our Technology track, answers and explanations included. The unit has 121 in total across 23 steps.

  • Fill the blankLevel 2

    1. Keeping your personal information from being shared without permission is called ____.

    • privacycorrect
    • gaming
    • charging
    • printing

    Privacy means having control over who can see and use your personal information.

  • True or falseLevel 1

    2. Computers and phones use electricity, which uses energy.

    Answer: True

    True, all our devices run on electricity, and making that electricity uses energy.

  • Type the answerLevel 2

    3. The gap between people who can access technology and those who cannot is called the digital ____.

    Answer: divide

    Closing the digital divide means helping more people get online.