Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Correlation Criteria

Cross-correlation itself can be computed two ways: directly in the spatial domain — literally sliding the kernel over the search area and summing a per-position inner product, as shown below — or in the Fourier domain via the fast Fourier transform (FFT), which is what locate actually does (see Fourier Domain, below). Both compute the same underlying quantity, but at very different cost: for the sliding sum, evaluated at every candidate offset, against for the FFT, with the number of pixels — a gap that widens sharply as images grow beyond this page's small teaching examples.

Spatial Domain

Cross Correlation (CC) walks through the geometry of locating a point: the kernel/search-area vector chain, solved by finding where their cross-correlation is maximized. This page covers what "cross-correlation" actually means as a formula — several related criteria are used in the spatial domain, differing in how each responds to brightness and contrast differences between the kernel and a candidate window — a same-sized window of the search area at one particular offset — summed pixelwise over index :

  • Cross-Correlation (CC)

  • Normalized Cross-Correlation (NCC)

  • Zero-mean Cross-Correlation (ZCC)

    where and likewise for .

  • Zero-mean Normalized Cross-Correlation (ZNCC)

    where and .

Invariance and Robustness

Invariance describes whether or not a correlation is robust or insensitive to changes in brightness and/or contrast.

  • For brightness, which is additive, subtracting each side's own mean cancels any constant added to that side, making "Zero-mean" approaches effective.
  • For contrast, which is multiplicative, dividing by each side's own norm cancels any constant scaling of that side, making "Normalized" approaches effective.

Whether a criterion performs each of those two cancellations determines its invariance:

MethodInvariant to brightness (additive)Invariant to contrast (multiplicative)Robustness
CC❌ No❌ NoLeast robust — neither cancellation
NCC❌ No✅ YesOnly robust to contrast changes
ZCC✅ Yes❌ NoOnly robust to brightness changes
ZNCC✅ Yes✅ YesMost robust

ZNCC combines ZCC's mean-subtraction (brightness invariance) with NCC's norm-division (contrast invariance), which is why it's the standard choice in most DIC implementations — including dictk.translation.locate's own underlying skimage.registration.phase_cross_correlation call (see Fourier Domain, below).

Neither cancellation helps against nonlinear or spatially-varying brightness/contrast (a shadow crossing part of the kernel, sensor saturation) — none of the four criteria above address that.

Brightness and Contrast Invariance in Practice

The table above is a formula-level guarantee, verified here on astronaut0 — the speckle-over-photograph image used from Multi-Point Motion onward — rather than taken on faith. Extract a kernel from astronaut0 unmodified, then compare it against a search area from a translated and brightness-shifted copy of the same image, using dictk.image.brightness with a small enough factor that no pixel clips at 255 (clipping is a genuine loss of information no correlation criterion can see past, and would contaminate this test):

from dictk.image import read, translate, brightness, PixelCoordinate, subimage
from dictk.correlation import cc, ncc, zcc, zncc

astronaut0 = read(path="astronaut0.png")
p0 = PixelCoordinate(x=100, y=100)
kernel_margin, search_margin = 25, 50
kernel = subimage(image=astronaut0, origin=PixelCoordinate(x=p0.x - kernel_margin, y=p0.y - kernel_margin), width=2 * kernel_margin, height=2 * kernel_margin)

dx, dy = -6, 8
current_baseline = translate(arr=astronaut0, dx=dx, dy=dy)
current_bright = brightness(arr=current_baseline, factor=1.01)  # +1.275 per pixel, no clipping here

search_origin = PixelCoordinate(x=p0.x - search_margin, y=p0.y - search_margin)
baseline_search = subimage(image=current_baseline, origin=search_origin, width=2 * search_margin, height=2 * search_margin)
bright_search = subimage(image=current_bright, origin=search_origin, width=2 * search_margin, height=2 * search_margin)

for name, fn in [("CC", cc), ("NCC", ncc), ("ZCC", zcc), ("ZNCC", zncc)]:
    baseline_peak = fn(kernel=kernel, search=baseline_search).max()
    bright_peak = fn(kernel=kernel, search=bright_search).max()
    pct_change = (bright_peak - baseline_peak) / abs(baseline_peak) * 100
    print(f"{name}: peak value change under brightness shift = {pct_change:+.4f}%")
CC: peak value change under brightness shift = +0.6431%
NCC: peak value change under brightness shift = -0.0007%
ZCC: peak value change under brightness shift = +0.0000%
ZNCC: peak value change under brightness shift = +0.0000%

ZCC and ZNCC come back at exactly +0.0000% — bit-for-bit unchanged, as the formula guarantees for any brightness shift small enough to avoid clipping. CC and NCC both drift, confirming they are not brightness invariant — even though, on astronaut0's strong, distinctive texture, that drift isn't large enough to move where the peak lands, only its value. That value-only distinction still matters in practice: it's what makes CC's raw magnitude unsafe to compare across different points or lighting conditions in a Multi-Point Motion grid, even on images where its peak still happens to land in the right place for any one point in isolation.

A parallel contrast test — dictk.image.contrast instead of brightness, same astronaut0 kernel/search pair — shows the other pairing:

CC: peak value change under contrast shift = +0.2194%
NCC: peak value change under contrast shift = -0.0052%
ZCC: peak value change under contrast shift = +1.8999%
ZNCC: peak value change under contrast shift = -0.0021%

NCC drifts about 40x less than CC does (-0.0052% vs +0.2194%), and ZNCC about 900x less than ZCC does (-0.0021% vs +1.8999%). Not perfectly bit-exact like the brightness case, because contrast scales around the image's own mean rather than performing a pure multiplicative gain, which mixes in a small secondary additive term — but the qualitative result matches the table: contrast invariance belongs to NCC and ZNCC, not CC or ZCC.

See Pan B, Xie H, Wang Z. "Equivalence of digital image correlation criteria for pattern matching." Applied Optics 2010;49(28):5501-9. [download]

dictk.correlation implements all four as standalone functions (cc, ncc, zcc, zncc), each returning the full correlation surface rather than just its peak — see Correlation Visualization for what those surfaces look like on the kernel and search area established in Cross Correlation (CC).

Next: Correlation Visualization visualizes these four correlation criteria in detail; the Fourier Domain section below explains the route locate itself actually takes.

Fourier Domain

Correlation Visualization computes CC directly in the spatial domain: a literal sliding sum, one value per candidate offset. The convolution theorem gives an equivalent route: multiplying the two images' Fourier transforms (one of them conjugated) and inverse-transforming the product yields that same correlation, all at once, for every offset — without ever explicitly sliding a window. This is exactly what dictk.translation.locate does internally, via skimage.registration.phase_cross_correlation. The appeal isn't a different answer — it's speed: a fast Fourier transform (FFT) costs per image, against the sliding sum's per candidate offset — decisive once images grow beyond this page's small teaching example.

reference_image, p0, current_image, kernel, and search are the same as in Correlation Visualization:

from dictk.image import read, translate, PixelCoordinate, subimage

reference_image = read(path="checkerboard0.png")
p0 = PixelCoordinate(x=100, y=75)
current_image = translate(arr=reference_image, dx=-6, dy=8)

kernel_margin = 25
kernel = subimage(
    image=reference_image,
    origin=PixelCoordinate(x=p0.x - kernel_margin, y=p0.y - kernel_margin),
    width=2 * kernel_margin,
    height=2 * kernel_margin,
)

search_margin = 50
search_center = p0
search = subimage(
    image=current_image,
    origin=PixelCoordinate(
        x=search_center.x - search_margin, y=search_center.y - search_margin
    ),
    width=2 * search_margin,
    height=2 * search_margin,
)

locate pads kernel to search's own shape before comparing them, though not quite with the padding used here — see the note below:

import numpy as np

pad_height = search.shape[0] - kernel.shape[0]
pad_width = search.shape[1] - kernel.shape[1]
kernel_padded = np.pad(kernel.astype(np.float64), ((0, pad_height), (0, pad_width)))

image_product = np.fft.fft2(search.astype(np.float64)) * np.fft.fft2(kernel_padded).conj()
fft_surface = np.fft.ifft2(image_product).real

dy, dx = np.unravel_index(np.argmax(fft_surface), fft_surface.shape)
print(f"FFT-domain peak offset (dx, dy) = ({dx}, {dy})")
FFT-domain peak offset (dx, dy) = (19, 33)

That peak, , matches exactly — the same offset Correlation Visualization's cc() surface and locate itself both find. That agreement is about the peak's location only. fft_surface here and locate's own computation differ in three ways, none of which change where the peak lands here, on this page's small, comfortably-within-bounds displacement:

  1. Shape. fft_surface is the circular correlation over the full padded extent (search's own shape, 100x100). cc() returns valid positions only (a smaller 51x51 array, no wraparound). The two arrays don't share a shape, so np.allclose between them wouldn't be meaningful.
  2. Normalization. fft_surface is a raw, unnormalized cross-power spectrum. locate instead passes normalization="phase" to phase_cross_correlation, dividing that spectrum by its own magnitude at every frequency before inverting it (see the extensive comment in locate's source for why).
  3. Padding anchor. kernel_padded above keeps kernel's content anchored at the padded array's top-left corner (np.pad's own default), matching cc()'s corner-offset convention above. locate centers it instead — a reason worth knowing once you've worked with locate a bit more: see Recoverable Displacement Range.