Skip to content

_promote_float casts 32/64-bit integer rasters to float32, losing integer precision above 2**24 #3684

Description

@brendancol

Description

_promote_float in xrspatial/convolution.py maps every integer dtype to float32:

def _promote_float(dtype):
    """Return at least float32; preserve float64."""
    if np.issubdtype(dtype, np.floating):
        return dtype
    return np.float32

float32 has a 24-bit mantissa, so it cannot represent integers above 2**24. Any consumer that feeds it a 32- or 64-bit integer raster with values past ~16.7M loses the low bits before the kernel runs. This produces silent wrong results through at least convolve_2d / convolution_2d and focal.mean (and apply / focal_stats, which follow the same contract per #2769).

#3680 hit this through edge_detection and scoped its fix to that module (pre-cast wide ints to float64 before convolve_2d). The shared helper still has the problem, and it is a contract question rather than a one-line patch: the int-to-float32 promotion was ratified three times (#1096, #2769, #3214), there is an explicit contract test (test_mean_int_promotes_to_float32_3214), and widening int32/int64 to float64 changes output dtype for integer inputs and doubles working memory in the focal paths whose budget guards use _promote_float(agg.dtype).itemsize.

Reproduction

import numpy as np
import xarray as xr
from xrspatial.convolution import convolution_2d
from xrspatial.focal import mean

base = 100_000_000
data = (np.arange(30).reshape(5, 6) % 6 + base).astype(np.int32)
kernel = np.array([[-1, 0, 1], [-2, 0, 2], [-1, 0, 1]], dtype=np.float64)

print(convolution_2d(xr.DataArray(data), kernel).data[2, 2])
# 0.0        (float64 reference: 8.0)

print(mean(xr.DataArray(data)).data[2, 2])
# 1e+08      (float64 reference: 100000002.0)

Options

  1. Widen _promote_float: int8/int16/uint8/uint16 to float32 (exact), int32/uint32/int64/uint64 to float64 (exact for 32-bit, exact up to 2**53 for 64-bit). Update the contract test, the focal memory-guard comments, and docstring examples that show float32 output for int input.
  2. Keep the contract and document the 2**24 limit loudly in convolve_2d, convolution_2d, focal.mean, apply, focal_stats.
  3. Keep float32 but warn at runtime when an integer input's max abs value exceeds 2**24.

Option 1 is the only one that fixes the results. The cost is a dtype change for integer inputs and 2x memory on those paths.

Follow-up from #3680, found by /sweep-security against edge_detection. Filed against the convolution module for its next sweep or a maintainer call on the contract.

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingseverity:highSweep finding: HIGHsweep-securityFound by /sweep-security

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions