Skip to content

fix: keep NaN distinct from NULL in float columns when pandas is enabled - #979

Open
maharanay22 wants to merge 1 commit into
databricks:mainfrom
maharanay22:fix-nan-float-null-conversion
Open

maharanay22 wants to merge 1 commit into
databricks:mainfrom
maharanay22:fix-nan-float-null-conversion

Conversation

@maharanay22

Copy link
Copy Markdown

What type of PR is this?

  • Bug Fix

Description

Fixes #978.

With pandas enabled (the default), ResultSet._convert_arrow_table maps float32/float64 columns to pandas' nullable Float32Dtype/Float64Dtype. That conversion turns IEEE NaN into pd.NA, and to_numpy(na_value=None) then returns None, so a NaN in a FLOAT or DOUBLE column could not be told apart from SQL NULL. With _disable_pandas=True the same data already came back as float('nan'). For example, [NaN, NULL, inf, 1.5] read back as [None, None, inf, 1.5] by default and [nan, None, inf, 1.5] with _disable_pandas=True, on both pandas 2.3.3 and 3.0.6.

After the object array is built, NaN is set back on the cells that are NaN in the Arrow column, using a vectorized pyarrow.compute.is_nan mask. NULL stays None and other values are unchanged. Tables without float columns do no extra work, and for float columns the cost is one boolean mask per column (about 5 ms per million rows locally, next to roughly 80 ms for the existing pandas conversion).

This is a user-visible change: code that treated a NaN result as None will now see float('nan'). Changelog entry added under # Unreleased.

How is this tested?

  • Unit tests

New tests in tests/unit/test_pandas_compatibility.py check exact row values for float32 and float64 columns containing NaN, NULL, inf, -inf and a normal value (split over two chunks, after a non-float column), and check that the default and _disable_pandas=True paths return the same rows. Both fail on main and pass with this change, under pandas 2.3.3 and 3.0.6. Full unit suite: 1022 passed, 5 skipped.

Related Tickets & Documents

Fixes #978. Related: #960 edits the same function for nested types. This change sits right after to_numpy and only uses the table that went through pandas, so either PR rebases onto the other with a one-name change (table_renamed to scalar_table). Both PRs also add a # Unreleased changelog section, so the second one to merge will need a trivial changelog rebase.

With pandas enabled (the default), _convert_arrow_table maps float32 and
float64 columns to pandas' nullable Float32Dtype/Float64Dtype. Converting
an Arrow float column to those dtypes turns IEEE NaN into pd.NA, and
to_numpy(na_value=None) then returns None for it, so a NaN value could
not be told apart from SQL NULL. With _disable_pandas=True the same data
already came back as float('nan').

After building the object array, set NaN back on the cells that are NaN
in the Arrow column, using a vectorized pyarrow.compute.is_nan mask.
NULL stays None, other values are unchanged, and tables without float
columns take no extra work.

Signed-off-by: Maha Rana Yadavalli <271375718+maharanay22@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

NaN in FLOAT/DOUBLE columns is returned as None (same as NULL) when pandas is enabled

1 participant