github unionai-oss/pandera v0.34.0

5 hours ago

⭐️ Highlights

This release adds TensorDict as a supported data container that can be validated with pandera:

import torch
from tensordict import TensorDict
import pandera.tensordict as pa

schema = pa.TensorDictSchema(
    keys={
        "observation": pa.Tensor(dtype=torch.float32, shape=(None, 10)),
        "action": pa.Tensor(dtype=torch.float32, shape=(None, 5)),
    },
    batch_size=(32,),
)

td = TensorDict(
    {"observation": torch.randn(32, 10), "action": torch.randn(32, 5)},
    batch_size=[32],
)
schema.validate(td)

Or using a class-based model:

class RL(pa.TensorDictModel):
    """Schema for reinforcement learning data."""

    # Use PyTorch dtypes in type annotations
    observation: torch.float32 = pa.Field(shape=(None, 10))
    action: torch.int64 = pa.Field(shape=(None,))
    reward: torch.float32 = pa.Field()

    class Config:
        batch_size = (32,)

# Validate using the model - schema is built automatically
td = TensorDict(
    {"observation": torch.randn(32, 10), "action": torch.randint(0, 4, (32,)), "reward": torch.randn(32)},
    batch_size=[32],
)
RL.validate(td)

What's Changed

  • feat(polars): add support for nested DataFrameModel validation by @Nikhil-jaiswal007 in #2439
  • ci: always run linters, regardless of changed paths by @cosmicbboy in #2462
  • docs: add a "Backward Compatibility" section to AGENTS.md by @Dev-iL in #2461
  • fix(api): raise informative TypeError for non-dataframe validate() input by @cognis-digital in #2421
  • Reject reversed bounds in Check.str_length by @chiruu12 in #2476
  • fix(polars): fail loudly on user-declared parsers instead of skipping them by @feiiiiii5 in #2473
  • fix(polars): honour ignore_na=False when the check output is null by @ebarkhordar in #2471
  • fix(config): coerce a string validation_depth instead of silently disabling gating by @shashvat-singham in #2470
  • docs: clarify parser and check execution order (#2045) by @IshanA2007 in #2466
  • Fix documented default for DataFrameSchema.reset_index drop by @VenishPaneliya in #2465
  • Report integer coercion overflow instead of wrapping silently by @shashvat-singham in #2443
  • fix: correct get_metadata return annotations by @Gagandeep-2003 in #2424
  • Fix implicit Optional on Column dtype in the Ibis backend by @VenishPaneliya in #2474
  • fix(polars): pass parser_output on Category.try_coerce's ParserError so lazy validation reports failure cases instead of crashing by @AmirF194 in #2469
  • fix(polars): do not select absent optional columns when adding missin… by @Moandco in #2460
  • fix(polars): add missing columns with string defaults as literals (#2463) by @Moandco in #2464
  • fix(pandas): match tz-aware datetime dtypes that render identically (#1839) by @cycsmail in #2416
  • docs: derived columns + System One (Jev) parsing proposal by @cosmicbboy in #2508
  • fix(io): keep renamed built-in checks in serialized schemas by @feiiiiii5 in #2501
  • Add PyTorch TensorDict validation backend by @cosmicbboy in #2403
  • fix: pass through unresolved TypeVar annotations in check_types by @sneha4175 in #2513
  • fix(pandas): stop KeyError for column dtypes outside the json-schema types by @feiiiiii5 in #2492
  • fix(pandas): propagate a column's own drop_invalid_rows through DataFrameSchema by @AmirF194 in #2521
  • fix(io): keep groupby and groups when serializing checks by @Rodrigo-Palma in #2490
  • fix(narwhals): handle scalar polars check outputs by @vgvr0 in #2542
  • fix(dtypes): reject negative Decimal scale with the documented error by @feiiiiii5 in #2507

New Contributors

Full Changelog: v0.33.1...v0.34.0

Don't miss a new pandera release

NewReleases is sending notifications on new releases.