github fivetran/great_expectations 1.24.0

6 hours ago

Compatibility: Python <3.14,>=3.10 → <3.15,>=3.10; numpy added (python_version >= "3.14"); pandas added (python_version >= "3.14"); pydantic added (python_version >= "3.14")

Highlights

  • Python 3.14 support — Great Expectations now installs and runs on Python 3.14, and the supported range is Python 3.10 through 3.14. On Python 3.14 installs resolve numpy 2.3.2 or later, pandas 2.3.3 or later, and pydantic 2.13.0 or later. The clickhouse and teradata extras are not supported on 3.14, and Python 3.15 and later still require GX_PYTHON_EXPERIMENTAL. (#12280)

    pip install great_expectations  # now works on Python 3.14
  • New Expectation: ExpectColumnTypeToBe — A schema-level Expectation that checks a column's declared type, separate from the row-level ExpectColumnValuesToBeOfType. It works on pandas, SQLAlchemy and Spark, returns only observed_value in its result, and reports success=False with observed_value=None for a missing column. Note that pandas 3 infers different default dtypes than pandas 2 for string and datetime columns, so re-check type_ when upgrading pandas. (#12186)

    import great_expectations.expectations as gxe
    
    batch.validate(gxe.ExpectColumnTypeToBe(column="order_id", type_="int64"))
  • Escape character for the LIKE pattern Expectations — The four LIKE pattern Expectations accept an optional single-character escape, which emits an ESCAPE clause so _ and % can be matched literally and identically across backends. Omitting it produces exactly the SQL generated before. BigQuery and ClickHouse raise a clear error naming the dialect, because neither supports a usable ESCAPE clause; escape wildcards with a backslash inside the pattern there instead. Both renderers now disclose the escape character. (#12154)

    import great_expectations.expectations as gxe
    
    gxe.ExpectColumnValuesToMatchLikePattern(column="c", like_pattern="a!_b", escape="!")
  • Objects now resolve through the Data Context that owns them — With more than one Data Context in a process, expectation suites, validation definitions, checkpoints, batch definitions, the validator and a checkpoint's actions now read, write and run against the context they came from rather than whichever context happens to be current. Data Docs are built into the owning context's site, ${VARIABLE} references in Slack, Teams and Email settings resolve from the owning context's configuration, and a "not found" error for an object that belongs to no context now names the context the lookup consulted. An object that belongs to no context still resolves through the current one. (#12297, #12298, #12302, #12303)

    c1 = gx.get_context(mode="ephemeral")
    suite = c1.suites.add(gx.ExpectationSuite(name="my_suite"))
    c2 = gx.get_context(mode="ephemeral")  # c2 is now current
    suite.add_expectation(...)  # still lands in c1
  • Database URLs are masked without being mangled — PasswordMasker.mask_db_url keeps every part of a database URL except the secrets: query strings, fragments, IPv6 brackets and empty authorities survive, absent usernames and hosts are no longer printed as None, and passwords containing /, ? or # no longer raise. Credentials passed as query parameters (password=, odbc_connect=, private_key_file_pwd=) are now masked on both code paths, with an allowlist of safe-to-display parameter names, and any value that is not a <scheme>://... URL is masked whole. (#12242)

  • Column names with spaces, hyphens, dots or keywords work in row conditions — On pandas and on Spark, a row_condition referring to a column whose name is not a bare identifier is now escaped before it reaches the query engine, so names such as "Total Amount", "a.b", "x-y", "from" and "Incident Number" filter the rows they name instead of raising a syntax error or silently validating the wrong rows. (#12238, #12285)

    gxe.ExpectColumnValuesToNotBeNull(
        column="status",
        row_condition=Column("Total Amount") > 15,
    )

Changes

Features
  • The four LIKE pattern Expectations accept an optional single-character escape that emits an ESCAPE clause so _ and % can be matched literally and portably; BigQuery and ClickHouse raise a clear error naming the dialect, and omitting escape produces exactly the SQL generated before. (#12154)
  • Added ExpectColumnTypeToBe, a batch-level Expectation that checks a column's declared type on pandas, SQLAlchemy and Spark, reporting only observed_value and returning success=False with observed_value=None for a missing column. (#12186)
  • Added support for Python 3.14: the supported range is now 3.10 to 3.14 with dependency floors for numpy, pandas and pydantic on 3.14, and four comparison fixes also change behavior on 3.10 to 3.13 — CaseInsensitiveString no longer compares equal to arbitrary non-string objects, ExceptionInfo inequality defers to the other operand, ExpectationConfiguration.isEquivalentTo returns False for non-configurations, and ExpectationValidationResult equality no longer raises when its expectation config is a plain dict. The clickhouse and teradata extras are not supported on 3.14. (#12280)
Bug fixes
  • On Spark, column names that are not bare SQL identifiers are quoted in generated filter strings, so a hyphenated name such as id-n no longer silently validates the wrong rows and a spaced name such as Incident Number no longer raises a parse error. (#12285)
  • A checkpoint's actions now run against the Data Context that owns the checkpoint: Data Docs are built into that context's site, Data Docs URLs are looked up there, and ${VARIABLE} references in Slack, Teams and Email settings resolve from its configuration. An action run directly, outside a checkpoint, still resolves through the current context. (#12303)
  • Validation definitions, checkpoints, batch definitions and the validator now run, save and check freshness against the Data Context that owns them, so with two contexts in one process an object no longer reports "has changed since it has last been saved", writes results to the wrong project, or validates another context's data source of the same name. (#12302)
  • Masked database URLs keep their query string, fragment, IPv6 brackets and empty authority, no longer invent None usernames or hosts, no longer raise on passwords containing /, ? or #, and now mask credentials passed as query parameters on both masking paths; values that are not URLs are masked whole. (#12242)
  • An Expectation Suite obtained from a Data Context now reads and writes through that context, so add_expectation, delete_expectation and save change the context the suite came from even while another is current, and suites.add_or_update of an existing name writes to the context whose factory was called. (#12298)
  • A "not found" error for an expectation suite, validation definition, checkpoint or batch definition that belongs to no Data Context now names the context the lookup consulted, and the batch definition miss message no longer nests a whole sentence inside the name. (#12297)
  • SQLite engines that GX creates now close their pooled database connections once the engine is collected, so on Python 3.13 and later users no longer see ResourceWarning: unclosed database from a GX SQL datasource or execution engine going out of scope. (#12301)
  • A row_condition may now contain single quotes, so conditions such as col == 'value' are accepted instead of raising an invalid-configuration error; the newline restriction is unchanged. (#12258)
  • ExpectColumnValuesToBeInSet on SQL backends ignores None entries in value_set, so values outside the set are flagged as unexpected instead of every row passing, matching pandas behavior; the unexpected-index query no longer renders a NULL literal. (#12277)
  • ExpectColumnValuesToBeIncreasing and ExpectColumnValuesToBeDecreasing on pandas now compare a datetime column's differences against a zero of the matching dtype, so a correctly ordered datetime column succeeds instead of reporting every row as unexpected. (#12270)
  • column.standard_deviation and ExpectColumnStdevToBeBetween now work on a SQLite database added with add_sql, returning the sample standard deviation instead of raising no such function: stddev_samp; add_sqlite behavior is unchanged. (#12212)
  • The unexpected_index_query produced for Spark results is now a valid Python expression with the condition rendered as a quoted SQL string, so it parses and evaluates, including for in-set and regex conditions; where Spark cannot resolve the condition the quoted display string is used as a fallback. (#12220)
  • On pandas, a row_condition referring to a column whose name is not a valid Python identifier — a space, dot or hyphen in the name, or a Python keyword — is now escaped with backticks, so the condition filters the intended rows instead of producing a syntax error reported as a failed validation. (#12238)
  • A query-metric or UnexpectedRowsExpectation query containing a literal brace — a JSON or array literal, a regex quantifier, or a LIKE pattern — now reports an error that quotes the query and explains that braces must be doubled, instead of a bare KeyError or ValueError from string formatting. (#12236)
Maintenance
  • Bumped the docs site's proxy-addr dependency from 2.0.7 to 2.0.8. (#12296)
  • Bumped the docs site's source-map-js dependency from 1.2.1 to 1.2.2. (#12295)
  • Bumped the docs site's joi dependency from 17.13.7 to 17.13.8. (#12294)
  • Bumped the docs site's compression dependency from 1.8.1 to 1.8.2. (#12293)
  • Bumped the docs site's dompurify dependency from 3.4.13 to 3.4.16. (#12284)
  • Bumped the docs site's fast-uri dependency from 3.1.7 to 3.1.8. (#12282)
  • Bumped the docs site's brace-expansion dependency from 1.1.18 to 1.1.21. (#12281)
  • Brought the tests/datasource suite under the type checker by removing its mypy exclusions and fixing the reported type errors. (#12259)
  • Brought the Data Docs and Slack renderer modules under the type checker and fixed the defects it exposed: a content-block renderer now renders no content for a failed or unrenderable validation result instead of raising, the Data Docs index skips a result the context cannot find with its existing warning, and a Data Docs site configured with cloud_mode or ge_cloud_mode now raises the GX Cloud shutdown error. (#12256)
  • Added explicit @override decorators to the expectation-core and renderer methods mypy's explicit-override check reported, with no behavioral change. (#12193)
  • CI now defines the supported Python versions once and derives every job's interpreter matrix from that single list, with no change to which interpreters run. (#12278)
  • The CLA status check retries a transient failure from the contributor-handle endpoint up to five times before giving up, so a signed contributor's pull request no longer loses its CLA gate to an intermittent blip; the check still fails closed when every attempt fails. (#12275)

Contributors

Thanks to @DhanushPillay (first contribution), @feiiiiii5, @iamdpsingh (first contribution), @ptimizeroracle, @Rodrigo-Palma (first contribution), @kamalesh404 (first contribution), @umanggarg754 (first contribution), @Star-cloud626, @krisnaparahita (first contribution), @LucasLeao18 (first contribution).

Don't miss a new great_expectations release

NewReleases is sending notifications on new releases.