Compatibility: Python <3.14,>=3.10 → <3.15,>=3.10; numpy added (python_version >= "3.14"); pandas added (python_version >= "3.14"); pydantic added (python_version >= "3.14")
Highlights
-
Python 3.14 support — Great Expectations now installs and runs on Python 3.14, and the supported range is Python 3.10 through 3.14. On Python 3.14 installs resolve numpy 2.3.2 or later, pandas 2.3.3 or later, and pydantic 2.13.0 or later. The
clickhouseandteradataextras are not supported on 3.14, and Python 3.15 and later still requireGX_PYTHON_EXPERIMENTAL. (#12280)pip install great_expectations # now works on Python 3.14
-
New Expectation: ExpectColumnTypeToBe — A schema-level Expectation that checks a column's declared type, separate from the row-level
ExpectColumnValuesToBeOfType. It works on pandas, SQLAlchemy and Spark, returns onlyobserved_valuein its result, and reportssuccess=Falsewithobserved_value=Nonefor a missing column. Note that pandas 3 infers different default dtypes than pandas 2 for string and datetime columns, so re-checktype_when upgrading pandas. (#12186)import great_expectations.expectations as gxe batch.validate(gxe.ExpectColumnTypeToBe(column="order_id", type_="int64"))
-
Escape character for the LIKE pattern Expectations — The four LIKE pattern Expectations accept an optional single-character
escape, which emits anESCAPEclause so_and%can be matched literally and identically across backends. Omitting it produces exactly the SQL generated before. BigQuery and ClickHouse raise a clear error naming the dialect, because neither supports a usable ESCAPE clause; escape wildcards with a backslash inside the pattern there instead. Both renderers now disclose the escape character. (#12154)import great_expectations.expectations as gxe gxe.ExpectColumnValuesToMatchLikePattern(column="c", like_pattern="a!_b", escape="!")
-
Objects now resolve through the Data Context that owns them — With more than one Data Context in a process, expectation suites, validation definitions, checkpoints, batch definitions, the validator and a checkpoint's actions now read, write and run against the context they came from rather than whichever context happens to be current. Data Docs are built into the owning context's site,
${VARIABLE}references in Slack, Teams and Email settings resolve from the owning context's configuration, and a "not found" error for an object that belongs to no context now names the context the lookup consulted. An object that belongs to no context still resolves through the current one. (#12297, #12298, #12302, #12303)c1 = gx.get_context(mode="ephemeral") suite = c1.suites.add(gx.ExpectationSuite(name="my_suite")) c2 = gx.get_context(mode="ephemeral") # c2 is now current suite.add_expectation(...) # still lands in c1
-
Database URLs are masked without being mangled —
PasswordMasker.mask_db_urlkeeps every part of a database URL except the secrets: query strings, fragments, IPv6 brackets and empty authorities survive, absent usernames and hosts are no longer printed asNone, and passwords containing/,?or#no longer raise. Credentials passed as query parameters (password=,odbc_connect=,private_key_file_pwd=) are now masked on both code paths, with an allowlist of safe-to-display parameter names, and any value that is not a<scheme>://...URL is masked whole. (#12242) -
Column names with spaces, hyphens, dots or keywords work in row conditions — On pandas and on Spark, a
row_conditionreferring to a column whose name is not a bare identifier is now escaped before it reaches the query engine, so names such as"Total Amount","a.b","x-y","from"and"Incident Number"filter the rows they name instead of raising a syntax error or silently validating the wrong rows. (#12238, #12285)gxe.ExpectColumnValuesToNotBeNull( column="status", row_condition=Column("Total Amount") > 15, )
Changes
Features
- The four LIKE pattern Expectations accept an optional single-character
escapethat emits anESCAPEclause so_and%can be matched literally and portably; BigQuery and ClickHouse raise a clear error naming the dialect, and omittingescapeproduces exactly the SQL generated before. (#12154) - Added
ExpectColumnTypeToBe, a batch-level Expectation that checks a column's declared type on pandas, SQLAlchemy and Spark, reporting onlyobserved_valueand returningsuccess=Falsewithobserved_value=Nonefor a missing column. (#12186) - Added support for Python 3.14: the supported range is now 3.10 to 3.14 with dependency floors for numpy, pandas and pydantic on 3.14, and four comparison fixes also change behavior on 3.10 to 3.13 —
CaseInsensitiveStringno longer compares equal to arbitrary non-string objects,ExceptionInfoinequality defers to the other operand,ExpectationConfiguration.isEquivalentToreturnsFalsefor non-configurations, andExpectationValidationResultequality no longer raises when its expectation config is a plain dict. Theclickhouseandteradataextras are not supported on 3.14. (#12280)
Bug fixes
- On Spark, column names that are not bare SQL identifiers are quoted in generated filter strings, so a hyphenated name such as
id-nno longer silently validates the wrong rows and a spaced name such asIncident Numberno longer raises a parse error. (#12285) - A checkpoint's actions now run against the Data Context that owns the checkpoint: Data Docs are built into that context's site, Data Docs URLs are looked up there, and
${VARIABLE}references in Slack, Teams and Email settings resolve from its configuration. An action run directly, outside a checkpoint, still resolves through the current context. (#12303) - Validation definitions, checkpoints, batch definitions and the validator now run, save and check freshness against the Data Context that owns them, so with two contexts in one process an object no longer reports "has changed since it has last been saved", writes results to the wrong project, or validates another context's data source of the same name. (#12302)
- Masked database URLs keep their query string, fragment, IPv6 brackets and empty authority, no longer invent
Noneusernames or hosts, no longer raise on passwords containing/,?or#, and now mask credentials passed as query parameters on both masking paths; values that are not URLs are masked whole. (#12242) - An Expectation Suite obtained from a Data Context now reads and writes through that context, so
add_expectation,delete_expectationandsavechange the context the suite came from even while another is current, andsuites.add_or_updateof an existing name writes to the context whose factory was called. (#12298) - A "not found" error for an expectation suite, validation definition, checkpoint or batch definition that belongs to no Data Context now names the context the lookup consulted, and the batch definition miss message no longer nests a whole sentence inside the name. (#12297)
- SQLite engines that GX creates now close their pooled database connections once the engine is collected, so on Python 3.13 and later users no longer see
ResourceWarning: unclosed databasefrom a GX SQL datasource or execution engine going out of scope. (#12301) - A
row_conditionmay now contain single quotes, so conditions such ascol == 'value'are accepted instead of raising an invalid-configuration error; the newline restriction is unchanged. (#12258) ExpectColumnValuesToBeInSeton SQL backends ignoresNoneentries invalue_set, so values outside the set are flagged as unexpected instead of every row passing, matching pandas behavior; the unexpected-index query no longer renders a NULL literal. (#12277)ExpectColumnValuesToBeIncreasingandExpectColumnValuesToBeDecreasingon pandas now compare a datetime column's differences against a zero of the matching dtype, so a correctly ordered datetime column succeeds instead of reporting every row as unexpected. (#12270)column.standard_deviationandExpectColumnStdevToBeBetweennow work on a SQLite database added withadd_sql, returning the sample standard deviation instead of raisingno such function: stddev_samp;add_sqlitebehavior is unchanged. (#12212)- The
unexpected_index_queryproduced for Spark results is now a valid Python expression with the condition rendered as a quoted SQL string, so it parses and evaluates, including for in-set and regex conditions; where Spark cannot resolve the condition the quoted display string is used as a fallback. (#12220) - On pandas, a
row_conditionreferring to a column whose name is not a valid Python identifier — a space, dot or hyphen in the name, or a Python keyword — is now escaped with backticks, so the condition filters the intended rows instead of producing a syntax error reported as a failed validation. (#12238) - A query-metric or
UnexpectedRowsExpectationquery containing a literal brace — a JSON or array literal, a regex quantifier, or a LIKE pattern — now reports an error that quotes the query and explains that braces must be doubled, instead of a bareKeyErrororValueErrorfrom string formatting. (#12236)
Maintenance
- Bumped the docs site's proxy-addr dependency from 2.0.7 to 2.0.8. (#12296)
- Bumped the docs site's source-map-js dependency from 1.2.1 to 1.2.2. (#12295)
- Bumped the docs site's joi dependency from 17.13.7 to 17.13.8. (#12294)
- Bumped the docs site's compression dependency from 1.8.1 to 1.8.2. (#12293)
- Bumped the docs site's dompurify dependency from 3.4.13 to 3.4.16. (#12284)
- Bumped the docs site's fast-uri dependency from 3.1.7 to 3.1.8. (#12282)
- Bumped the docs site's brace-expansion dependency from 1.1.18 to 1.1.21. (#12281)
- Brought the
tests/datasourcesuite under the type checker by removing its mypy exclusions and fixing the reported type errors. (#12259) - Brought the Data Docs and Slack renderer modules under the type checker and fixed the defects it exposed: a content-block renderer now renders no content for a failed or unrenderable validation result instead of raising, the Data Docs index skips a result the context cannot find with its existing warning, and a Data Docs site configured with
cloud_modeorge_cloud_modenow raises the GX Cloud shutdown error. (#12256) - Added explicit
@overridedecorators to the expectation-core and renderer methods mypy's explicit-override check reported, with no behavioral change. (#12193) - CI now defines the supported Python versions once and derives every job's interpreter matrix from that single list, with no change to which interpreters run. (#12278)
- The CLA status check retries a transient failure from the contributor-handle endpoint up to five times before giving up, so a signed contributor's pull request no longer loses its CLA gate to an intermittent blip; the check still fails closed when every attempt fails. (#12275)
Contributors
Thanks to @DhanushPillay (first contribution), @feiiiiii5, @iamdpsingh (first contribution), @ptimizeroracle, @Rodrigo-Palma (first contribution), @kamalesh404 (first contribution), @umanggarg754 (first contribution), @Star-cloud626, @krisnaparahita (first contribution), @LucasLeao18 (first contribution).