github dlt-hub/dlt 1.31.0

4 hours ago

dlt 1.31.0 Release Notes

Breaking Changes

  1. pendulum helpers removed from dlt.common.time; pendulum>=3 required (#4033 @rudolfix)
    • What changed: the minimum pendulum is now 3.0.0 on all Python versions. dlt still works with pendulum as before. documented usage and
      pendulum helpers work. Internally pendulum is not used.

      Unofficial helpers below were removed, with no compatibility aliases.

      Removed Replacement
      ensure_pendulum_datetime_utc(v) ensure_datetime_in_tz(v, timezone.utc)
      ensure_pendulum_datetime_non_utc(v) ensure_datetime(v)
      ensure_datetime_utc(v) ensure_datetime_in_tz(v, timezone.utc)
      PY_DATETIME_DECODERS DECODERS
    • No more forced UTC: ensure_pendulum_datetime no longer forces UTC. It uses the context timezone, which is UTC unless you set one.

  2. Filesystem timestamps are stdlib datetime (#4388 @rudolfix): modification_date and other timestamps from the fsspec layer and the filesystem source are now stdlib datetime objects (tz-aware UTC). Code that calls pendulum-only methods on them breaks.
  3. Datetime cursor values follow the declared cursor type (#4406 @rudolfix)
    • start_value, last_value and get_current_range() return pendulum.DateTime only for pendulum cursors. datetime.datetime cursors now get stdlib datetimes.
    • The lag unit now comes from incremental[date] / [datetime] or from the type of initial_value. It is no longer guessed from the string format of the values, which used to turn lag=3 into 3 seconds for ISO datetime strings.
  4. JSON datetimes serialize as +00:00 instead of Z (#4033 @rudolfix): applies to jsonl loader files and json columns, the same with every JSON backend. Old state and packages still load.
  5. Incremental.last_value is the current cursor position (#4033 @rudolfix)
    Note: this is enforcement of documented usage. last_value is the actual cursor value at the time of the call and now it works like that in every case.
    • Before the first batch it equals start_value (no change!); after that it is the running value.
    • lag no longer applies to it, and it keeps advancing when end_value is set.
    • This affects rest_api {incremental.last_value} placeholders on dependent resources.
  6. Snowflake: tz-aware timestamps default to TIMESTAMP_LTZ (#4397 @rudolfix). Set use_timestamp_tz=True to keep TIMESTAMP_TZ.
  7. Delta upsert honors hard_delete (#4305 @adrienModeo)
    • Rows flagged as deleted are now deleted, and never inserted.
    • With nested tables, this combination raises SchemaCorruptedException before the load.
  8. timestamp_timezone on the parquet writer is deprecated as a zone selector (#4033 @rudolfix): use the context timezone instead. "" still means "no UTC adjustment".

Highlights

  • cdc merge strategy (#4305 @adrienModeo)
    • Loads a full snapshot as an upsert and deletes destination rows that are missing from it.
    • Optionally deletes only within the loaded merge_key partitions.
    • Supported on duckdb (1.4+), motherduck, ducklake, Snowflake, Postgres, BigQuery, Databricks, MSSQL, Fabric, Athena (Iceberg) and Delta.
  • Skip unchanged rows on cdc and upsert (#4305 @adrienModeo)
    • skip_unchanged_rows: True updates only rows that changed, so unchanged rows keep their _dlt_load_id. Snowflake Streams and similar consumers then see only real changes.
    • row_version_column_name compares one version or hash column instead of all columns.
  • Merge conditions: source_filter and destination_scope (#4305 @adrienModeo)
    • source_filter chooses which loaded rows to merge. destination_scope chooses which destination rows can be deleted or retired.
    • They work with delete-insert, scd2 and cdc (upsert takes source_filter only).
    • Together they enable:
      • partition replace (BigQuery and Snowflake can prune partitions);
      • keyless delete-insert;
      • scd2 retiring only within a scope;
      • {table} / {staging_table} placeholders.
  • Stateful Relation.incremental() (#4033 @rudolfix)
    • Incremental reads from a dataset now push the range down to SQL and keep state.
    • The range end is locked to MAX(cursor), so runs cover the cursor range without gaps or duplicates.
    • Supports date cursors, range_start / range_end overrides, and JSONPath or qualified cursors.
  • Context timezone (#4033 @rudolfix)
    • A single process-wide timezone (UTC by default) that decides how naive and aware timestamps are stored for the timezone column hint, on both the object and arrow paths.
    • Set it with TimezoneContext, or with require.timezone on a job.

Core Library

  • cdc merge strategy, skip_unchanged_rows, source_filter / destination_scope: see Highlights (#4305 @adrienModeo)
  • Hard deletes from change feeds (#4305 @adrienModeo): a text hard_delete column treats any non-NULL value, such as "D", as deleted. Merge options are validated when the resource is defined and again before the load.
  • Stateful Relation.incremental() and context timezone: see Highlights (#4033 @rudolfix)
  • New Incremental helpers (#4033 @rudolfix):
    • with_cursor() copies an incremental with a different cursor;
    • get_current_range(apply_lag=True) works on bound and unbound instances;
    • advance() pins the range end;
    • when the primary key is the cursor, boundary rows are loaded eagerly and never replayed. A warning appears when the boundary row is deferred.
  • Cron and interval helpers in core dlt.common.interval (#4033 @rudolfix): TTimeInterval is now a NamedTuple, so .start and .end work alongside tuple unpacking. croniter is now a core dependency.
  • Snowflake Workload Identity Federation (#4362 @richacode007-byte): new workload_identity_provider credential field (AWS, AZURE, GCP, OIDC). Needs snowflake-connector-python>=3.17.0.
  • Globbing over http/https fsspec (#4388 @rudolfix): http(s) listings can be globbed. A missing size or modified date can be fetched with a HEAD request per file.
  • ClickHouse staging-optimized replace on Replicated / Shared database engines (#4392 @rudolfix)
  • Descriptions for variant columns created by dlt (#3713 @aditypan)
  • Table and column descriptions for internal _dlt tables (#3719 @aditypan)
  • Load a module from a folder as a private package (#4417 @rudolfix): import_folder_module in dlt.common.reflection.ref, which does not touch sys.path.
  • Fix: _dlt_loads.inserted_at and _dlt_version.inserted_at are written in UTC (#4033 @rudolfix): they were written in the loading machine's local time.
  • Fix: ClickHouse datetime literals use IANA zone names (#4033 @rudolfix): fixed-offset values used to produce zones like UTC+02:00, which ClickHouse rejects.
  • Fix: merge with hard_delete and no keys raised UnboundLocalError on nested tables (#4305 @adrienModeo): this also left the package partially loaded. Fixed in sql_jobs.py and the sqlalchemy merge job.
  • Fix: incremental lag unit follows the declared cursor type (#4406 @rudolfix)
  • Fix: incremental last_value stays forward-only when the cursor is exactly 0 (#4367 @chjnett)
  • Fix: bind_query generated wrong aliases on case-folding destinations (#4359 @rudolfix)
  • Fix: raise ValueError when coercing a non-finite decimal to bigint (#4348 @eeshsaxena)
  • Fix: PipelineTrace last-step accessors typed as Optional (#4344 @bunnysayzz)
  • Fix: Pipeline.last_trace typed as possibly None (#4469 @VioletM)
  • Fix: filesystem destination works on PyPy and pyodide (#4385 @tushardev-365): the hard orjson import was removed, and the simplejson backend gets the proper error. Adds a lint-emscripten import check.
  • Fix: read_csv_duckdb no longer uses the global duckdb instance (#4434 @jtcurlin)
  • Fix: BigQuery table descriptions are escaped in generated SQL (#4473 @rooperuu)
  • Fix: Fabric nvarchar precision scaled for UTF-8 byte semantics (#4259 @sdebruyn)
  • Fix: Fabric accepts time columns on the parquet load path (#4260 @sdebruyn)
  • Fix: DeletingResourcesNotSupported error message repaired (#4438 @simpleqt)
  • Fix: CLI errors no longer end with a generic "refer to our docs" note (#4136 @anxkhn)

Docs

Chores

  • Docs build tooling upgrade (#4420 @zilto): ruff 0.16, prek and ty; removes docs_tools and the old docs workflows.
  • mdsmith formatting rules for docs (#4450 @zilto)
  • Lint CI pinned to Python 3.10 (#4350 @zilto)
  • Remove the dead test_examples.yml workflow (@zilto, direct commit 23ecc8fee)
  • Release-highlights trigger workflow (#4419 @ShreyasGS)
  • agentic-docs: pass the Actions run URL to the job (#4364 @ShreyasGS)
  • Tests no longer write to ~/.dlt (#4368 @arose26)
  • Bump to 1.31.0 (#4518 @rudolfix)

dltHub Runtime, Jobs, Deployments and dlthub transformations

Breaking Changes

  1. @job / @pipeline_run: allow_external_schedulers and refresh deprecated (#4033 @rudolfix)
    • Use incremental_mode="interval" | "pipeline" and refresh_propagation instead.
    • The old keyword arguments still work but emit a DltDeprecationWarning. They are no longer in the typed signatures, so mypy reports errors for typed callers.
  2. TimeIntervalContext(allow_external_schedulers=False) is no longer a kill switch (#4033 @rudolfix)
  3. Deployment manifest engine v2 (#4033, #4417 @rudolfix)
    • v1 manifests migrate automatically when loaded.
    • Every job definition now has an engine_version. Jobs that take arguments now carry inputs, so every workspace's manifest hash changes once. Expect a one-time diff on the next deploy.
  4. MCP tools that declare no RequiresAccess are treated as needing full access (#4417 @rudolfix): plugin tools without the annotation are hidden from callers with a narrower grant, such as agent jobs. Interactive clients are not affected.
  5. dlthub ai toolkit install --overwrite removes stale files (#4417 @rudolfix)
    • Files from a previous install that the new toolkit version no longer ships are deleted. Installed files are tracked, with their hashes, in .dlt/.toolkits.
    • An existing full clone of the toolkit repo is re-cloned once as a shallow, sparse checkout.
  6. Relation.incremental() advances pipeline state by default (#4033 @rudolfix)
    • Pass advance=False to get the old stateless filter.
    • Calling .incremental() twice now combines both filters with AND instead of raising.
  7. Incremental.allow_external_schedulers defaults to None (#4033 @rudolfix): an explicit value on the incremental always wins, and a context-level False no longer force-disables it.

Highlights

  • Background agent jobs (#4417 @rudolfix)
    • What it is: a new job kind whose body is an agent loop, defined in an AGENT.md or with a decorated function.
    • What a definition holds: a system prompt, typed inputs and output (JSON Schema), MCP tool groups, skills, rules and an access grant.
    • How it runs: declared with run.agent(...), run locally with dlthub local run or on the runner.
    • Results: a structured result with status, summary, result and a full trace.
  • Pluggable agent loops (#4417 @rudolfix): pydantic-ai is the default and claude-agent-sdk is the alternative. Third-party loops register through the plug_agent_loop hook. Each loop has its own dependency group, installed by the runner.
  • Access model for jobs (#4417 @rudolfix)
    • access declares local (read/write/execute/network), data (read/write) and context (read).
    • local maps to a standard tool set (Read/Glob/Grep, Write/Edit, Bash, WebFetch/WebSearch) that is limited to the workspace and the temp folder.
    • data and context are enforced by the workspace MCP server: a tool the grant doesn't cover is never shown to the model.
    • Credential files (*secrets.toml, .env) are never readable by file tools.
  • Agent jobs are callable and testable like Python functions (#4417 @rudolfix)
    • You can call or await an agent job directly.
    • job.last_job_result exposes the trace, token counts and entities.
    • agent.py next to AGENT.md provides validate_input / validate_output hooks; JobAbortedException aborts a run cleanly.
  • Job-level incremental and refresh control (#4033 @rudolfix)
    • incremental_mode, refresh_propagation and auto_refresh_pipeline_mode, plus a [jobs] config section.
    • dlt.current.interval can be changed from inside a job (set, update, apply_lag, apply_full_days).
    • require.timezone runs a job in its declared timezone.

Core Library

  • Background agent jobs, loops and the access model: see Highlights (#4417 @rudolfix)
  • Structured results for every job (#4417 @rudolfix)
    • run.result(...) declares a typed result.
    • The launcher delivers an envelope shaped like Activity Streams 2.0 (type, job_ref, object, result) to telemetry as a job_result.
    • inputs / output JSON Schema is available on every job kind.
  • Entity-typed inputs (#4417 @rudolfix)
    • Inputs can be marked as references to platform entities, e.g. Annotated[str, Entity("job-runs")] or entity_type: in AGENT.md. Supported types: job-runs, job, pipeline, dataset, workspace.
    • This fills the result's object and expose.object_input, so the UI can offer an agent on a failed run.
  • job.fail: / job.success: triggers accept selectors (#4417 @rudolfix): selectors expand at manifest time to every matching job. They never target the declaring job or interactive jobs.
  • Agent settings through config (#4417 @rudolfix)
    • agent.model, agent.instructions, agent.verbosity, agent.max_turns and agent.max_tokens can be set per run with -c, env or toml.
    • The model endpoint is one set: agent.api_key / api_url / api_version, with runtime_* twins supplied by the runtime. Azure, LiteLLM and xAI are mapped automatically.
    • Token limits are counted by dlt, so they behave the same on both loops.
  • Workspace MCP server: --features, --no-default-features, --access (#4417 @rudolfix): starts a server that exposes only the requested tool groups and only the tools the grant covers. Tools declare their needs with RequiresAccess.
  • Toolkits ship agents (#4417 @rudolfix)
    • Agents install under .claude/dlthub/agents/<toolkit>/<name>/ (or .cursor/, .agents/), separate from native subagents.
    • The .dlt/.toolkits index ships with deployments, so <toolkit>:<agent> refs resolve on the runner.
  • Scheduler interval reaches every launcher (#4033 @rudolfix)
    • The interval is injected in-process for job, and passed as DLT_INTERVAL_* env vars to the module, streamlit, marimo, mcp and dashboard launchers.
    • Profile, interval and refresh env are set before the user module is imported, so pipelines created at import time see refresh runs.
  • Fix: MCP execute_sql_query read-only check hardened (#4417 @rudolfix): it now accepts exactly one statement, and the whole statement is checked for mutating statements and host functions.
  • Fix: deployment file selection no longer follows symlinks (#4287 @Sanjays2402): fixes a RecursionError on symlink loops.
  • Fix: dlthub ai status warning no longer mentions MCP (#4445 @AstrakhantsevaAA)

Docs

Chores

  • Disable anonymous telemetry in the platform connection test (#4347 @tetelio): avoids a fork segfault on macOS.

New Contributors

Don't miss a new dlt release

NewReleases is sending notifications on new releases.