dlt 1.31.0 Release Notes
Breaking Changes
- pendulum helpers removed from
dlt.common.time;pendulum>=3required (#4033 @rudolfix)-
What changed: the minimum
pendulumis now 3.0.0 on all Python versions.dltstill works withpendulumas before. documented usage and
pendulumhelpers work. Internallypendulumis not used.Unofficial helpers below were removed, with no compatibility aliases.
Removed Replacement ensure_pendulum_datetime_utc(v)ensure_datetime_in_tz(v, timezone.utc)ensure_pendulum_datetime_non_utc(v)ensure_datetime(v)ensure_datetime_utc(v)ensure_datetime_in_tz(v, timezone.utc)PY_DATETIME_DECODERSDECODERS -
No more forced UTC:
ensure_pendulum_datetimeno longer forces UTC. It uses the context timezone, which is UTC unless you set one.
-
- Filesystem timestamps are stdlib
datetime(#4388 @rudolfix):modification_dateand other timestamps from the fsspec layer and thefilesystemsource are now stdlibdatetimeobjects (tz-aware UTC). Code that calls pendulum-only methods on them breaks. - Datetime cursor values follow the declared cursor type (#4406 @rudolfix)
start_value,last_valueandget_current_range()returnpendulum.DateTimeonly for pendulum cursors.datetime.datetimecursors now get stdlib datetimes.- The
lagunit now comes fromincremental[date]/[datetime]or from the type ofinitial_value. It is no longer guessed from the string format of the values, which used to turnlag=3into 3 seconds for ISO datetime strings.
- JSON datetimes serialize as
+00:00instead ofZ(#4033 @rudolfix): applies to jsonl loader files and json columns, the same with every JSON backend. Old state and packages still load. Incremental.last_valueis the current cursor position (#4033 @rudolfix)
Note: this is enforcement of documented usage.last_valueis the actual cursor value at the time of the call and now it works like that in every case.- Before the first batch it equals
start_value(no change!); after that it is the running value. lagno longer applies to it, and it keeps advancing whenend_valueis set.- This affects rest_api
{incremental.last_value}placeholders on dependent resources.
- Before the first batch it equals
- Snowflake: tz-aware timestamps default to
TIMESTAMP_LTZ(#4397 @rudolfix). Setuse_timestamp_tz=Trueto keepTIMESTAMP_TZ. - Delta
upserthonorshard_delete(#4305 @adrienModeo)- Rows flagged as deleted are now deleted, and never inserted.
- With nested tables, this combination raises
SchemaCorruptedExceptionbefore the load.
timestamp_timezoneon the parquet writer is deprecated as a zone selector (#4033 @rudolfix): use the context timezone instead.""still means "no UTC adjustment".
Highlights
cdcmerge strategy (#4305 @adrienModeo)- Loads a full snapshot as an upsert and deletes destination rows that are missing from it.
- Optionally deletes only within the loaded
merge_keypartitions. - Supported on duckdb (1.4+), motherduck, ducklake, Snowflake, Postgres, BigQuery, Databricks, MSSQL, Fabric, Athena (Iceberg) and Delta.
- Skip unchanged rows on
cdcandupsert(#4305 @adrienModeo)skip_unchanged_rows: Trueupdates only rows that changed, so unchanged rows keep their_dlt_load_id. Snowflake Streams and similar consumers then see only real changes.row_version_column_namecompares one version or hash column instead of all columns.
- Merge conditions:
source_filteranddestination_scope(#4305 @adrienModeo)source_filterchooses which loaded rows to merge.destination_scopechooses which destination rows can be deleted or retired.- They work with
delete-insert,scd2andcdc(upserttakessource_filteronly). - Together they enable:
- partition replace (BigQuery and Snowflake can prune partitions);
- keyless
delete-insert; scd2retiring only within a scope;{table}/{staging_table}placeholders.
- Stateful
Relation.incremental()(#4033 @rudolfix)- Incremental reads from a dataset now push the range down to SQL and keep state.
- The range end is locked to
MAX(cursor), so runs cover the cursor range without gaps or duplicates. - Supports date cursors,
range_start/range_endoverrides, and JSONPath or qualified cursors.
- Context timezone (#4033 @rudolfix)
- A single process-wide timezone (UTC by default) that decides how naive and aware timestamps are stored for the
timezonecolumn hint, on both the object and arrow paths. - Set it with
TimezoneContext, or withrequire.timezoneon a job.
- A single process-wide timezone (UTC by default) that decides how naive and aware timestamps are stored for the
Core Library
cdcmerge strategy,skip_unchanged_rows,source_filter/destination_scope: see Highlights (#4305 @adrienModeo)- Hard deletes from change feeds (#4305 @adrienModeo): a text
hard_deletecolumn treats any non-NULL value, such as"D", as deleted. Merge options are validated when the resource is defined and again before the load. - Stateful
Relation.incremental()and context timezone: see Highlights (#4033 @rudolfix) - New
Incrementalhelpers (#4033 @rudolfix):with_cursor()copies an incremental with a different cursor;get_current_range(apply_lag=True)works on bound and unbound instances;advance()pins the range end;- when the primary key is the cursor, boundary rows are loaded eagerly and never replayed. A warning appears when the boundary row is deferred.
- Cron and interval helpers in core
dlt.common.interval(#4033 @rudolfix):TTimeIntervalis now aNamedTuple, so.startand.endwork alongside tuple unpacking.croniteris now a core dependency. - Snowflake Workload Identity Federation (#4362 @richacode007-byte): new
workload_identity_providercredential field (AWS, AZURE, GCP, OIDC). Needssnowflake-connector-python>=3.17.0. - Globbing over http/https fsspec (#4388 @rudolfix): http(s) listings can be globbed. A missing size or modified date can be fetched with a HEAD request per file.
- ClickHouse
staging-optimizedreplace onReplicated/Shareddatabase engines (#4392 @rudolfix) - Descriptions for variant columns created by dlt (#3713 @aditypan)
- Table and column descriptions for internal
_dlttables (#3719 @aditypan) - Load a module from a folder as a private package (#4417 @rudolfix):
import_folder_moduleindlt.common.reflection.ref, which does not touchsys.path. - Fix:
_dlt_loads.inserted_atand_dlt_version.inserted_atare written in UTC (#4033 @rudolfix): they were written in the loading machine's local time. - Fix: ClickHouse datetime literals use IANA zone names (#4033 @rudolfix): fixed-offset values used to produce zones like
UTC+02:00, which ClickHouse rejects. - Fix:
mergewithhard_deleteand no keys raisedUnboundLocalErroron nested tables (#4305 @adrienModeo): this also left the package partially loaded. Fixed insql_jobs.pyand the sqlalchemy merge job. - Fix: incremental
lagunit follows the declared cursor type (#4406 @rudolfix) - Fix: incremental
last_valuestays forward-only when the cursor is exactly0(#4367 @chjnett) - Fix:
bind_querygenerated wrong aliases on case-folding destinations (#4359 @rudolfix) - Fix: raise
ValueErrorwhen coercing a non-finite decimal to bigint (#4348 @eeshsaxena) - Fix:
PipelineTracelast-step accessors typed asOptional(#4344 @bunnysayzz) - Fix:
Pipeline.last_tracetyped as possiblyNone(#4469 @VioletM) - Fix: filesystem destination works on PyPy and pyodide (#4385 @tushardev-365): the hard
orjsonimport was removed, and thesimplejsonbackend gets the proper error. Adds alint-emscriptenimport check. - Fix:
read_csv_duckdbno longer uses the global duckdb instance (#4434 @jtcurlin) - Fix: BigQuery table descriptions are escaped in generated SQL (#4473 @rooperuu)
- Fix: Fabric nvarchar precision scaled for UTF-8 byte semantics (#4259 @sdebruyn)
- Fix: Fabric accepts time columns on the parquet load path (#4260 @sdebruyn)
- Fix:
DeletingResourcesNotSupportederror message repaired (#4438 @simpleqt) - Fix: CLI errors no longer end with a generic "refer to our docs" note (#4136 @anxkhn)
Docs
- Release highlights pages for dlt 1.22–1.30 (#4400, #4404, #4408, #4410, #4411, #4412, #4413, #4414, #4416, #4418 @dlt-oss-docs-agent)
- Merge strategies:
cdc, merge conditions and the support matrix (#4305 @adrienModeo); support table formatting (#4508 @rudolfix) - Incremental cursor, lag and context timezone docs (#4033, #4406 @rudolfix)
- Expand
<DocCardList />in the llms-txt output (#4078 @ShreyasGS) - Warn about cursor column mismatch after SQL reflection (#4300 @ShreyasGS)
- SQLAlchemy destination: Oracle env var query params and merge staging schema privileges (#4485 @dlt-oss-docs-agent)
- README examples unified around one Spotify dataset (#4407 @elviskahoro)
- Add the
hotdatadestination to the community page (#4423 @eddietejeda) - Intro page: "AI Workbench" renamed to "AI harness" (#4433 @AstrakhantsevaAA)
- Serve the 404 page with a real 404 status (#4488 @martinibach)
- API reference: drop private modules,
noindeximplementation details, add page descriptions (#4490 @martinibach) - Drop links to the removed
examples/directory (#4439 @simpleqt) - Fix the fundamentals course (#4451) and education lessons (#4475) (@AstrakhantsevaAA)
- Docstring fixes (#4184 @anxkhn, #4437 @simpleqt, #4440 @simpleqt)
Chores
- Docs build tooling upgrade (#4420 @zilto): ruff 0.16,
prekandty; removesdocs_toolsand the old docs workflows. mdsmithformatting rules for docs (#4450 @zilto)- Lint CI pinned to Python 3.10 (#4350 @zilto)
- Remove the dead
test_examples.ymlworkflow (@zilto, direct commit23ecc8fee) - Release-highlights trigger workflow (#4419 @ShreyasGS)
- agentic-docs: pass the Actions run URL to the job (#4364 @ShreyasGS)
- Tests no longer write to
~/.dlt(#4368 @arose26) - Bump to 1.31.0 (#4518 @rudolfix)
dltHub Runtime, Jobs, Deployments and dlthub transformations
Breaking Changes
@job/@pipeline_run:allow_external_schedulersandrefreshdeprecated (#4033 @rudolfix)- Use
incremental_mode="interval" | "pipeline"andrefresh_propagationinstead. - The old keyword arguments still work but emit a
DltDeprecationWarning. They are no longer in the typed signatures, so mypy reports errors for typed callers.
- Use
TimeIntervalContext(allow_external_schedulers=False)is no longer a kill switch (#4033 @rudolfix)- Deployment manifest engine v2 (#4033, #4417 @rudolfix)
- v1 manifests migrate automatically when loaded.
- Every job definition now has an
engine_version. Jobs that take arguments now carryinputs, so every workspace's manifest hash changes once. Expect a one-time diff on the next deploy.
- MCP tools that declare no
RequiresAccessare treated as needing full access (#4417 @rudolfix): plugin tools without the annotation are hidden from callers with a narrower grant, such as agent jobs. Interactive clients are not affected. dlthub ai toolkit install --overwriteremoves stale files (#4417 @rudolfix)- Files from a previous install that the new toolkit version no longer ships are deleted. Installed files are tracked, with their hashes, in
.dlt/.toolkits. - An existing full clone of the toolkit repo is re-cloned once as a shallow, sparse checkout.
- Files from a previous install that the new toolkit version no longer ships are deleted. Installed files are tracked, with their hashes, in
Relation.incremental()advances pipeline state by default (#4033 @rudolfix)- Pass
advance=Falseto get the old stateless filter. - Calling
.incremental()twice now combines both filters with AND instead of raising.
- Pass
Incremental.allow_external_schedulersdefaults toNone(#4033 @rudolfix): an explicit value on the incremental always wins, and a context-levelFalseno longer force-disables it.
Highlights
- Background agent jobs (#4417 @rudolfix)
- What it is: a new job kind whose body is an agent loop, defined in an
AGENT.mdor with a decorated function. - What a definition holds: a system prompt, typed
inputsandoutput(JSON Schema), MCP tool groups, skills, rules and anaccessgrant. - How it runs: declared with
run.agent(...), run locally withdlthub local runor on the runner. - Results: a structured result with
status,summary,resultand a full trace.
- What it is: a new job kind whose body is an agent loop, defined in an
- Pluggable agent loops (#4417 @rudolfix): pydantic-ai is the default and claude-agent-sdk is the alternative. Third-party loops register through the
plug_agent_loophook. Each loop has its own dependency group, installed by the runner. - Access model for jobs (#4417 @rudolfix)
accessdeclareslocal(read/write/execute/network),data(read/write) andcontext(read).localmaps to a standard tool set (Read/Glob/Grep, Write/Edit, Bash, WebFetch/WebSearch) that is limited to the workspace and the temp folder.dataandcontextare enforced by the workspace MCP server: a tool the grant doesn't cover is never shown to the model.- Credential files (
*secrets.toml,.env) are never readable by file tools.
- Agent jobs are callable and testable like Python functions (#4417 @rudolfix)
- You can call or await an agent job directly.
job.last_job_resultexposes the trace, token counts and entities.agent.pynext toAGENT.mdprovidesvalidate_input/validate_outputhooks;JobAbortedExceptionaborts a run cleanly.
- Job-level incremental and refresh control (#4033 @rudolfix)
incremental_mode,refresh_propagationandauto_refresh_pipeline_mode, plus a[jobs]config section.dlt.current.intervalcan be changed from inside a job (set,update,apply_lag,apply_full_days).require.timezoneruns a job in its declared timezone.
Core Library
- Background agent jobs, loops and the access model: see Highlights (#4417 @rudolfix)
- Structured results for every job (#4417 @rudolfix)
run.result(...)declares a typed result.- The launcher delivers an envelope shaped like Activity Streams 2.0 (
type,job_ref,object,result) to telemetry as ajob_result. inputs/outputJSON Schema is available on every job kind.
- Entity-typed inputs (#4417 @rudolfix)
- Inputs can be marked as references to platform entities, e.g.
Annotated[str, Entity("job-runs")]orentity_type:inAGENT.md. Supported types:job-runs,job,pipeline,dataset,workspace. - This fills the result's
objectandexpose.object_input, so the UI can offer an agent on a failed run.
- Inputs can be marked as references to platform entities, e.g.
job.fail:/job.success:triggers accept selectors (#4417 @rudolfix): selectors expand at manifest time to every matching job. They never target the declaring job or interactive jobs.- Agent settings through config (#4417 @rudolfix)
agent.model,agent.instructions,agent.verbosity,agent.max_turnsandagent.max_tokenscan be set per run with-c, env or toml.- The model endpoint is one set:
agent.api_key/api_url/api_version, withruntime_*twins supplied by the runtime. Azure, LiteLLM and xAI are mapped automatically. - Token limits are counted by dlt, so they behave the same on both loops.
- Workspace MCP server:
--features,--no-default-features,--access(#4417 @rudolfix): starts a server that exposes only the requested tool groups and only the tools the grant covers. Tools declare their needs withRequiresAccess. - Toolkits ship agents (#4417 @rudolfix)
- Agents install under
.claude/dlthub/agents/<toolkit>/<name>/(or.cursor/,.agents/), separate from native subagents. - The
.dlt/.toolkitsindex ships with deployments, so<toolkit>:<agent>refs resolve on the runner.
- Agents install under
- Scheduler interval reaches every launcher (#4033 @rudolfix)
- The interval is injected in-process for
job, and passed asDLT_INTERVAL_*env vars to the module, streamlit, marimo, mcp and dashboard launchers. - Profile, interval and refresh env are set before the user module is imported, so pipelines created at import time see refresh runs.
- The interval is injected in-process for
- Fix: MCP
execute_sql_queryread-only check hardened (#4417 @rudolfix): it now accepts exactly one statement, and the whole statement is checked for mutating statements and host functions. - Fix: deployment file selection no longer follows symlinks (#4287 @Sanjays2402): fixes a
RecursionErroron symlink loops. - Fix:
dlthub ai statuswarning no longer mentions MCP (#4445 @AstrakhantsevaAA)
Docs
- Agents section for background agent jobs (#4449 @lis365b) and an agents onboarding page (#4479 @AstrakhantsevaAA)
- Workspace environment variables (#4358 @tetelio; master backport #4366 @tetelio)
- Workspace API keys usage and default access (#4177 @anuunchin)
- Platform alerts (#4391 @AstrakhantsevaAA); updated email and Slack alert instructions (#4499 @schrodervictor)
- dltHub CI/CD guide (#4452 @zilto)
- How a job waits for several upstream jobs (#4513 @ShreyasGS)
- Deploying a marimo notebook to the platform (#4384 @dlt-oss-docs-agent)
- dltHub onboarding page (#4296 @AstrakhantsevaAA); playground workspace page moved (#4355 @kaliole)
- Streamlit page moved under a renamed "Data Apps" category (#4372 @ShreyasGS)
- Fix the
timeout/grace_periodexamples to use the nested form (#4380 @AstrakhantsevaAA) - dltHub release highlights for 0.26 (#4424), 0.27 (#4432) and September 24, 2026, with pages now named by date (#4498) (@ShreyasGS)
Chores
- Disable anonymous telemetry in the platform connection test (#4347 @tetelio): avoids a fork segfault on macOS.
New Contributors
- @anxkhn made their first contribution in #4136
- @sdebruyn made their first contribution in #4259
- @Sanjays2402 made their first contribution in #4287
- @adrienModeo made their first contribution in #4305
- @bunnysayzz made their first contribution in #4344
- @eeshsaxena made their first contribution in #4348
- @chjnett made their first contribution in #4367
- @arose26 made their first contribution in #4368
- @tushardev-365 made their first contribution in #4385
- @eddietejeda made their first contribution in #4423
- @jtcurlin made their first contribution in #4434
- @simpleqt made their first contribution in #4437
- @rooperuu made their first contribution in #4473
- @schrodervictor made their first contribution in #4499