github ClickHouse/clickhouse-connect v1.10.0

4 hours ago

clickhouse-connect v1.10.0

Rust codec performance

  • NumPy and Pandas queries no longer require PyArrow. This covers buffered and streamed results and query(..., use_numpy=True). Strings use Arrow conversion when available and Rust object conversion otherwise. Invalid UTF-8 keeps its hex rendering. Explicit Arrow output and Arrow storage still require PyArrow.
  • Supported numeric, Boolean, BFloat16, Interval, date, timestamp, and time columns now convert directly from decoded buffers. Selected nullable, Array, and LowCardinality shapes also use this path. Existing output dtypes stay unchanged, apart from the nanosecond correction described below.
  • Large response chunks now feed the decoder in bounded slices, reducing temporary copies and peak memory for large NumPy, Pandas, and Python results.

Client improvements

  • DB-API connections and cursors now support with blocks. They close on normal and exceptional exit. Connection contexts don't manage transactions thus commit() and rollback() are still no-ops.
  • Async Arrow inserts move DataFrame conversion and Arrow encoding to a dedicated client worker, reducing event-loop stalls. #1054
  • Large async external-data uploads use bounded writes to reduce event-loop stalls during TLS encryption. Upload payloads are replayable for eligible retries. #1057

Query and insert fixes

  • QueryResult.result_rows works after materializing result_columns, including async queries with column_oriented=True.
  • Dynamic shared-storage values and JSON paths beyond max_dynamic_types decode supported compound and scalar values to Python objects. Nullable defaults and typed JSON null paths work, and DataFrame results retain the decoded objects. #1070
  • Internal DESCRIBE TABLE requests get their own query ID. Inserts retain the caller's ID, avoiding QUERY_WITH_SAME_ID_IS_ALREADY_RUNNING errors. #1066
  • After a remote close or HTTP 429/503/504, SQL execution retries only recognized reads. Writes, DDL, mutations, SET, USE, and CHECK TABLE aren't replayed. This also applies through DB-API and SQLAlchemy. Read retries and insert API retries are still available. #1079

Streaming and temporal fixes

  • Synchronous Native response cleanup waits for active reads and drains partially consumed HTTP chunks through the same parser. Early Rust stream closure drains the response before releasing its iterator, avoiding premature closure and subsequent SESSION_IS_LOCKED errors.
  • Rust cleanup releases each response source once, handles closure during read-ahead startup, and releases decoder workers after cancelled pending reads.
  • Async Rust stream cleanup finishes before cancellation is re-raised.
  • Rust NumPy and Pandas results preserve nanoseconds in nullable scalar DateTime64(9) columns, including SimpleAggregateFunction aliases. Unsupported scalar precisions now fail consistently.
  • Nullable SimpleAggregateFunction aliases of Time64 return correct durations and NaT for NULL values.

SQLAlchemy and Alembic fixes

  • Importing the dialect preserves another provider's clickhouse:// registration. Generated schema metadata and Alembic revisions use clickhousedb_* options so copies and migrations work with both drivers installed. #1074
  • GROUP BY uses aliases only for matching top-level expressions selected by the query. Unselected and nested expressions render in full. #1062
  • Nullable() and LowCardinality() preserve the wrapped type in their return annotations. #1033
  • Alembic accepts reflected AggregateFunction state versions when metadata omits the version, including nested types. Explicit versions still detect changes; legacy unversioned states match explicit version 0.

Installation

pip install clickhouse-connect

For the experimental Rust codec:

pip install "clickhouse-connect[rust]"

The Rust extra requires clickhouse-connect-core>=0.2.1,<0.3. The driver checks both binding and column-buffer APIs and reports upgrade guidance for older wheels.

clickhouse-connect-core` releases independently of the driver. To pick up compatible codec fixes without changing your driver version:

pip install --upgrade "clickhouse-connect-core>=0.2.1,<0.3"

The NumPy/Pandas improvements in 1.10.0 also require upgrading clickhouse-connect so a core-only upgrade won't enable them.

For users upgrading both:

pip install --upgrade "clickhouse-connect[rust]" "clickhouse-connect-core>=0.2.1,<0.3"

Upgrade notes

  • In mixed installs with clickhouse-sqlalchemy, use clickhousedb:// or clickhousedb+connect://. Rename ClickHouse Connect schema options to clickhousedb_* and register custom defaults under clickhousedb. Generated metadata uses this prefix. Read engine and Dictionary options through their canonical keys, including in cc-only installs; public reflection dictionaries retain their existing clickhouse_* keys.
  • Type checkers infer list[String] for [Nullable(String)]. Annotate mixed type collections as list[ChSqlaType].
  • Rust NumPy object columns for nullable scalar DateTime64(9) now contain UTC numpy.datetime64 cells instead of Python datetime cells. SQL NULL is still None. Pandas output preserves nanoseconds. All-null block object inference is unchanged. Use query_df to retain named timezone metadata.
  • Insert API retries provide at-least-once delivery and can duplicate rows without server deduplication. Check the outcome before retrying an ambiguous SQL write in your application.

Don't miss a new clickhouse-connect release

NewReleases is sending notifications on new releases.