github confluentinc/librdkafka v2.16.0

3 hours ago

librdkafka v2.16.0 is a feature release:

  • Fix re-bootstrap cases that never reached a bootstrap broker while the learned brokers were still connected, or kept an already connected bootstrap broker without re-resolving its address (#5560).
  • The ALL_BROKERS_DOWN error is now reported only once every reconnect.backoff.max.ms or when the outage restarts (#5600).
  • Avoid duplicate FETCH_STOP for the same toppar during assignment removal (#5574).
  • Fix rd_kafka_clusterid(), rd_kafka_query_watermark_offsets() and rd_kafka_offsets_for_times() waiting until their timeout and accessing the freed client instance when it's destroyed during the call (#5616).
  • Upgraded bundled OpenSSL to 3.5.8 and libcurl to 8.22.0 (#5598).

Security considerations

Bundled dependencies were further upgraded, beyond what v2.15.1 already
covers, for source/autoconf builds:
OpenSSL 3.5.7 → 3.5.8 (LTS); libcurl 8.21.0 → 8.22.0.

Upgrade considerations

  • Admin requests in flight on a decommissioned broker now fail with
    RD_KAFKA_RESP_ERR__TRANSPORT instead of RD_KAFKA_RESP_ERR__DESTROY_BROKER, so
    callers retry them instead of treating them as a hard failure.
  • The ALL_BROKERS_DOWN error is now reported only once every reconnect.backoff.max.ms. In case there are multiple re-bootstrap attempts, caused
    by no available broker connection, this reduces the amount of events while still signalling that the outage is ongoing.

Fixes

General fixes

  • Issues: #5600.
    Fix re-bootstrap cases that never reached a bootstrap broker while the learned brokers were still connected, or kept an already connected bootstrap broker.
    The client kept asking the very brokers that reported its metadata as stale. Learned and bootstrap brokers are now decommissioned when a re-bootstrap sequence starts, so the bootstrap servers are re-created and connected again, re-resolving their addresses, and only they are used until a Metadata response rebuilds the broker list. The transaction coordinator is reset too when its broker is decommissioned, so the coordinator connection is re-established and its address re-resolved, as already done for the group coordinator. Queued messages are handed back to their partitions and re-sent once new leaders are known. Admin requests in flight on a decommissioned broker now fail with RD_KAFKA_RESP_ERR__TRANSPORT (previously
    RD_KAFKA_RESP_ERR__DESTROY_BROKER) and should be retried.
    Happening since 2.11.0 (#5600).
  • Issues: #5546.
    ALL_BROKERS_DOWN was reported on every re-bootstrap cycle during a sustained
    outage. It is now reported once per outage, re-armed when a broker connection
    comes up, and at most once every reconnect.backoff.max.ms while it lasts.
    Happening since 2.11.1 (#5600).
  • rd_kafka_clusterid(), rd_kafka_query_watermark_offsets() and
    rd_kafka_offsets_for_times() now return immediately, with NULL or
    RD_KAFKA_RESP_ERR__DESTROY, when the client is destroyed during the call.
    Previously they kept waiting for metadata until their timeout and then
    accessed the client instance after it was freed. rd_kafka_destroy() now
    waits for these calls to return before freeing it.
    Happening since 0.11.3 (#5616).

Consumer fixes

  • Issues: #5585.
    A consumer with enable.auto.commit=true no longer sends an OffsetCommit
    for an assignment it has already lost. On a client-side session timeout, or
    when max.poll.interval.ms is exceeded, the member id is reset, the
    assignment is marked lost, and its partitions are revoked. The revoke-time
    auto-commit of those partitions was still being sent, because the
    assignment-lost flag was cleared inside rd_kafka_cgrp_unassign() and
    rd_kafka_cgrp_incremental_unassign() before the removed partitions were
    served, defeating the guard that skips commits for a lost assignment. The
    commit went out with an empty member id and the previous generation, which a
    broker rejects with UNKNOWN_MEMBER_ID, or, for a static member
    (group.instance.id set), with the fatal FENCED_INSTANCE_ID that stops the
    consumer. The flag is now kept set until the removal has been served and the
    unassign completes (rd_kafka_cgrp_unassign_done(),
    rd_kafka_cgrp_incr_unassign_done() and, for the KIP-848 consumer protocol,
    rd_kafka_cgrp_consumer_incr_unassign_done()), so the offsets of a lost
    assignment are never committed while the flag is still cleared as soon as
    the revoke is done, keeping commits, close() and unsubscribe() working
    for a member that retains other partitions, and rd_kafka_assignment_lost()
    reporting false again by the time the next assignment is delivered.
    Happening since 1.6.0 (#5585).
  • Issues: #5573.
    Prevents duplicate FETCH_STOP requests during assignment removal
    by introducing an assignment-owned rktp_wait_stop flag (#5574).

Checksums

Release asset checksums:

  • v2.16.0.zip SHA256 df41723517183306a004502d2ecbc05888e0f26d1dd75d5cb9f0cb750de9ce58
  • v2.16.0.tar.gz SHA256 e6b61de61d3282879a88e4ee3d9f634a8b05bf78a9198dfc25c4820c0c9ed231

Don't miss a new librdkafka release

NewReleases is sending notifications on new releases.