github dapr/dapr v1.17.15
Dapr Runtime v1.17.15

4 hours ago

Dapr 1.17.15

This update contains the following bug fix:

  • Azure Event Hubs checkpointed messages regardless of whether the application processed them

Azure Event Hubs checkpointed messages regardless of whether the application processed them

Problem

The Azure Event Hubs pub/sub and input binding components advanced (checkpointed) their Event Hubs offsets without regard to whether the application actually handled the event successfully.
processEvents() invoked the application through handleAsync(), but the error handleAsync() returned was discarded at the call site, leaving UpdateCheckpoint() with no condition to check before running.

Both delivery modes were affected, in different ways:

  • With the default enableInOrderMessageDelivery: false, checkpointing happened fire-and-forget, independent of message delivery to the application.
  • With enableInOrderMessageDelivery: true, the component waited for the handler to return before checkpointing, but checkpointed anyway when the handler responded RETRY.

Impact

You were affected if you consumed messages through the Azure Event Hubs pub/sub component or received events through the Azure Event Hubs input binding.
An event your application returned RETRY for, or one still in flight when the sidecar crashed or partition ownership was lost, was checkpointed and never redelivered, silently downgrading Dapr's documented at-least-once delivery guarantee ("Dapr attempts to redeliver the message until successful delivery") to at-most-once.
No error was surfaced to the publisher or subscriber when this happened.
Sibling pub/sub components such as Kafka and Service Bus already conditioned their commit or checkpoint on successful handling; only Event Hubs did not.

Root Cause

The error returned by handleAsync() had no guard condition attached to it at its call site in processEvents(), so UpdateCheckpoint() ran unconditionally after every event or batch, regardless of whether the application had acknowledged it successfully.

Solution

The Event Hubs pub/sub and input binding components now advance a partition's checkpoint only through contiguous, successfully-handled events or batches:

  • A handler failure is retried with the configured backoff until it succeeds or the subscription is cancelled, instead of being dropped and checkpointed anyway.
  • Under concurrent handling, a partition's checkpoint advances only through the newest contiguous run of successful batches, and checkpoint writes for a partition are serialized so they can no longer race each other out of order.
  • Checkpointing is skipped once the subscription is cancelled or partition ownership loss is observed, and a checkpoint-store failure now surfaces as an error instead of being silently swallowed.
  • A new maxConcurrentHandlers option (default 100) bounds how many received-but-not-yet-checkpoint-eligible events can be in flight at once, pausing further receives once the bound is reached.
  • Both delivery modes now start from the earliest retained event when a consumer group has no stored checkpoint yet; an existing stored checkpoint always takes precedence.
  • A new subscription-level checkPointFrequencyPerPartition option controls how often the checkpoint is persisted for a partition (starting with the first batch, then every N contiguous successfully-processed batches, or never if set to 0); this only affects how far a restart may replay already-successful batches, not the correctness guarantee itself.

enableInOrderMessageDelivery: false remains the default and continues to control ordering only — it no longer has any bearing on whether a checkpoint requires success.

Don't miss a new dapr release

NewReleases is sending notifications on new releases.