Features
- Head IDs are now generated as UUIDv7.
catalog.DefaultIDGeneratorreturns a UUIDv7 instead of a random UUIDv4, falling back to v4 if generation fails. A UUIDv7 starts with a millisecond timestamp, so head directories and catalog records now sort lexicographically in creation order, which makes reading the data directory and debugging much easier. The format is unchanged (RFC 4122, 36 characters), so existing heads, catalog records and log files need no migration (#519). - Prom++ can now report that it is running with a feature set other than the expected one. The new optional
PROMPP_FEATURES_DEFAULTenvironment variable holds the expected set; at startup it is compared withPROMPP_FEATURESand the result is exposed as theprompp_features_differ_from_defaultgauge —1when they differ,0when they match — registered only when the variable is set. A mismatch is also logged at warn level withadded/removed/changedfields, so it is clear what exactly drifted. The comparison ignores order, whitespace and empty tokens and compares values by meaning:default_sample_age_limit=60mequals=1h,head_read_concurrencyequalshead_read_concurrency=1(#523).
Fixes
- A scrape target serving a malformed native histogram over protobuf could crash Prom++.
ProtobufParser.Histogram()callsCompact(0)before anything validates the payload, andcompactBucketsassumed the span lengths add up to the bucket count — a histogram with buckets behind only zero-length spans walked past the end of the span slice and panicked (found byFuzzParseProtobuf). Compaction is now skipped whenever the spans don't cover exactly as many buckets as the histogram has, and the histogram goes on toValidate(), which rejects it withErrHistogramSpansBucketsMismatch; nothing is dropped or rewritten silently. Upstream Prometheus guards only the case where every span is zero-length (#520). - A
head.logmigration would have wiped the catalog.createSwapFileopened the swap fileO_WRONLY, but after a migration or a compaction that file becomes the workingFileLogand has to be readable too:catalog.syncread right afterNewFileLog, gotbad file descriptor, concluded the log was corrupted and overwrotehead.logwith an empty one — losing the active head, the rotated heads not yet persisted and the remote write positions. It never fired in production, because the log is opened as V2 and no migration takes place, but it blocked movinghead.logto V3 (#518). momentupdated to 2.31.0 in the web UI, closing GHSA-4p3w-j4w9-5jqw — a path traversal through a crafted non-string locale name.
Performance
- A rotated head no longer carries its ingestion lookup structures until retention. They are released in two stages, each at the earliest safe point. On rotation, once the old head is read-only, the label set → ls id hash set and the relabeler input LSS are freed: only relabeling on the active head and the shrink during the added-series copy need them, and both are finished by then. After a successful persist, the sorted ls id set is freed as well — it has to survive until the block is written, because
ChunkRecoderandRevertableLoaderiterate it, and it is kept for the retry if the persist or the status update fails. Rotated heads loaded from disk on restart skip the rotator but still free both. These structures cost several bytes per series per LSS and used to stay allocated while the new active head was already growing, so this lowers both the peak right after rotation and the steady-state memory of every rotated head (#522, #527). - Rotation no longer makes the first commit of the new head expensive. The first WAL segment of a new head holds the label sets of every series copied from the old head, so finalizing it costs far more than a regular commit — and it used to happen on the first
Commitof the already active head, through fastcgo (which occupies the P for the whole call, blocking the Go scheduler and delaying GC stop-the-world) while holding the WAL encoder lock and the LSS read lock, i.e. blocking both the ingest path and new series. That segment is now committed, flushed and synced for every shard right after the added series are copied, before the new head goes live, through a regular cgo call that releases the P. Shards run concurrently, up toGOMAXPROCSat a time; a failure is logged and rotation continues. Rotation itself takes longer as a result, while the old head keeps accepting writes. The newprompp_rotator_head_wal_encoder_long_finalize_duration_nanosecondsmetric reports the longest such call (#525). - Encoding became about 10% faster. The combined Gorilla path is gone from
DataStorage— every encoder now keeps timestamps and values in separate streams. The Gorilla codec itself stays and is still used elsewhere. Theprompp_data_storage_gorilla_countmetric is removed along with the path (#515). - Rebuilding the sorting index no longer scans the whole B-tree. The maximum ls id used to be found by walking the tree; the bound now comes from
QueryableEncodingBimap, which already knows it asnext_item_index(). Build, rebuild and sort take an exclusive ls id bound, so the index size always matches the id space thatupdate()appends to (#516).