5.23.0 was tagged but never published: its Ubuntu 20.04 package failed to
build, so no packages or release were made. 5.23.1 is that release without
focal.
Removed
- Ubuntu 20.04 (focal) packages are no longer built. Focal ships clang 10,
and thememory_writebacksampler's CO-RE enum relocations
(bpf_core_enum_value) need clang 12, so the focal package no longer
compiles. Focal's standard support ended in May 2025 and its GA kernel
(5.4) is below the 5.8 minimum. Its repository keeps serving 5.22.1, the
last focal release;install.shstill installs it there, with a warning.
Added
rezolus mcp installregisters the MCP server with Claude Code and
installs therezolus-mcpskill, which tells the client how to investigate
a recording (describe it and its metrics before querying, extract features
before forming hypotheses, pick one recording out of a multi-recording
archive). Registration runsclaude mcp addin user (default) or project
scope and replaces an existing entry in that scope; withoutclaudeon
PATH, project scope writes.mcp.jsonand user scope prints the command.
--dry-runchanges nothing and--no-skillregisters only. The skill file
is replaced only when it is this skill and never through a symlink, and a
warning names the removal command when another scope'srezolusentry
would win. Backported from 6.0 without the server flags, since 5.x has
only the read tools.memory_slabinfosampler: sizes of the dentry, VFS/ext4/XFS inode,
ext4 extent-status, jbd2 journal-head, buffer-head and page-cache-index
slab caches from/proc/slabinfo(memory_slab_cache_objects{cache,state},
memory_slab_cache_bytes{cache}), read at most once perinterval
(60 s default) off the scrape cycle. Answers whether the metadata caches
fit;ext4_inode_cachehere is the cause behindext4_inode_loads.
Caches the kernel lacks are absent. The Memory dashboard's Kernel subgroup
shows the caches when the recording has them.filesystemsampler: for each ext4 filesystem,filesystem_errors(the
superblock's persisted error count, from sysfserrors_count) and
filesystem_written_bytes(lifetime bytes written to the block device,
journal included, from sysfslifetime_write_kbytes), read on the existing
60 s sweep and absent on other filesystem types. The Filesystem dashboard
gains Device Writes and Errors cards.syscall_counts/syscall_latency: asyncclass (op="sync") for
fsync,fdatasync,sync,syncfsandmsync, split out of the
filesystemandmemoryclasses so their latency, a device round trip,
is no longer averaged into metadata calls that complete from cache. New
syscall{op="sync"},cgroup_syscall{op="sync"}and
syscall_latency{op="sync"}series; the Syscall dashboard gains a Sync
subgroup.ext4_allocsampler: BPF on ext4's block allocator and the metadata reads
around it. Per extent allocation, requested versus returned blocks, groups
scanned and the criterion reached (ext4_allocations,
ext4_allocation_blocks{kind},ext4_allocation_groups_scanned,
ext4_allocations_by_criterion{criterion},ext4_allocation_size);
blocks freed; inodes allocated and freed; writeback passes with pages
written, skipped and errors; discarded blocks and preallocation releases;
and the synchronous inode-table and bitmap reads (ext4_inode_loads,
ext4_bitmap_loads{kind}). Host-wide. The viewer's ext4 section gains
Allocator, Inodes, Writeback and Metadata Reads groups.memory_writebacksampler: BPF on the page-cache writeback tracepoints.
writeback_throttle_latency,writeback_throttle_checks,
writeback_throttle_eventsandwriteback_throttled_timefrom
balance_dirty_pages(the sleep the dirty-page throttle imposes on
writers);writeback_runs{reason}fromwriteback_start;
writeback_pages_written.balance_dirty_pageshas had two argument
lists, so the sampler reads the tracepoint's argument count from BTF
(kernel_btf_tracepoint_arg_count) and loads the matching program, or
disables the hook and reports degraded when neither matches. The viewer's
Memory section shows throttle latency, throttle rate, runs by reason and
pages written when the recording has them.memory_vmstatexportsmemory_pages_dirtiedandmemory_pages_written
(nr_dirtied,nr_written): the complete page-cache writeback counts,
including pages written by integrity syncs, which the flusher's
writeback_pages_writtentracepoint does not account for.memory_vmstatexports the rest of what/proc/vmstatsays about memory
pressure, from the file it already reads: the dirty limits the writeback
throttle uses (memory_dirty_threshold{kind}, gauges in bytes); reclaim
scanned and reclaimed by kswapd versus direct (memory_reclaim_scanned,
memory_reclaim_reclaimed),memory_allocation_stallssummed over
zones;memory_page_faultsandmemory_major_page_faults; swap in and
out; working-set refaults, activations and restores by file/anon (5.9+);
memory_oom_kills; THP faults by outcome, collapses and splits;
compaction stalls and outcomes; and the NUMA balancer's PTE updates, hint
faults and migrations. Absent lines leave their metric absent. The viewer's
Memory section gains Reclaim, Working Set and dirty-limit plots, gated on
the recording.memory_meminfoexports 30 more gauges from the/proc/meminfoit
already reads:memory_dirtyandmemory_writeback; the LRU lists
(memory_active/memory_inactivebykind,memory_unevictable,
memory_mlocked,memory_shmem,memory_mapped,memory_anon); kernel
memory (memory_slabbykind,memory_kernel_reclaimable,
memory_kernel_stack,memory_page_tables,memory_percpu); swap
(memory_swap_total/_free/_cached); overcommit (memory_commit_limit,
memory_committed); huge pages (memory_hugepages_anon/_shmem/_file,
memory_hugetlb,memory_hugetlb_pagesbystate); and
memory_hardware_corrupted. A line the kernel does not print leaves its
gauge absent rather than 0. The viewer's Memory section gains Writeback,
Page Cache, Anonymous, Kernel, Swap and Huge Pages subgroups, shown when
the recording has the metrics.ext4_journalsampler: BPF on the jbd2 and ext4 tracepoints. Every phase
of each journal commit (ext4_journal_commit_latency{phase}), commit,
handle and block counts, checkpoint latency and counts, lock-buffer stalls,
fsync/fdatasync counts and errors (ext4_sync_file,
ext4_sync_file_errors), andext4_errors/ext4_shutdowns. Host-wide.
jbd2 reports its phases in jiffies; the sampler converts with a tick
measured byclock_getres(CLOCK_MONOTONIC_COARSE). Thetp_btf/raw_tp
twin is chosen per hook bykernel_btf_has_tracepoints, which consults
module BTF as well as vmlinux, since ext4 is a module on some kernels.
Viewer gains an ext4 section. Design:docs/journal/2026-09-28-ext4-sampler.md.- The reader reads long tables in dendro archives, the 6.0 layout: one row
per tick and occupant, with each occupant's labels in a parquet stream
beside the table (<table>/occupants; the format is metriken-segment
0.1.0's, the relabel metriken-query 0.32.0's).
Series carry the occupant's labels plus an internal__occupant__, and
filters on those labels work. Checked against the same data written one
column per slot: every query agrees. Nothing writes this layout yet. rezolus view,rezolus mcpand the static-site viewer open dendro
archives, recognized by content like a.rez.RezReaderreads its
container through aCatalogtrait (crates/rez/src/catalog.rs) that
.rezv3 and dendro both implement, so an archive written by
recording upgrade --to dendroreads as the.rezit came from. Checked on
two real recordings (581 MB and 1.28 GB): the samemcp queryanswers from
the original and the conversion, per-thread and per-cgroup queries
included. Therecordingsubcommands still take.rezonly.recording upgrade --to dendro in.rez -o out.dendrowrites a copy of a
.rez(v1, v2 or v3) as a dendro archive, the container #1224 plans for
6.0. Recordings become sources and tables become streams; segment, WAL and
caller-row bytes are copied unchanged, and WAL rows a sealed segment
already holds are dropped. A timestamp abovei64::MAXis refused rather
than wrapped.-ois required and must not exist, because no rezolus
release reads the output yet; opening one anywhere a.rezis expected now
fails naming it as a dendro archive rather than withno such table: recordings. Measured on a 1.28 GB, 5,874-segment archive: 7.0 s, and the
output matches the input table by table.
Changed
- The archive reader moved to metriken-archive 0.1.0 (
ArchiveReader), with
no change in behaviour: every reader test passes unchanged.RezReader
is now a wrapper that recognizes the container (a.rezv1/v2/v3 or a
dendro archive) and derefs toArchiveReader, so call sites are
unchanged.RezDbimplementsmetriken_archive::Catalog, and the
identity index is read throughreader::IdentityIndex, an
metriken_archive::IndexRelabel.catalog::Container::of_pathreturns
Nonefor a file that is not a catalog container. - The WAL row format (
WalGroupRow,WalCell,WalValue) and WAL-tail
materialization moved fromcrates/rez/src/wal.rsto metriken-segment
0.1.2, andwal_group_row/group_approx_bytesto metriken-exposition
0.21.1, with no change in behaviour;rez::walandrez::rezre-export
them.WalValue::ofis nowrez::wal::wal_value. The unused pre-dendro
stream framing incrates/rez/src/wire.rs(StreamFrame,
encode_frame,decode_frame,encode_frame_filtered,
STREAM_CONTENT_TYPE) is deleted:/metrics/streamhas sent dendro's
replication frames since it shipped. - The wide segment format, meaning the table model, its parquet encoding
and decoding,TableBuilder/GroupTableBuilder,Windowand the
GroupSchemamirror, moved fromcrates/rezto metriken-segment 0.1.1,
with no change in behaviour.crates/rezre-exports it under the old
names.TableBuilder::push_entriesis now therez::rez::PushEntries
trait. Dependencies: metriken-exposition 0.21.0, metriken-query 0.33.0. - Agent:
syscall{op}is now per-CPU, one series per op per CPU with an
idlabel, ascpu_usageis (docs/principles.md principle 9). The BPF
map already kept a bank per CPU; the sampler summed them before export.
Dashboards sum over CPUs and are unchanged. A Prometheus scrape of the
exporter now gets per-CPU series, so a query that read the host total
needssum without (id). Measured on a 32-CPU host: refresh p50 43 µs and
p90 96 µs, against 45 µs and 85 µs before. - dendro 0.2.2 → 0.3.0:
Frameis#[non_exhaustive], so the recorder's
frame naming has a catch-all arm. Behaviour is unchanged: the recorder still
refuses any frame kind it was not written for. - Hindsight logs the address its HTTP endpoint actually bound, not the
configured one, solisten = "127.0.0.1:0"reports the port it got.
Fixed
- Agent: a
/metrics/streaminterval with no reading to send (before the
first sampling pass, or when a snapshot failed to encode) now gets an empty
Rowsframe with the nextseq, as every other interval does. It was
skipped, which left a gap a subscriber reads as lost intervals, and a
subscriber waiting on its first interval waited for the next one. - Agent:
cpu_usagelost CPU from the per-CPU and per-cgroup totals under
task churn, not only from the per-task view. Two causes, both in how a new
task is started in BPF:- A task whose
task_infoevent did not fit in the ring buffer (drained
only when a snapshot is taken, about 1,130 events) looked new on every
accounting hit and had its baseline re-zeroed each time, so every delta
was skipped. The baseline is now set once per task instance and only the
metadata send is retried. - Every task's first observation was skipped, which drops all the CPU of a
thread that lives about one tick, and each field's first non-zero value.
A task that started after the agent attached is now counted from zero; one
that predates it is counted from its first observation, so an agent
restart still does not credit running tasks' lifetime CPU to one tick.
Measured on a 32-core host with 60 s ofstress-ng --pthread(about 18,000
threads/s) against the kernel'scpu.stat: the cgroup's CPU went from 22%
of the kernel's figure (5.20.0) to 85%, system time from 20 to 80 of 94
core-seconds. Without churn, totals were already within 1%.
- A task whose
- Recorder: retention of the identity index kept too little. It cut each
stream'scaller_rowsat the latestFullat or before the row cutoff, but
a segment is evicted only when its newest row is older than the cutoff, so a
segment spanning the cutoff kept rows older than it and lost theFull
they depend on. The reader would then skip theDeltas before its first
Full, so up to one seal age (300 s) of rows at the old end of a rolling
buffer would read without their task and cgroup labels. The cut is now the
latestFullat or before the oldest row the stream still holds after
eviction. No shipped path hit this:hindsightis the only caller of
retention and it scrapes, so it writes no index, andrecord --stream
writes the index but never evicts. It would have appeared once the rolling
buffer took the stream. (introduced in #1281, v5.22.0)