The filesystem telemetry and the attribution options from 6.0, backported in
#1376.
Added
memory_pagecachesampler (opt-in): the page cache's traffic per mount,
and per cgroup withcgroup_attribution = true. Buffered read calls and
bytes from onefentryonfilemap_read, pages filled by the filling
task's context (read, write, fault, other) from the add tracepoint, pages
evicted, and mmap faults. Pages filled during reads over bytes read is the
read miss ratio. The memory dashboard's Page Cache group gains the rates
and the miss ratio when a recording has them.xfs_logsampler (opt-in): how long threads block on the XFS log, per
mount and per cgroup, with host-wide latency histograms. Waiting for log
space is thexfs_log_grant_sleep/_wakepair; a log force is
fentry/fexitonxfs_log_forceandxfs_log_force_seq(the fsync's
log write); CIL-full waits are counted. The counts equal
xfs_log_space_sleepsandxfs_log_forcesfromxfs_statsper mount;
the time and the cgroup are what the stats file cannot carry. Needs
kernels 5.12+ with XFS's BTF. The XFS dashboard gains a Blocked Time
group and the cgroups dashboard an XFS log blocked-time plot.xfs_statssampler: XFS's own per-mount counters from
/sys/fs/xfs/<dev>/stats/stats, one sysfs read per XFS mount off the
scrape cycle (1 s default): log writes, forces and force sleeps, in-core
log buffer stalls, log-space requests and sleeps, AIL pusher outcomes,
transactions, inode-cache lookups by outcome and reclaims, extents and
blocks allocated and freed, directory operations, file calls and bytes,
and the metadata buffer cache. Labeledmount,fstype,devnumand
block_devicethrough the same slot registry as the ext4 samplers, which
now also assigns slots to XFS mounts. The viewer gains an XFS section.ext4_opssampler (opt-in, never enabled by[defaults]): how long fsync,
unlink, write and rename held the calling thread inside ext4.ext4_op_latency{op}
histograms; per-filesystemext4_ops{op},ext4_op_time{op},
ext4_op_errors{op}andext4_write_bytes; per-cgroupcgroup_ext4_ops{op}
andcgroup_ext4_op_time{op}, the time a service's request threads are held
inside the filesystem (withcgroup_attribution = true). fsync and unlink from their tracepoints, write and
rename fromfentry/fexitonext4_file_write_iterandext4_rename2
(its arity confirmed from BTF), start timestamps in task local storage;
kernels 5.12+. The ext4 dashboard gains an Operations group and a
Write Path group that puts application bytes, writeback bytes, journal
bytes and device bytes on one axis; the cgroup dashboards gain ext4 Blocked
Time.
Changed
cpu_perfhonourscgroup_attribution(on by default for this
sampler). Withcgroup_attribution = falseitssched_switchprogram is
not loaded, so nothing reads the PMU per context switch;cpu_cyclesand
cpu_instructionsare read per CPU at scrape time and thecgroup_cpu_*
series are absent. On a KVM guest with an emulated PMU the program's two
counter reads measured 20 µs per context switch. A[defaults] cgroup_attribution = falsereachescpu_perftoo.cpu_usageno longer exportstask_cpu_usageunless its section (or
[defaults]) setstask_attribution = true. The per-task accounting the
host and cgroup totals are computed from still runs; what is dropped by
default is the export: the task-metadata and task-exit events, the walk of
the per-pid map's populated slots each time a snapshot is served, and one
series per thread in recordings. Measured under 27 K short threads/s:
refresh p50 78 µs off against 3,690 µs on, host totals unchanged.ext4_journalandext4_alloccounters are per filesystem: every counter
carriesmount,fstype,devnumandblock_device, the labels the
filesystemsampler gives the same mount, withmount="other"for a
device the agent's mount table does not know yet. The BPF programs look the
device up in adev_t → slotmap the agent keeps in step with
/proc/self/mountinfo(rescanned every 10 s, and sooner whenother
moves); counter banks are per CPU and per slot, 8 MiB and 12 MiB of
eagerly allocated map for the two samplers. jbd2 events from an ocfs2
mount get their own slot rather than being folded into the ext4 totals.
Histograms stay host-wide. The ext4 dashboard draws one line per mount.
Queries that sum these counters are unaffected; anything matching their
exact label set sees the new labels.ext4_opsandxfs_logattribute to cgroups only when their section (or
[defaults]) setscgroup_attribution = true. The per-cgroup path was
measured at half the end hook's cost (265 of 535 ns), so it is off by
default; when off, thecgroup_ext4_*andcgroup_xfs_log_*series are
absent and the path is folded out of the loaded program.