ui: Restructure flamegraph hash + merge for ~22% speedup on bottom-up
Two structural changes to the hash and merge_hashes pipeline that
together cut native-side hash+merge cost from ~31s to ~22s on a 2.5M
row bottom-up flamegraph (and proportionally more on WASM):
1. Skinny hash output. _viz_flamegraph_upwards_hash and
_viz_flamegraph_downwards_hash no longer pull in the (small, fixed)
|grouping| columns or |name| at the end of the walk. Only |grouped|
columns and the propagated value/cumulativeValue stay - everything
else is filled in later from a single per-merged-row JOIN against
|source| in resolve_groups, instead of a per-walk-row JOIN (~4x
fewer JOIN rows since the walk produces ~9.3M rows that merge into
~2.5M).
2. Drop the ORDER BY hash from the hash table; replace the per-row
LIMIT 1 subquery in merge_hashes with a self-join on a hash-indexed
grouped table. The old ORDER BY cost ~12s on 9.3M rows. The new
pipeline is:
hash (skinny, no ORDER BY) -> INDEX hash(hash)
-> group_hashes (one row per unique hash, with rep_id) -> INDEX grouped(hash)
-> resolve_groups (self-join for parentId + JOIN source for cols)
The materialized + indexed grouped table makes the parentId resolution
a clean indexed lookup rather than a per-row scan.
3. resolve_groups must ORDER BY g.id at the end. Without it the merged
table ends up stored in source-JOIN order rather than id-order, and
the downstream _graph_aggregating_scan in the trim propagation step
degrades by ~5x. Documented inline.
merge_hashes is removed in favor of the two-step group_hashes +
resolve_groups - there are no other callers of these macros.
Verified end-to-end: identical output (row count, root cumulative,
placeholders) and ~10s native saved on the example trace. WASM
should see proportionally more (the user reported 74s hash + 34s
merged; estimated savings 30-50s).
Notes journal at: /tmp/claude/flamegraph_perf_journal.md
Perfetto is an open-source suite of SDKs, daemons and tools which use tracing to help developers understand the behaviour of complex systems and root-cause functional and performance issues on client and embedded systems.
It is a production-grade tool that is the default tracing system for the Android operating system and the Chromium browser.
Perfetto is not a single tool, but a collection of components that work together:
Perfetto was designed to be a versatile and powerful tracing system for a wide range of use cases.
ftrace, allowing you to visualize scheduling, syscalls, interrupts, and custom kernel tracepoints on a timeline.chrome://tracing. Use it to debug and root-cause issues in the browser, V8, and Blink.We‘ve designed our documentation to guide you to the right information as quickly as possible, whether you’re a newcomer to performance analysis or an experienced developer.
New to tracing? If you're unfamiliar with concepts like tracing and profiling, start here:
Ready to dive in? Our “Getting Started” guide is the main entry point for all users. It will help you find the right tutorials and documentation for your specific needs:
Want the full overview? For a comprehensive look at what Perfetto is, why it's useful, and who uses it, see our main documentation page:
For users interested in the Debian distribution of Perfetto, the official source of truth and packaging efforts are maintained at Debian Perfetto Salsa Repository
Have questions? Need help?
We follow Google's Open Source Community Guidelines.