ui: Three flamegraph perf wins: skip empty walk, pre-floor trim, parentCum Three independent bit-identical optimizations that together shave ~20% off native bottom-up pipeline time (measured -16s on a 17M-walk-row heap-graph trace, 83s -> 67s). All three are surgical: same output, no algorithmic change, just removing wasted work. 1. Skip the empty UNION side in the hash walk. For BOTTOM_UP the downwards walk has empty inits (showDownward=FALSE) so it produces no rows, but graph_scan still does setup work. For TOP_DOWN the upwards walk is similarly empty (isPivot is never true). In the TS layer, only call whichever macro actually produces rows. 2. Pre-floor the trim propagation. The floor threshold (min_value) is now applied at the scan edge filter instead of inside the per-node IIF. Dest nodes below min_value are unreachable from the seeded roots, which is exactly the "dead" semantic the original implementation achieved via +inf propagation. On this trace the scan input shrinks from 8.4M rows to ~108K rows; the propagation step drops from 12.65s to 1.03s. Output bit-identical. 3. Precompute parentCumulativeValue. resolve_groups already does LEFT JOIN $grouped p to assign parentId; pulling p.cumulativeValue through as a new column is free. trim_with_placeholder now carries it forward (placeholders use MIN(d.parentCumulativeValue) since all dropped children of a parent share it). global_layout reads it directly from s instead of doing its own LEFT JOIN $merged p. global_layout drops from 13.2s to 1.0s. Also adds ORDER BY id to trim_with_placeholder. The merged UNION ALL wasn't emitting rows in dense id order, which the journal from the prior commit noted makes graph_scan 5x slower downstream in global_layout. Fixing it recovers global_layout's speedup from E.
Perfetto is an open-source suite of SDKs, daemons and tools which use tracing to help developers understand the behaviour of complex systems and root-cause functional and performance issues on client and embedded systems.
It is a production-grade tool that is the default tracing system for the Android operating system and the Chromium browser.
Perfetto is not a single tool, but a collection of components that work together:
Perfetto was designed to be a versatile and powerful tracing system for a wide range of use cases.
ftrace, allowing you to visualize scheduling, syscalls, interrupts, and custom kernel tracepoints on a timeline.chrome://tracing. Use it to debug and root-cause issues in the browser, V8, and Blink.We‘ve designed our documentation to guide you to the right information as quickly as possible, whether you’re a newcomer to performance analysis or an experienced developer.
New to tracing? If you're unfamiliar with concepts like tracing and profiling, start here:
Ready to dive in? Our “Getting Started” guide is the main entry point for all users. It will help you find the right tutorials and documentation for your specific needs:
Want the full overview? For a comprehensive look at what Perfetto is, why it's useful, and who uses it, see our main documentation page:
For users interested in the Debian distribution of Perfetto, the official source of truth and packaging efforts are maintained at Debian Perfetto Salsa Repository
Have questions? Need help?
We follow Google's Open Source Community Guidelines.