RFC-0027: Public Protos for Stack Sampling and Heap Profiling #6027
Replies: 3 comments 6 replies
|
Heap snapshots can use the StackSample proto as written, by using UNIT_BYTES. I also like that the follower_descriptors field can let us encode multiple sample types at the same point. Chrome could align the heap and CPU stack samples on the same timer, for both runtime and trace-size savings. A couple of questions, though:
Is there a drawback to that?
|
6 replies
|
📝 RFC Document Updated View changes: Commit History |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
📄 RFC Doc: 0027-public-profiling-protos.md
Public Protos for Stack Sampling and Heap Profiling
Authors: @LalitMaganti
Status: Implemented
Problem
Today, anyone consuming a Perfetto trace containing stack samples or heap
profiling data has to read protos that were not designed to be public API:
PerfSample(inprotos/perfetto/trace/profiling/profile_packet.proto)is shaped by the
perf_event_opensyscall. Fields likecpu_mode,timebase_count,follower_counts,kernel_records_lost,unwind_error,sample_skipped_reasonandproducer_eventmix threeunrelated concerns: the actual observation, the perf transport, and
producer-side diagnostics. The defaults submessage references
PerfEvents.Timebase, which is also perf-shaped.ProfilePacket(heap profiling) overloads its fields by mode: e.g.self_allocated/self_freedare populated in normal mode whileself_maxis populated when
dump_at_max=true, and they are not all setsimultaneously. It also leaks producer health (
ClientError,buffer_overran,buffer_corrupted,hit_guardrail,from_startup,rejected_concurrent) into the same message as the data. The aggregatedshape is the only one practically usable;
StreamingAllocation/StreamingFreeexist but lack callstacks and are documented as "only forlocal testing".
The consequence: the proto schema is tied to implementation details of how
Perfetto records this data today. We cannot evolve the recording side
(e.g. add new sampling backends, support self-emitting profilers, support
async-runtime-aware profiling) without either breaking consumers or growing
the existing protos with more mode flags.
Additionally, these protos do not naturally accommodate self-emitted
profilers - profilers that run inside the process being observed (e.g. a
language runtime sampling its own coroutines). The existing shape assumes an
external observer with kernel-level identity (
pid/tid/cpu), which isnot always the right model.
We need a new set of protos that:
them and not break.
perf_event_openconcepts leaking through.simpleperf) and self-emitted profilers (runtime samplers, async-aware
in-process profilers).
Decision
Proceed with the transport-neutral
StackSampleproto described below. Thestack-sampling proto and its Trace Processor importer have been implemented.
The heap allocation and free protos are deferred until there is a concrete need
for them. Their design below remains a proposal rather than an implemented or
committed API.
Design
Scope
This RFC proposes new public protos for two distinct data classes:
Stack sampling: point-in-time observations of where a thread (or
coroutine, or other execution context) was, along with the value of a
primary counter and optional follower counters.
Heap profiling in streaming form: per-event allocation and free
records. Heap snapshots (point-in-time set of live allocations, akin to
Java/ART heap dumps) are intentionally out of scope for v1 and will be a
follow-up; see Open Questions.
The protos live under
protos/perfetto/trace/profiling/but the load-bearingsurface is the new top-level fields added to
TracePacket.Top-level shape on
TracePacketThree new fields on
TracePacket:Each event is its own
TracePacket. Granular fields (rather than a singlewrapped
ProfilerSample) were chosen for clarity and because the two dataclasses have genuinely different shapes; see Alternatives.
No
TracePacketDefaultsadditions. Samples are self-describing throughinterned descriptors (
CounterDescriptor,HeapDescriptor, etc.) ratherthan via sequence-scoped defaults. This avoids the "samples are
uninterpretable without the defaults packet" hazard and keeps each sample
locally interpretable from
InternedDataalone.Task context and execution context
The "who" and "where" of a sample are split into two independently
inline-or-interned messages, because they have very different cardinality
profiles and very different update rates:
TaskContextidentifies the subject of the sample - the taskbeing observed (pid, tid, async_id). Slowly-varying (per-thread /
per-coroutine stable); benefits heavily from interning.
ExecutionContextdescribes how the task is executing at sampletime (cpu, mode). Varies per-sample but has very low cardinality
(cores × modes), so interning is still effective.
Every external-profiling event carries a
TaskContext(inline orinterned). Stack samples additionally carry an
ExecutionContext. Heapevents do not carry an
ExecutionContext- cpu/mode are notmeaningful for an allocation record.
Interning notes:
TaskContextper task and referenceit from every sample. Per-sample overhead becomes a single varint.
TaskContexts andO(cores × modes) unique
ExecutionContexts across the whole trace.Both intern very effectively across thousands of samples.
Interning strategy is a producer concern, not a wire-format concern. The
schema supports either inline or interned for both contexts independently.
Stack sampling
Heap profiling (streaming)
Descriptors
Each descriptor type is its own message and follows the same inline-or-
interned pattern. Inline is for simplicity / low-volume cases; interned
(via
InternedData) is for high-rate cases.InternedDatais extended with:unwind_error_iidreuses the existing interned string tables.Standardized enums
Both enums are append-only forever. Producers that need a unit not in
the enum use
unit_stron theCounterDescriptor. NewUnitvalues areadded only when broadly useful.
Callstack model
We reuse the existing
Callstack/Frame/Mappingmodel inprofile_common.protofor the interned form, and mirrorTrackEvent'sinline
Callstack(function_name / source_file / line_number) for theinline form. No new callstack representation is introduced.
For out-of-band sampling profilers, the interned form is the natural
mode (raw PCs + mappings, offline symbolization). The inline form is for
synthetic / low-volume / test cases where simplicity outweighs efficiency.
Semantic stability commitments
These protos commit to forever-stable semantics, not just wire stability.
Concretely:
of other field values or defaults. The
self_max/self_allocatedmode-flip in the current
ProfilePacketis exactly the pattern we willnot repeat. New states get new fields.
ModeandUnit(and any future enum)never repurpose existing values.
"the primary counter measures CPU cycles" but never "field X means Y
when default Z is set".
ring buffers, unwinder failures, guardrail trips, client errors) are not
in this proto. They go in a separate sidecar packet defined in a follow-
up RFC. The only producer-side hint we keep on the sample itself is
unwind_error, which describes the data (the stack is partial), notthe producer's internal state.
Relationship to existing protos
The existing
PerfSample,ProfilePacket,StreamingAllocation,StreamingFreeandStreamingProfilePacketare not deleted, deprecated,or affected by this RFC. They continue to work and continue to be the
emission format for heapprofd and traced_perf. There is no goal of
migrating those producers to the new protos. The new protos are aimed at
new producers - in particular OS-level samplers other than
perf_event_openand self-emitting in-process profilers - that have noincumbent format and need a stable public surface to target.
Relationship to
TrackEventThese protos are deliberately not part of
TrackEvent.TrackEventis in-band instrumentation ("this thing happened in my program, here's a
callstack of where"). External / out-of-band profiling is a separate
concept, sibling to
TrackEventand to the generic kernel-event protos.The two concepts share the
InternedDatainterning model (callstacks,frames, mappings), but the message shapes do not depend on each other.
Alternatives considered
A. Single unified
ExternalProfileSamplemessage with a oneof payloadPro:
TracePacketfield to extend later.Con:
different shapes.
these, it can join the family without retroactively requiring the
wrapper.
We prefer granular
TracePacketfields.B. Extend
TrackEventrather than introducing new top-level protosPro:
Con:
TrackEventis in-band instrumentation. External profilers are notthe instrumented program describing itself; they are an external (or
internal but out-of-band) observer describing the program. Conflating
these has the same shape as conflating
ftraceevents withTrackEventTrackEvent's per-event overhead (event name, debug annotations, trackreference) is unnecessary at kHz-scale sampling rates.
C. Window-aggregated heap profile (preserving today's
ProfilePacketshape)Pro:
Con:
live (
dump_at_maxmode flip,continuedchunking, peak vs windowambiguity).
true.
deserves its own message, not a third mode of the same one.
We commit to streaming-only in v1; a future heap-snapshot proto will be a
distinct message.
D. Batched / columnar (SoA) per-packet shape
Pro:
StreamingAllocation/StreamingProfilePacketalready usethis shape.
Con:
Decision: AoS (one message per event) is the canonical public shape. If
measured to matter, a sibling batched message can be added as a future
optimization without disturbing the canonical shape.
Open questions
Heap snapshots / dumps. A point-in-time "set of currently-live
allocations per callsite" is semantically distinct from streaming
alloc/free and from a window aggregate. It also has natural overlap with
Java/ART
heap_graph.proto. Out of scope for this RFC; follow-up.Stackless coroutines. C++20 / Rust async coroutines do not have a
contiguous stack; their "callstack" is an await chain / state machine.
v1 supports stackful coroutines (goroutines, Boost.Context, fibers)
using the existing
Callstackmodel. Stackless coroutine modelling isdeferred.
Producer diagnostics sidecar. What does the new "producer health"
packet look like (data loss, unwinder errors, guardrails)? Separate
follow-up RFC.
💬 Discussion Guidelines:
All reactions