RFC-0039: Unifying machine, GPU, and custom track dimensions #6698
Replies: 5 comments 7 replies
|
Few thoughts from my side:
Rough thoughts, but I think there is some direction here which will work. If you want to take antoher stab at iterating, that would be good or otherwise, I can also sit down and come up with an evolution of this proposal as well but it will take me some time to get to this. |
|
📝 RFC Document Updated View changes: Commit History |
|
Alright sorry for the delay on this finally got some time to sit down and look at this. Some thoughts:
Dimensions should also be exposed in TP of course and in the details panel so this is a bit too strong of a framing.
This is a tad confusint to me, specifically how dimensions fit within process and thread: Could you be a bit more specific in how you see custom dimensions interacting with these? Specifically, I'm a little concerned about the "implicit child inheritence concept".
The way I see it is that GPU is just a special dimension which happens to be presented differently at the UI level. At the TP level, it should be indistinguishable from any other dimension concept.
Going back once more to above: how I see it is simply that this is a process scoped track which happens to have extra dimensions. Becuase process scoping is somehting we have as a top level concept in Perfetto, we will treat it special but there shouldn't need to be anything special for this at the TP table/importer level.
We should just add this in
Dimension is great, let's do it.
I think this set seems reasonable.
I personally think we should error out if a parent has a dimension and a child decides to "override" it - it just breaks the model too much.
Yeah we should consider interning, it's cheap and could help with the merging problem.
Yeah this gets into the game of "is it per trace or per-trace and machine" concept. Personally I would say it's per trace and machine as that simplifies our life a lot.
Probably yes. |
|
📝 RFC Document Updated View changes: Commit History |
|
@LalitMaganti sorry for the delay. addressed your feedback and to be able to have a better idea what it looks like in code I also created this prototype change: https://github.com/google/perfetto/tree/dev/dreveman/rfc0039-dimensions-prototype ptal |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
📄 RFC Doc: 0039-track-grouping-dimensions.md
Unifying machine, GPU, and custom track dimensions
Authors: @dreveman
Status: Draft
PR: N/A
Problem
Perfetto has grown several independent mechanisms that do fundamentally the same
thing: take an identifier attached to a set of tracks and, only when more than
one distinct value is present, surface it in the UI — either as a name suffix or
as an extra level of track hierarchy. Each is hand-rolled, and each bakes in a
different, fixed choice of how it appears:
Machine. In a multi-machine trace, thread/process/CPU/GPU tracks get a
(machine N)(or(<name>)) suffix — a label, no dedicated tree node.Implemented by a shared machine-label helper, a dense per-machine index
(
label_index) in the preludemachineview, and applied at ~10 track-namingcall sites.
GPU. In a multi-GPU trace, GPU tracks get an extra hierarchy level plus a
GPU N/ name label. Implemented indev.perfetto.Gpu, with a separate,re-implemented per-process gate in
dev.perfetto.GpuByProcess. The GPU codealready reuses the machine-label helper and folds machine into its sort order.
Custom, workload-defined dimensions (the gap). There is real demand for
another identifier that is generic and specific to the workload — declared by
the producer rather than derived by the trace processor. The motivating example
is the rank of a PyTorch distributed-training process, but "shard",
"replica", "worker", "stage" are all the same shape. There is no way for a
producer to declare such a dimension today, and adding one the current way would
be a third hand-rolled copy of the same collapse-and-label logic.
The data model is already partly generic — arbitrary track dimensions live in the
track's dimension arg set and are read via
extract_arg— but there is noproducer surface for a generic labeled dimension (
TrackDescriptoronly has typedprocess/thread/counter/state), and no shared labeling layer abovedimensions.
We define one concept — a track dimension, in trace_processor's existing
sense — that machine, GPU, process/thread, and custom identifiers are all
instances of, and surface it through a single labeling layer. The goal is:
machine, GPU, process/thread, and custom identifiers.
independent of whether it came from a typed descriptor, import context, or a
producer-declared custom value.
machine and custom dimensions use them directly, GPU's existing hierarchy
consumes the same label metadata, and details panels show the same values.
A dimension is first a trace_processor/query concept. Track subtitles are its
normal timeline presentation, but not its only surface: dimensions are also
available to SQL, details panels, and specialized UI. A well-known dimension can
have specialized presentation — GPU hierarchy today — without becoming a
special data model. Hierarchy for system-wide concepts comes from
trace_processor identity/merging, not a generic UI grouping mode (see
Well-known identity vs producer-local structure and Alternatives).
Throughout this RFC, "custom dimension" means a producer-declared,
workload-specific dimension; rank is used only as a concrete example of one.
Decision
Adopt
Dimension/TrackDescriptor.dimensionsas a producer surface andnormalize custom and well-known dimensions into one typed, track-keyed
trace_processor relation. Dimensions declared on process/thread tracks resolve
through existing
upid/utidassociation; dimensions on ordinary tracks resolvethrough
parent_uuid. Conflicting child overrides are invalid. The standard UIpresentation is a multi-label track subtitle, details tabs show the same resolved
values, and specialized consumers such as GPU hierarchy use the same TP data and
label helpers. The initial well-known set is
machine,gpu,cpu,process,and
thread; custom string data is interned; numbering is per source trace andmachine.
Design
Terminology: dimensions
We align on trace_processor's existing concept. A dimension is a named,
typed key/value in a track's identity. There is one vocabulary and one resolved
trace_processor representation regardless of whether the value originated in a
typed descriptor, import context, or
TrackDescriptor.dimensions.Dimensions fall into two categories:
sources: initially
machine,gpu,cpu,process, andthread. Theirdefining property is that the same canonical value in different data sources
refers to the same real thing. This permits validation, merging, and specialized
presentation. The set is declared centrally; custom producers cannot redefine
these reserved names.
rank,shard,stage, …). Their identity is local to the producer/source trace andthey are not cross-data-source merge keys.
A dimension declaration carries:
rank,shard, …).worker-east.A dimension has no independent
scopeenum. It is declared on a track. Existingprocess/thread association and explicit
parent_uuidrelationships determinewhich other tracks resolve that declaration as part of their effective dimension
set. “Process-scoped rank” is shorthand for “rank declared on the process track,”
not a separate storage or importer concept.
Presentation metadata adds a label template (
machine %d,GPU %d,rank %d)and stable numbering. Numbering is deterministic and gap-free within a
(source trace, machine, dimension name)partition. A producer-supplieddisplay_nameoverrides only the rendered numbered label; it does not replacethe canonical typed value used by SQL or merging.
Well-known identity vs producer-local structure
There is a useful distinction in how a producer's track event relates to the rest
of the trace. It shapes how dimensions participate in identity and presentation:
something the trace already knows about globally, so track-event tracks want to
merge with other data sources that carry the same dimension. GPU is the
example: GPU track-event tracks should sit alongside GPU counter-descriptor
tracks for the same GPU. Thread/process are the same shape. This merging is a
trace_processor responsibility, keyed on well-known dimensions — and it is
what produces GPU hierarchy today.
owns the structure entirely via
TrackDescriptor.parent_uuid. Custom dimensionsare here. Grouping is the producer's; we only label.
Hierarchy that comes from merging system-wide concepts belongs in trace_processor,
not in a UI mode layered over arbitrary producer trees. This RFC builds the
shared producer and trace_processor dimension model plus its subtitle/details
presentation. Existing GPU hierarchy remains a specialized UI consumer of the
well-known
gpudimension. Generalizing identity merging to more well-knowndimensions remains future work.
Presentation: subtitles, details, and specialized UI
The normal timeline presentation of a dimension is a label: a secondary
annotation on the affected node, not a mutation of its canonical name. This RFC
adds a first-class track/group subtitle affordance (secondary text under the
name), modeled on Chrome's existing behavior. It supports multiple labels in a
stable order, separated by
·, with truncation and the complete values availablethrough tooltip/accessibility text.
The shared collapse rule is unchanged: if the applicable tracks have a single
distinct value for a dimension, its subtitle label is invisible; if they have
more than one, the label is shown. Dimensions collapse independently. Labeling
never reparents tracks.
A process
traineron machine 1 with customrank 3in a multi-machine,multi-rank trace:
The same effective dimensions are shown in details tabs as canonical name, raw
typed value, and optional display name. SQL exposes them independently of UI
presentation.
No generic LEVEL mode. GPU hierarchy remains specialized presentation in
dev.perfetto.Gpu/dev.perfetto.GpuByProcess. At the trace_processor level,gpuis indistinguishable in shape from another dimension. The GPU UI consumesthe shared dimension metadata and label helper, applies the resulting
GPU N/name as its existing group-node title, and keeps its hierarchy behavior:
Resolution through
parent_uuidand process/thread associationLabeling never reparents anything, but effective dimension resolution is explicit
and shared by SQL, subtitles, details, and specialized UI:
the same
upid, including its threads, process tracks, and per-process GPUtracks.
the same
utid.parent_uuiddescendants.declared on itself or its explicit ancestors, plus synthesized well-known
dimensions such as machine.
Repeating the same name and typed value is redundant and deduplicated. A
different value for an inherited name is invalid producer input: trace_processor
records an import error and does not apply the override.
For example:
But:
Sibling subtrees may carry different values when their common parent did not
declare that dimension. This keeps producer-defined grouping in
parent_uuidwhile making inheritance deterministic and preventing a child from silently
breaking the identity model.
Data model and pipeline
Producer surface (new) —
TrackDescriptor. Add repeated customdimensions to
TrackDescriptor(field numbers are illustrative until theproto change is prepared):
Dimension names, string values, and display names are interned in the
packet-sequence incremental state. The exact
InternedDataentry or sharedinterned-string namespace is an implementation detail to settle with the
proto change. IIDs are sequence-local transport details: trace_processor
resolves them immediately to canonical typed values, and unknown IIDs produce
an import diagnostic. Integer values remain inline.
The message intentionally contains no scope. The dimension is declared on
the track described by the containing
TrackDescriptor; the resolution rulesabove determine its effective descendants. Producers use typed descriptors
for well-known dimensions and cannot declare a custom dimension with a
reserved well-known name.
trace_processor import and resolution. All dimensions use one generic
track-dimension path. Typed process/thread descriptors, GPU descriptors, and
machine import context synthesize the same canonical rows as custom
declarations. A post-import resolver applies process/thread association and
parent_idinheritance, deduplicates repeated identical values, and reports aproducer error for a conflicting inherited value. There are no separate
custom process-, thread-, or GPU-dimension storage paths.
trace_processor query surface. Expose effective dimensions through one
typed long-form relation (public name to be finalized):
A small registry exposes dimension metadata such as the well-known bit and
label template. The generic relation is keyed by
track_id, so slices,counters, GPU work, and other track-backed domains use the same join. Raw
IIDs never escape into SQL.
Stable numbering. Generalize the existing dense machine index into a
deterministic index per
(source trace, machine, dimension name). Values areordered by canonical type and value rather than packet arrival. Tracks with
no machine use a source-trace-global synthetic machine bucket. Merged input
traces retain their source-trace identity for numbering unless future merging
explicitly unifies that dimension.
UI. Factor presentation into: (i) a label helper that maps a canonical
value to its template/numbering/display override; (ii) a generic collapse pass
that emits subtitle labels; and (iii) a details renderer for all effective
dimensions. Machine and custom dimensions use all three. GPU uses the same TP
rows and label helper but retains its specialized group presentation.
Custom dimensions and GPU work
GPU is a well-known dimension with the same trace_processor shape as machine,
CPU, process, thread, or a custom dimension. Its special behavior is limited to
identity merging and UI presentation.
A per-process GPU track is associated with a
upid, so a custom dimensionattached to that process track — for example
rank— resolves onto the GPU trackthrough the ordinary process-association rule. No GPU-specific custom-dimension
code or
_process_dimensionjoin is needed. A cross-process GPU group does notreceive a process dimension when its children disagree:
rankremains on theper-process child tracks while the group carries only dimensions that have one
unambiguous value, such as
gpuandmachine.Querying by a dimension in trace_processor
Dimensions are a first-class query axis for ad-hoc SQL and batch analysis. The
long-form
track_dimensionrelation preserves value types, so joins do not relyon string coercion:
Process and thread remain typed top-level Perfetto concepts. Their canonical
tracks declare or synthesize dimensions, and the resolver propagates those
values to tracks with matching
upid/utid; they do not require parallelper-scope dimension tables. Machine continues to reach global tracks because TP
synthesizes it from each track's import context. A global track receives a custom
dimension only from its own descriptor or an explicit
parent_uuidancestor.Details and non-track surfaces
When a selected slice, counter, or other row has a backing
track_id, the detailspanel resolves and displays that track's effective dimensions. Process/thread
details use their canonical tracks. Each entry shows the canonical name, raw
typed value, and optional display name; well-known dimensions such as machine no
longer need a separate raw-ID-only presentation. This avoids copying custom
columns into every event table while making dimensions consistently visible in
both SQL and details UI.
Migration
label moves from a name suffix to the new subtitle immediately; the machine
table and canonical track names remain.
hierarchy and the hardcoded multi-GPU presentation gate remain, but both GPU
plugins consume the shared dimension metadata and label helper.
edges for effective-dimension resolution. This RFC does not replace
upidorutidstorage.thread, or ordinary track. Adding another custom dimension is then data rather
than new TP/UI code.
Labels still collapse when only one distinct value is present. The deliberate
presentation change is that machine labels move from name suffixes to subtitles.
Alternatives considered
Option 1 — Generic TP dimensions + shared presentation; GPU hierarchy untouched (recommended)
Adopt trace_processor's dimension vocabulary, add the producer surface for custom
dimensions, resolve all effective values into one typed track-keyed relation, and
extract subtitle/details/label helpers shared by machine, custom dimensions, and
GPU presentation. Do not add a generic hierarchy mode; leave GPU's existing
merging-based grouping as specialized presentation.
Pro:
dimensions; no per-domain custom-dimension tables.
parent_uuidtrees.decoupled from canonical names and details UI reads the same source.
Con:
GPU grouping stays specialized until merging work is generalized.
needs a diff-test sweep.
Option 2 — Generic LEVEL presentation mode
Model machine and GPU as instances of one grouping concept with LABEL/LEVEL modes,
and drive GPU's hierarchy through a generic UI grouping pass.
Pro:
Con:
parent_uuidsubtrees. Hierarchy from system-wide concepts belongs intrace_processor merging, not a UI mode.
Option 3 — Add each custom dimension the current way (do nothing generic)
Give a custom dimension (e.g. rank) its own dimension and its own hardcoded
collapse/label gate, like GPU got.
Pro:
Con:
guarantees a fourth.
Future work (non-goals of this RFC)
already merges process/thread across data sources to the initial well-known set
(
machine,gpu,cpu,process,thread), so hierarchy for system-wideconcepts falls out of trace_processor identity rather than any UI grouping mode.
This is the path to eventually folding GPU's specialized grouping into a shared
presentation. Replacing the current typed process/thread storage with the generic
identity model is not required by this RFC.
surfaced (e.g. promote a label to its own subtree) and reorder dimensions at
runtime.
Remaining implementation questions
The review resolves the model-level questions: use
Dimension/dimensions,add subtitles now, reserve
machine/gpu/cpu/process/thread, rejectconflicting inherited values, intern string data, number per source trace and
machine, and expose dimensions in SQL and details UI. Implementation still needs
to settle:
namespaces or an existing generic interned-string entry; define incremental
state reset behavior and validation for unknown IIDs.
whether
declaring_track_id/ inheritance provenance are public or internal.value. Follow TP conventions to decide whether to reject only the declaration
or packet while recording the diagnostic; do not fail the whole trace merely
because one producer emitted an invalid dimension.
component source-trace identifier used with
machine_idfor stable numbering.multiple labels; this is UI implementation detail rather than an optional
feature.
💬 Discussion Guidelines:
All reactions