Skip to content

perf: outpace CPython across the benchmark suite - #83

Merged
owenthcarey merged 96 commits into
mainfrom
perf/outpace-cpython
Oct 5, 2026
Merged

owenthcarey merged 96 commits into
mainfrom
perf/outpace-cpython

Conversation

@owenthcarey

@owenthcarey owenthcarey commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Makes WeavePy faster than CPython 3.14 on 21 of the 24 bench-gate fixtures (with dict_ops at parity), and matches or beats main on every fixture. Across the suite, the geometric mean is 0.50× CPython's time and 0.75× main's. The full report, with method, per-fixture tables, imports, memory, and what still trails, is in docs/PERFORMANCE-CPYTHON-PARITY.md.

What changed

  • Frame compiler (frame_jit): a baseline Cranelift compiler for loops tier 2 leaves behind (method calls, container operations, generator bodies). It admits a loop only when most of its instructions stay native, now including whether each attribute read has hit a field.
  • Tier 2 call paths: native calls, field-update methods, constructor fields, keyword leaf calls, global list lanes, dict.items() loops, object compares, and deque lanes. Cranelift's mid-end stays on for tier 2.
  • Leaf plans: compile thresholds by the rent-or-buy rule, subscripts, in-line caches, and in-place evaluation when native code calls a pure leaf (key=lambda t: t[1]).
  • Native accelerators: pickle (one-pass decode and a new memo), collections.deque, datetime (packed values with a per-class pool), compiled regular expressions and their matches, and map and filter.
  • Exceptions and generators: lazy tracebacks, catching in the quiet loop, and generator fast steps that compile and avoid reference churn.
  • Classes, imports, and startup: class bodies and metaclass attribute reads run directly; native code cache format 3 with lazy column tables; an Fx-hashed intern pool; the opcode shim's tables built on first use.
  • Smaller wins: sized set(), big-int arithmetic in the core loop, isinstance against tuples, cheaper function creation, and keyword leaf calls that reuse their **kwargs dictionary (which also removes an allocator page churn that made Windows call_overhead layout-sensitive).
  • Toolchain and CI: satisfies Rust 1.99's deprecations and Clippy lints under -D warnings; the bench gate builds its merge-base reference with lints capped, since main predates those lints.

Bench gate (local run, macOS x86-64, vs main and CPython 3.14.8)

CI's A/B gates pass on Linux, macOS ARM64, and Windows.

fixture vs main vs CPython
deltablue 0.83 3.03
generators 0.83 1.10
dict_ops 0.83 1.06
pickle_bench 0.61 0.95
float_math 0.77 0.94
deque_ops 0.49 0.88
json_bench 0.82 0.79
list_ops 0.94 0.79
startup 0.94 0.79
fannkuch 0.69 0.75
nbody 0.76 0.74
str_methods 0.59 0.69
call_overhead 0.62 0.69
attr_access 0.59 0.66
pyaes 1.02 0.64
pidigits 0.94 0.61
richards 0.58 0.55
datetime_ops 0.36 0.51
jitkernels 0.62 0.37
fib 0.96 0.36
spectral_norm 0.97 0.20
nested_loops 0.99 0.08
jitloop 0.99 0.08
sumvm 0.99 0.05

Still trailing CPython

  • deltablue (3.0×), generators (1.1×), and dict_ops (0.92 to 1.06× across runs): class attribute reads through an instance, isinstance, and Python-to-Python calls still leave compiled code for general helpers.
  • Imports of large stdlib packages retire 1.3 to 1.8 times CPython's instructions (down 14% to 19% on this branch), and peak memory runs 1.2 to 2.5 times CPython's.

Validation

  • CI: all 28 checks pass, including the blocking A/B bench gates on Linux, macOS ARM64, and Windows, the regression suites on all three platforms, and the free-threaded lanes.
  • Curated CPython 3.14 gate (616 suites): passes except where main fails the same way on the development host (test_capi, test_genexps, test_tempfile, test_utf8_mode, and the multiprocessing spawn and forkserver timeouts).
  • Bundled regrtest: 274 of 274 pass; VM unit tests: 418 of 418; cargo test --workspace --all-targets --all-features, cargo test --workspace --doc, Clippy with -D warnings, cargo fmt --check, the tools' tests, and the artifact check pass.
  • Wrapping up found and fixed these, each with a regression test where it applies:
    • a plan JIT instruction that was invalid IR, which crashed with SIGILL on ARM64;
    • a frame that caught an exception, parked without resetting its handler state;
    • an unsigned underflow in the regex prefix search (debug builds);
    • OSError printed as OSError: 9 in unraisable reports (test_signal);
    • interrupt_main() never landing during call-and-catch recursion (test_threading);
    • the frame compiler admitting loops whose attribute reads always leave native code.

…r round trips

A plain instance (no native payload, `object.__hash__`) now hashes by
identity directly instead of probing the weakref type and re-entering the
interpreter. `LeafProbe` accepts such instances, so `d[obj]` stays in
the core loop, and `member_eq` settles two plain instances without
resolving `__eq__` four times. Instance-keyed dict reads drop from 2,272
to 918 instructions.
… by refcount

The collector's tracked handles and the weakref registry held strong
references, so every drop site had to emulate refcount death: suspect
maps, prompt-reap cascades, parked drops and copy-resurrected finalizers.
Both now hold weak references and objects die when their last reference
goes, as in CPython.

- `WeakObject` gives every heap variant a non-owning handle; collections
  upgrade a snapshot, seed `gc_refs` from strong counts and prune dead
  handles incrementally.
- Death hooks in `Drop` clear weakrefs and queue callbacks. Types without
  a hook are swept at collections.
- `Rc`'s last release of an instance or generator that owes a finalizer
  moves the reference itself into the pending queue, so `__del__` sees
  the same object (identity preserved, no copy).
- A trashcan bounds recursive drops of deep container chains.
- The core loop checks for queued work after the releases that can end
  a binding (stores, pops, container ops, helper calls), at the
  instruction CPython would run the finalizer.
- The JIT constructor fallback no longer finalizes the instance it
  allocated and abandoned.

Removes roughly 5,000 lines of reap machinery. Micro costs drop sharply
(instance construction 5,040 to 2,953 instructions, bound method calls
2,055 to 1,068), and deltablue, nbody, fannkuch and pyaes get 6-17%
faster.
Leaf plans (the frameless evaluator's register form of small methods)
now compile with Cranelift once warm. Registers become SSA variables;
moves, constants, branches, and integer compares and arithmetic run in
line; other operations call helpers that share the interpreter's own
implementation (factored out of the runner). A method call whose site
last resolved a pure leaf compiles the callee's plan in line, guarded
on the function's identity and code, so `self.output().value` runs
without nested evaluations. A guard failure or any undecidable operation
declines exactly as the runner does.

Other changes:
- Property getters run as inline activations in the core loop instead
  of a nested interpreter, with the getter cached per site and class
  version (3,844 to 2,166 instructions per access).
- `getattr(obj, name[, default])` on a plain instance takes a leaf
  lane (3,614 to 1,543 instructions).
- Comparisons of two instances of one class call the class's Python
  dunder directly, without a bound method.
- Tier 2's compile-time probes no longer materialize split instance
  dictionaries, which had left probed instances on the slow dict path.
- Weakref callbacks count toward the shared pending-work gate, so one
  thread's drain can't strand another thread's queued callbacks.
- `_testinternalcapi.instance_layout` reports an instance's attribute
  layout without changing it.

deltablue retires 4.6% fewer instructions.
… fast paths

A returning inline activation released its locals in line only when
every heap local had escaped (a check the old prompt-reap discipline
needed); now any unshared locals vector is released in line, and an
object that dies with it queues its finalizer for the caller's next
check. The exit-reap check is gone, and `inline_finish` takes the same
shortcut.

Native leaf plans gain in-line float compares and float add, subtract,
multiply and divide (NaN results and zero divisors take the helper),
and attribute, method and global helpers that receive their site's
cache slot directly.

deltablue retires 10% fewer instructions than before plan compilation;
a non-leaf method call costs 832 instructions instead of 1,069.
A generator qualified for fast steps only if every instruction in its
body was one a step runs, so a body that builds its iterator first
(`for j in range(k): yield j`) never did. Eligibility now covers what
a resume reaches from a yield; the first resume's step stops at the
prologue and the general loop runs it once. Resuming such a generator
costs 647 instructions instead of 1,187.
…s handlers directly

While a traceback holds a frame's object, the frame can't take the quiet
loop, so its handler ran every instruction through the full prologue.
When nothing that prologue serves is pending, each instruction now only
brings the frame's `lasti` current and steps. An `except` clause naming
the raised exception's own class matches by identity, without walking
either class's MRO.

A raise caught in the same frame costs about 7,000 instructions instead
of 8,375.
…ons per namespace state

Each attribute site in a native plan keeps the split-field positions of
up to four receiver classes, so methods a base class shares (the
constraint classes in deltablue all read `self.direction`) stop missing
the single-class field slot. Each global site remembers which namespace
holds its value and at what index, under the globals' and builtins'
process-unique stamps, and reads the current value there on a hit.

deltablue retires about 3% fewer instructions.
A tier-2 attribute read of a `__slots__` field took the general path
(closures, a guard re-check, a named lookup) every time. Scalar slot
fields now read at their usual position alongside indexed instance
fields. attr_access retires about 5% fewer instructions.
repr(u8) fixes the tag byte at offset 0 and each payload at its natural
alignment, so code compiled at run time can read and write values in
place. The size stays 16 bytes (asserted).
- A cached property whose getter is a pure leaf is evaluated on the
  receiver with no activation (t_property: 2026 -> 1276 instructions
  per read).
- Newborn instances wait in a young set of weak handles instead of being
  entered in the collector's index: most die before the next collection,
  which first hands the survivors over, as do the reflective APIs and the
  finalization pass. Their births still pace gen-0 collections.
- gc_trace::track takes a reference (callers no longer clone to register).
- The cached __del__ verdict no longer pays for the MRO walk's prologue.
Compiled code now reads a pinned instance's field, and stores a scalar
over a scalar field, without calling a helper when the instance's values
are split over its class's shared names: the pin, the guard's class
version, the split block's names and the field's index are checked in
line, and anything else takes the helper as before. The VM measures the
offsets with offset_of! on the owning types and checks them against the
first live instance a probe sees before publishing them
(WEAVEPY_JIT_NO_INLINE_ATTRS=1 turns this off).

A loop of two reads and a write drops from 432 to 171 instructions per
iteration; float_math retires 14% fewer instructions.
A native plan's attribute read now checks its site's first field-cache
entry in line (the receiver's class version, its split values over the
class's names, the cached position) and produces the field's register
words without calling h_field; anything else calls it as before. The
object layout tier 2 measures is shared with the plans, and the plan
runner publishes it at its first instance read. The in-line checks are
grouped so each site adds few blocks.

deltablue: -2.8% instructions per iteration.
Compiled code now reads and writes an int or float element of a pinned
list of that lane, appends a scalar when the list has room, and steps a
list loop over int, float or bool elements without calling a helper:
the pin's lane, the list cell's borrow counter, the index and the
element's tag are checked in line, and anything else takes the helper.
The object layout is now measured on a private instance of object at
the first compile, so code that never reads an instance attribute gets
the in-line paths too. An attribute site with no split position (a
__slots__ member) leaves for its helper before touching the pin.

nbody -9%, float_math -10%, list_ops -4% instructions.
WEAVEPY_JIT_QUICK=1 compiles with Cranelift's quick settings (a tuning
knob: 13-89M fewer instructions of compilation, 6-32% slower code).
…lace

`str` payloads get their own header word: the code-point count, computed
on first use or known at birth. It replaces the thread-local length and
ASCII caches (whose one-slot ASCII cache missed on every new string, and
whose length cache held strong references to up to 1,024 strings), so
`len(s)`, index and slice bounds, and every ASCII check are O(1) reads.
Strings derived from ASCII text (split fields, case mappings, slices,
joins of counted parts, replacements) are born knowing their counts.

The string builders write results straight into the final allocation
through a sequential writer, with no intermediate `String`, zeroing, or
second copy, and copy short pieces inline rather than calling `memcpy`:

- `upper`/`lower`/`casefold`/`title`/`capitalize`/`swapcase` of ASCII
  text map whole buffers with branch-free, vectorized loops.
- `join` of exact strings sizes the result once and stores a one-byte
  separator directly.
- `replace` records its matches in one `memchr` scan and builds an
  exact-size result, returning the receiver itself when nothing matched
  (as CPython does); `split` with a one-byte separator scans with
  `memchr`, and an unsplit receiver is returned as the list's item.
- `split()` presizes for prose, settles most bytes with one comparison,
  and defers its all-`str` result list from GC tracking at any length.
- `%`-formatting borrows an exact `str` template, copies literal runs
  at once, and counts the error index only when raising (which now names
  a non-ASCII conversion character by its code point, as CPython does).

A text or bytes payload's last owner frees it with plain loads while the
reference-count bias holds, instead of `Arc`'s two locked decrements.

str_methods (WORK=15000 vs 3000, retired instructions per unit):
41,839 -> 27,002. Wall time on a loaded machine (min of 9, interleaved):
python3.14 97.2 ms, before 114.0 ms (1.17x), after 67.6 ms (0.70x).
The decoder no longer parses each stream twice (a preflight probe, then
the object build). It builds objects directly and checks the stream's
shape as it goes: a node table records each container's height and
whether a container, tuple, or BUILD already holds it. Only a container
that nothing holds may receive items, so an item can never reach its own
container and heights never propagate. Any reason to return to the Python
unpickler drops what was built; only classes proven to have no finalizer
or other hook are instantiated, so that runs no Python code. Short
streams still check just the dispatch entries they use, after decoding.

A BUILD state that only the memo can still reach gives its storage to
the instance (its memo slot is spent, and reading it falls back), and the
state's keys are interned in place through a per-stream cache. Reachable
containers join the collector's deferred set once the stream is
accepted, the decoder's vectors are reused across calls, and the reader
checks one limit per read.

pickle_bench (macOS, JIT off): about 2.58M to 1.98M retired instructions
per iteration.
Lists and dictionaries are written in place under short borrows: leaves
are saved directly, and only a nested container is cloned so that the
borrow can be released before recursing. A dictionary that changes
meanwhile (a new mutation stamp) returns to the Python pickler, as a
resized list already did.

Strings, bytes, lists, and dicts look up and claim their memo slot with
one table probe. While one thread owns every object, memoized values are
no longer pinned, since nothing else can release them. Identities are
computed in line, frames commit through an in-line check, the empty
argument tuple of NEWOBJ is written directly, slot state reuses one
vector, and the output reservation follows the larger of the last two
outputs.

pickle_bench (macOS, JIT off): about 1.98M to 1.69M retired instructions
per iteration.
A value built in a temporary, with a byte store for its variant, and then
copied with word loads stalls store forwarding, and the decoder did that
on almost every instruction. Each opcode arm now writes its entry straight
into the stack's slot, node numbers are word-sized like the rest of an
entry, a GET copies the memo's stored words, and MEMOIZE rebuilds the
value from its fields.

A MEMOIZE that follows the instruction that pushed the value is taken
without a dispatch, and TUPLE2, MEMOIZE, BUILD (a state with __slots__)
runs without creating the pair, whose memo slot is spent from the start.
A moved state dictionary is reused, with its table's capacity, by the
next EMPTY_DICT: a __dict__ state trades tables with the instance's empty
dictionary, and slot state is stored in one step. SETITEMS inserts string
keys with their cached hash.

The dispatch table is now always checked after decoding, for the
instructions the stream used, and the unchanged class version stands in
for looking up __init__, load, and find_class.

loads of pickle_bench's two values, interleaved runs on a loaded host:
0.91 to 0.72 ms per pair (CPython 3.14: 0.56 ms).
Memo positions live in a linear-probing table keyed by address, whose
entries carry the generation of the call that wrote them, so the next
`dumps` clears it by starting a new generation. A slot is read by its
declaration position first, which is where a plain class keeps it.

dumps of pickle_bench's two values, interleaved runs on a loaded host:
0.51 to 0.44 ms per pair (CPython 3.14: 1.05 ms).
The fused MEMOIZE and the slot-state BUILD peek at the next one or two
bytes of the stream. A comparison against a constant array, byte by
byte and in line, replaces a slice comparison that called memcmp on
every pushed value.
A baseline compiler for the core loop: a code object whose hot loop runs
mostly in instructions it handles is compiled with Cranelift into one
function over its basic blocks, entered at any block start. It keeps a
virtual operand stack (locals and constants are referenced, not copied;
int, float and bool arithmetic and compares stay unboxed), reads cached
instance fields in place, runs local method calls the core loop would
run in place through a helper, and hands every other instruction back to
the core loop with the frame exactly as the loop expects it.

Admission keeps compile costs off code that wouldn't gain: only a loop
that got hot compiles, and only when its instructions run natively at
least eight to one against hand-overs, counting a local method call as
native only when its site resolved to a builtin or a leaf. Cranelift
runs without its optimizer and with the single-pass allocator, since
the lowering does the optimizing that pays.

Whole-run instructions, frame JIT off vs on: deque_ops -15%, dict_ops
-8%, datetime_ops -4%, generators -3%, deltablue +1%, the rest within
noise.
A plan compiled after 64 evaluations: each compile costs Cranelift
millions of instructions, and a small method's native code saves tens
per call, so short runs paid for code they never amortized (deltablue
compiled a dozen plans for about 110M instructions). Plans now compile
after 20,000 evaluations, the rent-or-buy point.

deltablue: -12% at the default work, -2% at 4x; long runs of every
fixture within noise.
A method whose body is `self.n += k; return self.n` now runs in compiled
code without a call once wpjit_call_method has run it through the helper:
the helper arms the method entry with the field's split index, the most
split values a receiver can hold without shadowing the method's name, the
increment's form and the function's code pointer, and compiled code checks
those, the receiver's class version and split names, and the observer,
dict-watcher and exotic-key gates in line before the store. Anything else
takes the helper as before.

A list loop, a list read or an object-lane method result that pins an
object an activation already pinned reuses that pin (a 16-entry memo by
address, checked against the pin it names), so a loop over the same few
objects no longer grows the pin table until the pin-pressure exit.

Per-unit instructions: richards 4,465 -> 1,851 (the memo alone: 4,010),
call_overhead 6,473 -> 6,162, attr_access 3,023 -> 2,783.
guards_hold settles the namespaces' stamps in line and leaves the
name-by-name revalidation, the callee table and the class-constructor
probes out of line; try_native_call fills the cached child context's pin
table in place instead of moving it in and out, reads the pure-leaf hint
before calling the classifier, and note_callee_exit keeps its round-trip
accounting out of line. A list built from a short range is deferred with
the collector without scanning its integers first.

Per-unit instructions: fannkuch 1,759 -> 1,562, float_math 7,563 ->
7,347, attr_access 2,783 -> 2,713.
A call that leaves trailing or keyword-skipped parameters out of a
compiled scalar leaf now enters it directly from native code too: the
leaf carries the function's scalar defaults (each in its parameter's
lane), keyword values land in the slots the site's permutation names,
and the rest bind the burned-in defaults. The caller's guard snapshot
lists every callee whose defaults it may have burned in, and guards_hold
fails once one of them has its `__defaults__` replaced; a declined entry
or a deopting leaf still runs the ordinary call from the start.

call_overhead's per-unit instructions drop from 6,183 to 4,952.
A method that adds to one of its own `__slots__` members and returns it
now gets the same callback-free update plan as one over an instance-dict
field: the member's slot descriptor can't acquire hooks the class version
doesn't track, so the store guard certifies it like a plain class value.
wpjit_call_method then runs such a method through
native_scalar_field_update instead of a framed native call.

attr_access's per-unit instructions drop from 2,713 to 1,799.
try_native_call keeps its common path small: the under-arity default
check, the scalar-leaf entry, a fresh child context and everything after
a deopt or an object result move out of line, and a clean scalar return
puts the child back without the general result translation. Framed
activations reuse the last finished activation's pin table allocation, so
a loop that leaves at the pin limit doesn't regrow it from empty.

guards_hold remembers, per burned-in class constructor, the class version
at which the full construction probe last held; an unchanged class then
rechecks only its metaclass and its `__init__` code instead of rebuilding
the probe after every call that ran Python.

Compiled code no longer writes back, at every exit and call, the locals
its body never assigns: the frame already holds them, and keeping each
one live to every exit only cost registers and compile time.
WEAVEPY_JIT_CLIF_STATS=1 prints each compiled function's size.

Per-unit instructions: call_overhead 4,952 -> 4,671, fannkuch 1,562 ->
1,510, float_math 7,347 -> 7,291; compile times drop 3-17% (maximize
8.2 -> 6.8 ms).
A dynamic keyword call of a warm pure-leaf function (or a bound method
over one) now binds its marshaled values straight onto the leaf's
argument list through the site's cached CallPyKwNames permutation,
building any `**kwargs` dictionary in its cell, and delivers the result
without revalidating the activation's guards (a pure leaf runs no Python).
It used to stage the values in scratch vectors, fill a full locals vector
and move the filled dictionary into its cell three times.

A loop of `f(i, delta=2)` calls into `def f(a, **kw)` drops from 312 to
244 ns per iteration (CPython: 149); call_overhead's per-unit instructions
drop from 4,671 to 3,953. WEAVEPY_JIT_CLIF_STATS=1 also prints Cranelift's
pass timings.
…e slow exit

A new-key attribute store (the constructor pattern, `self.x = ...` in
`__init__`) of a scalar now appends the field in compiled code when the
fresh instance's split values already hold exactly the class's names
before it and have room: its guard carries the name's position among the
class's shared names, and the class version, the block's names and
length and capacity, the native body and the dict-watcher gate are
checked in line before anything is written. The first store of an
instance still allocates the values through the helper.

A native callee that returns `None` on the object lane (a procedure, a
constructor's `__init__`) or an instance where the site takes an object
(a method returning `self`) now puts its context back without going
through the general result translation.

float_math's per-unit instructions drop from 7,291 to 6,834.
The lowering already emits the code it wants (its guards' loads aren't
pure, so the e-graph pass merges almost nothing), and opt_level "none"
keeps the backtracking register allocator. Measured over the five call
fixtures, per-unit instructions are unchanged or slightly lower and the
compile-inclusive totals drop 3-6M instructions; wall time vs CPython
improves by 1-4 points (float_math 0.981 -> 0.959, richards 0.494 ->
0.458). WEAVEPY_JIT_OPT=1 restores opt_level "speed".
Two costs every constructor or call of these shapes paid:

- A `__slots__` class's `__init__` (`self.x = x; self.y = y`) never ran
  as a frameless leaf: the leaf store path only knew instance dicts, so
  each evaluation declined until the body was demoted to an inline
  activation, whose stores then each left the core loop for the helper.
  The core loop's store, and a leaf's buffered or constructor store,
  now write a validated slot member (the site's `StoreAttrSlot`, keyed
  by the class version) directly, as the helper does.
- The inline call and return pushed the activation, the parked slot and
  the result through out-of-line `Vec::push` calls, emptied the locals
  through a `Drain`, and called two small helpers out of line. They now
  push with the room check in line and release the locals in place.

Measured (retired instructions per call, a local call of each case):
`P(i, 2)` for a two-slot class 6,219 -> 4,027 (CPython 1,460);
`p = P(i, 2); return p.x + p.y` 7,098 -> 4,891 (CPython 1,762); a call of
a one-line non-leaf function from a loop 2,179 -> 2,044, a method call
2,493 -> 2,362; `with cm` 4,426 -> 4,213. Deltablue 33.18M -> 32.62M
per unit (32.07M -> 31.50M once warm, 40 -> 120 units).
…riables frameless

Two more costs of closures and functions made in a loop:

- Every new function copied its code's template of name slots
  (`__module__`, `__name__`, `__qualname__`): a small dict built, three
  keys and values cloned, then freed with the function. A function now
  holds the shared template as a seed and copies it into its own slots
  only when something reads or writes them; most functions never do.
- A function with free variables (`lambda x: x * k`, a nested `def` that
  reads its maker's locals) was never a leaf, so each call took an
  inline activation. Leaf plans now read a free variable straight from
  the function's closure cell (a leaf has no `STORE_DEREF`, so nothing
  can rebind it during the evaluation; an empty cell declines to the
  full path, which raises), and such functions are frameless leaves.

Measured (retired instructions per call, a local call of each case): a
call of an existing closure `g(i)` 2,894 -> 1,466 (CPython 920);
`mk(i)(1)` 8,544 -> 6,574 (CPython 2,502); building a nested `def`
4,926 -> 4,152; `lambda` built 3,514 -> 2,806, built and called 4,275 ->
3,490 (CPython 1,549). Importing argparse, typing, dataclasses,
email.message, json, unittest, logging and decimal 727M -> 716M.
Deltablue is unchanged within noise.
Tier 2 compiles code during start-up and imports only after sixteen times
its normal threshold, since that code is mostly run once; the frame
compiler had no such budget. Importing several modules that each compile
regular expressions made `re._compiler._optimize_charset` hot, and
compiling its 559 instructions cost about 100M instructions, a tenth of
the whole import. Code that gets hot while start-up or an import runs now
compiles only after sixteen times the frame compiler's threshold too.

`import typing, json, dataclasses, argparse, logging, asyncio, inspect,
textwrap, http.client, unittest` in one process: 978M -> 874M
instructions (python3.14: 447M). Single-module imports are unchanged.
`_opcode` (imported by `opcode`, and so by `dis`, `inspect` and
`asyncio`) built a lambda per opcode for its popped and pushed counts
and a set per flag predicate when it was imported, though most programs
never ask for a stack effect. The generator now emits the stack tables
behind `_stack_tables()`, built on the first `stack_effect` call, with
a plain int for every count that doesn't depend on the oparg, and the
predicates test literal frozensets of the opcodes that have each flag.
`_weave_flowgraph` reads the tables the same way.

Import-only instructions: opcode 27M -> 13M (python3.14 3M), dis 57M ->
43M (20M). The regenerated module answers all 21,210 checks of every
predicate and stack effect the same as before.
Classes built by a metaclass (`enum`, `abc`, `typing`) spent most of
their time outside the class body's own code:

- A `__prepare__` namespace with a Python `__setitem__` (enum's
  `EnumDict`) took a bound method and a fresh, unhashed key string per
  `STORE_NAME`. The function is now called with the namespace as its
  first argument and the code's interned name as the key; `LOAD_NAME`
  probes such a namespace without building a key object. `__set_name__`
  hooks that are plain functions are called the same way.
- Zero-argument `super().name` handed every name it couldn't resolve
  to a plain function (dunders such as `super().__setitem__` and
  `super().__new__`, and methods of built-in bases) to the full path,
  which builds a `super` proxy instance, introspects the frame and runs
  the generic attribute lookup on the proxy. When nothing observes the
  call and the receiver passes the basic `supercheck`, `LOAD_SUPER_ATTR`
  now walks the MRO from the class and receiver already on the stack,
  sharing the walk with the proxy's own lookup.
- Attribute reads on a class whose metaclass is a Python subclass of
  `type` with the default `__getattribute__` went to the full path.
  When the metaclass's MRO has no entry for the name, the class's MRO
  answers exactly as for a plain metaclass, so the leaf paths take them
  (uncached by site, as the class's version doesn't follow the
  metaclass). `cls.__dict__` and class-level dicts and lists
  (`_member_map_`, `_member_names_`) read in the leaf paths too.

Creating a three-member `Enum` costs 1.82M -> 1.63M instructions
(python3.14 0.37M), an `IntFlag` 2.56M -> 2.27M (0.54M), and a class
whose metaclass `__prepare__` returns a dict subclass 216K -> 151K
(83K). Import-only instructions, before -> after: typing 67M -> 63M,
json 86M -> 84M, dataclasses 186M -> 182M, argparse 80M -> 77M,
logging 281M -> 276M, asyncio 447M -> 432M, inspect 183M -> 176M,
http.client 228M -> 218M; the ten modules imported together 680M ->
670M (python3.14 447M).
CI builds with `-D warnings` on the current stable toolchain, which now
deprecates the `std::u64` module (used by the vendored lzma-sys) and
`AtomicU64::fetch_update`, and adds Clippy lints that flagged code on
this branch. The vendored crate uses the associated constant, the type
version counter its own compare-exchange loop (`try_update` is newer
than the MSRV), and the workspace allows the new `assert_is_empty`
lint, which fires on tests throughout. The rest is `cargo fmt` and
`cargo clippy --fix`.
Compiling tier-2 code without the mid-end optimizer saved a few percent
of compile time but cost numeric loops dearly: the A/B bench gate
against main showed `jitloop` 3.7x and `spectral_norm` 1.36x slower.
Tier 2 optimizes again by default (`WEAVEPY_JIT_OPT=0` turns it off);
leaf plans stay unoptimized unless `WEAVEPY_JIT_OPT=1`.

Bench gate vs main (macOS, A/B): geometric mean 0.77x, no fixture
regressed; 0.52x CPython.
…iles

The frame compiler's admission rule counted every `LOAD_ATTR` as native,
but its code serves only split-layout fields and natively served
instances' fields. A loop reading a class attribute through an instance
(`o.K`) compiled anyway and left for the core loop on every iteration,
which cost about 1,000 instructions per read against the interpreter's
550. A site now counts as native only once the core loop has recorded a
field there.
`core_return`'s in-line return parked every clean-stack activation through
`inline_park_clean`, including one that caught an exception and so had
its handler state set. That path skips resetting the frame's handler,
saved-exception, and frame-object fields, and its debug assertion failed
in every debug build that imported the standard library. Such a frame
now goes through `inline_park`, as `inline_deliver` already does.
When the literal prefix is skipped in full, `prefix_len - prefix_skip - 1`
is -1. CPython computes it with signed sizes; the Rust port overflowed,
which panicked in debug builds (release builds wrapped to the right
value).
Natively built exceptions no longer store a separate message, so the
report's Rust-side rendering fell back to `args[0]` and printed a failed
system call's error as `OSError: 9`. `exception_message` now uses
`OSError.__str__`'s errno form, shared with `oserror_str`, which fixes
`test_signal`'s wakeup-fd write report.
…tion

The quiet loop resumed a handler without the per-instruction prologue's
GIL countdown and hot-gate check. Code that only calls and catches, such
as recursion retried from its `RecursionError` handler, then never gave
up the GIL, so `_thread.interrupt_main()` from another thread never
landed and `test_threading` timed out. Every catch site now takes the
prologue's checks, as the full step's already did.
The parity report now carries the bench gate's A/B results against
`main` and CPython 3.14.8, current import and memory costs, per-operation
costs for what still trails, and the validation the branch passed.
@owenthcarey
owenthcarey force-pushed the perf/outpace-cpython branch from aed1889 to 46a97fc Compare October 3, 2026 09:32
@owenthcarey owenthcarey changed the title perf: outpace CPython across the benchmark suite (WIP) perf: outpace CPython across the benchmark suite Oct 3, 2026
@owenthcarey
owenthcarey marked this pull request as ready for review October 3, 2026 09:32
The A/B gate builds the merge-base with the workflow's `-D warnings`, so
a lint a newer stable toolchain added fails the reference build before
anything is timed: Rust 1.99 deprecates `std::u64`, which `main`'s
vendored lzma-sys still uses. The reference binary is only timed, so its
build now caps lints at warnings.
`inline_method` widened the split block's length with `uextend.i64`, but
`uload32` already yields an `i64`. Release builds skip Cranelift's
verifier; x86-64 tolerated the instruction, but AArch64 encoded an
invalid one, so compiled plans with in-line method loads died with
SIGILL on macOS ARM (`test_jit_plan_caches`, `deltablue`). A new
regression test runs the plan-cache tests with both JIT verifiers on.
The Windows gate has measured `call_overhead` at 1.2x the merge-base on
every PR run, while the same build, environment, and gate pass at 0.65x
on separately dispatched runners. `bench_cold_jit.py` now takes its
fixture list as an option, and the failure-only Windows diagnostic adds
`call_overhead`, so the next failure records its timings and JIT traces
from inside the failing job.
In the last failing run, the gate timed the PR binary's
`call_overhead` at 89 ms while the cold-JIT diagnostic, in the same job,
timed it at 38 ms (the merge-base and CPython agreed between the two at
60 ms). A failure-only step now times direct runs that each change one
thing the two differ in: the environment, the caches, and the binary
path.
…e fails

Direct runs in the failing job time the PR binary's `call_overhead` at
0.65x the merge-base while the gate times it at 1.2x, whatever the
environment, caches, or binary path. The failure-only step now also
times it with an absolute fixture path, launched from Python, and through
`weavepy-bench run` with each binary alone.
… fails

The last failing run pinned the slowdown to the fixture path: the PR
binary times `call_overhead` at 75 ms from a relative path and 138 ms
from an absolute one, however it's launched, while the merge-base is
unaffected. The failure-only step now separates the caches involved and
records JIT and burst statistics for both paths.
The slow configuration (an absolute fixture path with warm caches) has
the same tier-2 compiles as the fast one, and any change to the caches
or instrumentation makes it fast again. The failure-only step now runs
that configuration with each tier switched off or reconfigured, twice
each, to show which one it depends on.
Every JIT tier switch leaves the absolute-path run slow, and with the JIT
off both paths time alike, so the difference is in how compiled code
runs rather than what compiles. The failure-only step now records the
VM's tier-2 counters (native calls, fallbacks, deopts) for both paths.
The slow and fast configurations run the same tier-2 paths, and every
in-process statistic tried so far shifts the layout enough to hide the
slowdown. The failure-only step now records sampling profiles of both
from outside the process and uploads them with the diagnostic artifact.
The slow run's profile spends a fifth of its samples in an unnamed
cluster of `python314.dll` functions, which samply couldn't symbolicate
on the runner. The failure-only step now uploads the binaries and their
PDBs beside the profiles so they can be symbolicated offline.
Each tier-2 keyword call of a pure leaf that collects `**kwargs`
allocated a dictionary and freed it on return. When those blocks were
alone on their mimalloc page, the page went back to the arena on every
free and was fetched fresh on the next call. A Windows profile of the
bench gate's `call_overhead` put a quarter of its samples there, which
made it 1.2x the merge-base whenever the heap happened to fall that way.

The emptied dictionary of a call that nothing kept now serves the next
call; clearing restamps it, so the callee still sees a new dictionary.
A regression test checks that returned, stored, and captured
dictionaries stay intact and that no keys leak between calls.
The Windows bench gate passes now that keyword leaf calls reuse their
`**kwargs` dictionary, so the failure-only step that profiled and timed
the slowdown has done its job. The cold-JIT diagnostic keeps
`call_overhead` in its fixture list.
@owenthcarey
owenthcarey merged commit 18a9ea4 into main Oct 5, 2026
28 checks passed
@owenthcarey
owenthcarey deleted the perf/outpace-cpython branch October 5, 2026 03:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant