Repository navigation
perf: outpace CPython across the benchmark suite - #83
Merged
Merged
Conversation
…r round trips A plain instance (no native payload, `object.__hash__`) now hashes by identity directly instead of probing the weakref type and re-entering the interpreter. `LeafProbe` accepts such instances, so `d[obj]` stays in the core loop, and `member_eq` settles two plain instances without resolving `__eq__` four times. Instance-keyed dict reads drop from 2,272 to 918 instructions.
… by refcount The collector's tracked handles and the weakref registry held strong references, so every drop site had to emulate refcount death: suspect maps, prompt-reap cascades, parked drops and copy-resurrected finalizers. Both now hold weak references and objects die when their last reference goes, as in CPython. - `WeakObject` gives every heap variant a non-owning handle; collections upgrade a snapshot, seed `gc_refs` from strong counts and prune dead handles incrementally. - Death hooks in `Drop` clear weakrefs and queue callbacks. Types without a hook are swept at collections. - `Rc`'s last release of an instance or generator that owes a finalizer moves the reference itself into the pending queue, so `__del__` sees the same object (identity preserved, no copy). - A trashcan bounds recursive drops of deep container chains. - The core loop checks for queued work after the releases that can end a binding (stores, pops, container ops, helper calls), at the instruction CPython would run the finalizer. - The JIT constructor fallback no longer finalizes the instance it allocated and abandoned. Removes roughly 5,000 lines of reap machinery. Micro costs drop sharply (instance construction 5,040 to 2,953 instructions, bound method calls 2,055 to 1,068), and deltablue, nbody, fannkuch and pyaes get 6-17% faster.
Leaf plans (the frameless evaluator's register form of small methods) now compile with Cranelift once warm. Registers become SSA variables; moves, constants, branches, and integer compares and arithmetic run in line; other operations call helpers that share the interpreter's own implementation (factored out of the runner). A method call whose site last resolved a pure leaf compiles the callee's plan in line, guarded on the function's identity and code, so `self.output().value` runs without nested evaluations. A guard failure or any undecidable operation declines exactly as the runner does. Other changes: - Property getters run as inline activations in the core loop instead of a nested interpreter, with the getter cached per site and class version (3,844 to 2,166 instructions per access). - `getattr(obj, name[, default])` on a plain instance takes a leaf lane (3,614 to 1,543 instructions). - Comparisons of two instances of one class call the class's Python dunder directly, without a bound method. - Tier 2's compile-time probes no longer materialize split instance dictionaries, which had left probed instances on the slow dict path. - Weakref callbacks count toward the shared pending-work gate, so one thread's drain can't strand another thread's queued callbacks. - `_testinternalcapi.instance_layout` reports an instance's attribute layout without changing it. deltablue retires 4.6% fewer instructions.
… fast paths A returning inline activation released its locals in line only when every heap local had escaped (a check the old prompt-reap discipline needed); now any unshared locals vector is released in line, and an object that dies with it queues its finalizer for the caller's next check. The exit-reap check is gone, and `inline_finish` takes the same shortcut. Native leaf plans gain in-line float compares and float add, subtract, multiply and divide (NaN results and zero divisors take the helper), and attribute, method and global helpers that receive their site's cache slot directly. deltablue retires 10% fewer instructions than before plan compilation; a non-leaf method call costs 832 instructions instead of 1,069.
A generator qualified for fast steps only if every instruction in its body was one a step runs, so a body that builds its iterator first (`for j in range(k): yield j`) never did. Eligibility now covers what a resume reaches from a yield; the first resume's step stops at the prologue and the general loop runs it once. Resuming such a generator costs 647 instructions instead of 1,187.
…s handlers directly While a traceback holds a frame's object, the frame can't take the quiet loop, so its handler ran every instruction through the full prologue. When nothing that prologue serves is pending, each instruction now only brings the frame's `lasti` current and steps. An `except` clause naming the raised exception's own class matches by identity, without walking either class's MRO. A raise caught in the same frame costs about 7,000 instructions instead of 8,375.
…ons per namespace state Each attribute site in a native plan keeps the split-field positions of up to four receiver classes, so methods a base class shares (the constraint classes in deltablue all read `self.direction`) stop missing the single-class field slot. Each global site remembers which namespace holds its value and at what index, under the globals' and builtins' process-unique stamps, and reads the current value there on a hit. deltablue retires about 3% fewer instructions.
A tier-2 attribute read of a `__slots__` field took the general path (closures, a guard re-check, a named lookup) every time. Scalar slot fields now read at their usual position alongside indexed instance fields. attr_access retires about 5% fewer instructions.
repr(u8) fixes the tag byte at offset 0 and each payload at its natural alignment, so code compiled at run time can read and write values in place. The size stays 16 bytes (asserted).
- A cached property whose getter is a pure leaf is evaluated on the receiver with no activation (t_property: 2026 -> 1276 instructions per read). - Newborn instances wait in a young set of weak handles instead of being entered in the collector's index: most die before the next collection, which first hands the survivors over, as do the reflective APIs and the finalization pass. Their births still pace gen-0 collections. - gc_trace::track takes a reference (callers no longer clone to register). - The cached __del__ verdict no longer pays for the MRO walk's prologue.
Compiled code now reads a pinned instance's field, and stores a scalar over a scalar field, without calling a helper when the instance's values are split over its class's shared names: the pin, the guard's class version, the split block's names and the field's index are checked in line, and anything else takes the helper as before. The VM measures the offsets with offset_of! on the owning types and checks them against the first live instance a probe sees before publishing them (WEAVEPY_JIT_NO_INLINE_ATTRS=1 turns this off). A loop of two reads and a write drops from 432 to 171 instructions per iteration; float_math retires 14% fewer instructions.
A native plan's attribute read now checks its site's first field-cache entry in line (the receiver's class version, its split values over the class's names, the cached position) and produces the field's register words without calling h_field; anything else calls it as before. The object layout tier 2 measures is shared with the plans, and the plan runner publishes it at its first instance read. The in-line checks are grouped so each site adds few blocks. deltablue: -2.8% instructions per iteration.
Compiled code now reads and writes an int or float element of a pinned list of that lane, appends a scalar when the list has room, and steps a list loop over int, float or bool elements without calling a helper: the pin's lane, the list cell's borrow counter, the index and the element's tag are checked in line, and anything else takes the helper. The object layout is now measured on a private instance of object at the first compile, so code that never reads an instance attribute gets the in-line paths too. An attribute site with no split position (a __slots__ member) leaves for its helper before touching the pin. nbody -9%, float_math -10%, list_ops -4% instructions. WEAVEPY_JIT_QUICK=1 compiles with Cranelift's quick settings (a tuning knob: 13-89M fewer instructions of compilation, 6-32% slower code).
…lace `str` payloads get their own header word: the code-point count, computed on first use or known at birth. It replaces the thread-local length and ASCII caches (whose one-slot ASCII cache missed on every new string, and whose length cache held strong references to up to 1,024 strings), so `len(s)`, index and slice bounds, and every ASCII check are O(1) reads. Strings derived from ASCII text (split fields, case mappings, slices, joins of counted parts, replacements) are born knowing their counts. The string builders write results straight into the final allocation through a sequential writer, with no intermediate `String`, zeroing, or second copy, and copy short pieces inline rather than calling `memcpy`: - `upper`/`lower`/`casefold`/`title`/`capitalize`/`swapcase` of ASCII text map whole buffers with branch-free, vectorized loops. - `join` of exact strings sizes the result once and stores a one-byte separator directly. - `replace` records its matches in one `memchr` scan and builds an exact-size result, returning the receiver itself when nothing matched (as CPython does); `split` with a one-byte separator scans with `memchr`, and an unsplit receiver is returned as the list's item. - `split()` presizes for prose, settles most bytes with one comparison, and defers its all-`str` result list from GC tracking at any length. - `%`-formatting borrows an exact `str` template, copies literal runs at once, and counts the error index only when raising (which now names a non-ASCII conversion character by its code point, as CPython does). A text or bytes payload's last owner frees it with plain loads while the reference-count bias holds, instead of `Arc`'s two locked decrements. str_methods (WORK=15000 vs 3000, retired instructions per unit): 41,839 -> 27,002. Wall time on a loaded machine (min of 9, interleaved): python3.14 97.2 ms, before 114.0 ms (1.17x), after 67.6 ms (0.70x).
The decoder no longer parses each stream twice (a preflight probe, then the object build). It builds objects directly and checks the stream's shape as it goes: a node table records each container's height and whether a container, tuple, or BUILD already holds it. Only a container that nothing holds may receive items, so an item can never reach its own container and heights never propagate. Any reason to return to the Python unpickler drops what was built; only classes proven to have no finalizer or other hook are instantiated, so that runs no Python code. Short streams still check just the dispatch entries they use, after decoding. A BUILD state that only the memo can still reach gives its storage to the instance (its memo slot is spent, and reading it falls back), and the state's keys are interned in place through a per-stream cache. Reachable containers join the collector's deferred set once the stream is accepted, the decoder's vectors are reused across calls, and the reader checks one limit per read. pickle_bench (macOS, JIT off): about 2.58M to 1.98M retired instructions per iteration.
Lists and dictionaries are written in place under short borrows: leaves are saved directly, and only a nested container is cloned so that the borrow can be released before recursing. A dictionary that changes meanwhile (a new mutation stamp) returns to the Python pickler, as a resized list already did. Strings, bytes, lists, and dicts look up and claim their memo slot with one table probe. While one thread owns every object, memoized values are no longer pinned, since nothing else can release them. Identities are computed in line, frames commit through an in-line check, the empty argument tuple of NEWOBJ is written directly, slot state reuses one vector, and the output reservation follows the larger of the last two outputs. pickle_bench (macOS, JIT off): about 1.98M to 1.69M retired instructions per iteration.
A value built in a temporary, with a byte store for its variant, and then copied with word loads stalls store forwarding, and the decoder did that on almost every instruction. Each opcode arm now writes its entry straight into the stack's slot, node numbers are word-sized like the rest of an entry, a GET copies the memo's stored words, and MEMOIZE rebuilds the value from its fields. A MEMOIZE that follows the instruction that pushed the value is taken without a dispatch, and TUPLE2, MEMOIZE, BUILD (a state with __slots__) runs without creating the pair, whose memo slot is spent from the start. A moved state dictionary is reused, with its table's capacity, by the next EMPTY_DICT: a __dict__ state trades tables with the instance's empty dictionary, and slot state is stored in one step. SETITEMS inserts string keys with their cached hash. The dispatch table is now always checked after decoding, for the instructions the stream used, and the unchanged class version stands in for looking up __init__, load, and find_class. loads of pickle_bench's two values, interleaved runs on a loaded host: 0.91 to 0.72 ms per pair (CPython 3.14: 0.56 ms).
Memo positions live in a linear-probing table keyed by address, whose entries carry the generation of the call that wrote them, so the next `dumps` clears it by starting a new generation. A slot is read by its declaration position first, which is where a plain class keeps it. dumps of pickle_bench's two values, interleaved runs on a loaded host: 0.51 to 0.44 ms per pair (CPython 3.14: 1.05 ms).
The fused MEMOIZE and the slot-state BUILD peek at the next one or two bytes of the stream. A comparison against a constant array, byte by byte and in line, replaces a slice comparison that called memcmp on every pushed value.
A baseline compiler for the core loop: a code object whose hot loop runs mostly in instructions it handles is compiled with Cranelift into one function over its basic blocks, entered at any block start. It keeps a virtual operand stack (locals and constants are referenced, not copied; int, float and bool arithmetic and compares stay unboxed), reads cached instance fields in place, runs local method calls the core loop would run in place through a helper, and hands every other instruction back to the core loop with the frame exactly as the loop expects it. Admission keeps compile costs off code that wouldn't gain: only a loop that got hot compiles, and only when its instructions run natively at least eight to one against hand-overs, counting a local method call as native only when its site resolved to a builtin or a leaf. Cranelift runs without its optimizer and with the single-pass allocator, since the lowering does the optimizing that pays. Whole-run instructions, frame JIT off vs on: deque_ops -15%, dict_ops -8%, datetime_ops -4%, generators -3%, deltablue +1%, the rest within noise.
A plan compiled after 64 evaluations: each compile costs Cranelift millions of instructions, and a small method's native code saves tens per call, so short runs paid for code they never amortized (deltablue compiled a dozen plans for about 110M instructions). Plans now compile after 20,000 evaluations, the rent-or-buy point. deltablue: -12% at the default work, -2% at 4x; long runs of every fixture within noise.
A method whose body is `self.n += k; return self.n` now runs in compiled code without a call once wpjit_call_method has run it through the helper: the helper arms the method entry with the field's split index, the most split values a receiver can hold without shadowing the method's name, the increment's form and the function's code pointer, and compiled code checks those, the receiver's class version and split names, and the observer, dict-watcher and exotic-key gates in line before the store. Anything else takes the helper as before. A list loop, a list read or an object-lane method result that pins an object an activation already pinned reuses that pin (a 16-entry memo by address, checked against the pin it names), so a loop over the same few objects no longer grows the pin table until the pin-pressure exit. Per-unit instructions: richards 4,465 -> 1,851 (the memo alone: 4,010), call_overhead 6,473 -> 6,162, attr_access 3,023 -> 2,783.
guards_hold settles the namespaces' stamps in line and leaves the name-by-name revalidation, the callee table and the class-constructor probes out of line; try_native_call fills the cached child context's pin table in place instead of moving it in and out, reads the pure-leaf hint before calling the classifier, and note_callee_exit keeps its round-trip accounting out of line. A list built from a short range is deferred with the collector without scanning its integers first. Per-unit instructions: fannkuch 1,759 -> 1,562, float_math 7,563 -> 7,347, attr_access 2,783 -> 2,713.
A call that leaves trailing or keyword-skipped parameters out of a compiled scalar leaf now enters it directly from native code too: the leaf carries the function's scalar defaults (each in its parameter's lane), keyword values land in the slots the site's permutation names, and the rest bind the burned-in defaults. The caller's guard snapshot lists every callee whose defaults it may have burned in, and guards_hold fails once one of them has its `__defaults__` replaced; a declined entry or a deopting leaf still runs the ordinary call from the start. call_overhead's per-unit instructions drop from 6,183 to 4,952.
A method that adds to one of its own `__slots__` members and returns it now gets the same callback-free update plan as one over an instance-dict field: the member's slot descriptor can't acquire hooks the class version doesn't track, so the store guard certifies it like a plain class value. wpjit_call_method then runs such a method through native_scalar_field_update instead of a framed native call. attr_access's per-unit instructions drop from 2,713 to 1,799.
try_native_call keeps its common path small: the under-arity default check, the scalar-leaf entry, a fresh child context and everything after a deopt or an object result move out of line, and a clean scalar return puts the child back without the general result translation. Framed activations reuse the last finished activation's pin table allocation, so a loop that leaves at the pin limit doesn't regrow it from empty. guards_hold remembers, per burned-in class constructor, the class version at which the full construction probe last held; an unchanged class then rechecks only its metaclass and its `__init__` code instead of rebuilding the probe after every call that ran Python. Compiled code no longer writes back, at every exit and call, the locals its body never assigns: the frame already holds them, and keeping each one live to every exit only cost registers and compile time. WEAVEPY_JIT_CLIF_STATS=1 prints each compiled function's size. Per-unit instructions: call_overhead 4,952 -> 4,671, fannkuch 1,562 -> 1,510, float_math 7,347 -> 7,291; compile times drop 3-17% (maximize 8.2 -> 6.8 ms).
A dynamic keyword call of a warm pure-leaf function (or a bound method over one) now binds its marshaled values straight onto the leaf's argument list through the site's cached CallPyKwNames permutation, building any `**kwargs` dictionary in its cell, and delivers the result without revalidating the activation's guards (a pure leaf runs no Python). It used to stage the values in scratch vectors, fill a full locals vector and move the filled dictionary into its cell three times. A loop of `f(i, delta=2)` calls into `def f(a, **kw)` drops from 312 to 244 ns per iteration (CPython: 149); call_overhead's per-unit instructions drop from 4,671 to 3,953. WEAVEPY_JIT_CLIF_STATS=1 also prints Cranelift's pass timings.
…e slow exit A new-key attribute store (the constructor pattern, `self.x = ...` in `__init__`) of a scalar now appends the field in compiled code when the fresh instance's split values already hold exactly the class's names before it and have room: its guard carries the name's position among the class's shared names, and the class version, the block's names and length and capacity, the native body and the dict-watcher gate are checked in line before anything is written. The first store of an instance still allocates the values through the helper. A native callee that returns `None` on the object lane (a procedure, a constructor's `__init__`) or an instance where the site takes an object (a method returning `self`) now puts its context back without going through the general result translation. float_math's per-unit instructions drop from 7,291 to 6,834.
The lowering already emits the code it wants (its guards' loads aren't pure, so the e-graph pass merges almost nothing), and opt_level "none" keeps the backtracking register allocator. Measured over the five call fixtures, per-unit instructions are unchanged or slightly lower and the compile-inclusive totals drop 3-6M instructions; wall time vs CPython improves by 1-4 points (float_math 0.981 -> 0.959, richards 0.494 -> 0.458). WEAVEPY_JIT_OPT=1 restores opt_level "speed".
Two costs every constructor or call of these shapes paid: - A `__slots__` class's `__init__` (`self.x = x; self.y = y`) never ran as a frameless leaf: the leaf store path only knew instance dicts, so each evaluation declined until the body was demoted to an inline activation, whose stores then each left the core loop for the helper. The core loop's store, and a leaf's buffered or constructor store, now write a validated slot member (the site's `StoreAttrSlot`, keyed by the class version) directly, as the helper does. - The inline call and return pushed the activation, the parked slot and the result through out-of-line `Vec::push` calls, emptied the locals through a `Drain`, and called two small helpers out of line. They now push with the room check in line and release the locals in place. Measured (retired instructions per call, a local call of each case): `P(i, 2)` for a two-slot class 6,219 -> 4,027 (CPython 1,460); `p = P(i, 2); return p.x + p.y` 7,098 -> 4,891 (CPython 1,762); a call of a one-line non-leaf function from a loop 2,179 -> 2,044, a method call 2,493 -> 2,362; `with cm` 4,426 -> 4,213. Deltablue 33.18M -> 32.62M per unit (32.07M -> 31.50M once warm, 40 -> 120 units).
…riables frameless Two more costs of closures and functions made in a loop: - Every new function copied its code's template of name slots (`__module__`, `__name__`, `__qualname__`): a small dict built, three keys and values cloned, then freed with the function. A function now holds the shared template as a seed and copies it into its own slots only when something reads or writes them; most functions never do. - A function with free variables (`lambda x: x * k`, a nested `def` that reads its maker's locals) was never a leaf, so each call took an inline activation. Leaf plans now read a free variable straight from the function's closure cell (a leaf has no `STORE_DEREF`, so nothing can rebind it during the evaluation; an empty cell declines to the full path, which raises), and such functions are frameless leaves. Measured (retired instructions per call, a local call of each case): a call of an existing closure `g(i)` 2,894 -> 1,466 (CPython 920); `mk(i)(1)` 8,544 -> 6,574 (CPython 2,502); building a nested `def` 4,926 -> 4,152; `lambda` built 3,514 -> 2,806, built and called 4,275 -> 3,490 (CPython 1,549). Importing argparse, typing, dataclasses, email.message, json, unittest, logging and decimal 727M -> 716M. Deltablue is unchanged within noise.
Tier 2 compiles code during start-up and imports only after sixteen times its normal threshold, since that code is mostly run once; the frame compiler had no such budget. Importing several modules that each compile regular expressions made `re._compiler._optimize_charset` hot, and compiling its 559 instructions cost about 100M instructions, a tenth of the whole import. Code that gets hot while start-up or an import runs now compiles only after sixteen times the frame compiler's threshold too. `import typing, json, dataclasses, argparse, logging, asyncio, inspect, textwrap, http.client, unittest` in one process: 978M -> 874M instructions (python3.14: 447M). Single-module imports are unchanged.
`_opcode` (imported by `opcode`, and so by `dis`, `inspect` and `asyncio`) built a lambda per opcode for its popped and pushed counts and a set per flag predicate when it was imported, though most programs never ask for a stack effect. The generator now emits the stack tables behind `_stack_tables()`, built on the first `stack_effect` call, with a plain int for every count that doesn't depend on the oparg, and the predicates test literal frozensets of the opcodes that have each flag. `_weave_flowgraph` reads the tables the same way. Import-only instructions: opcode 27M -> 13M (python3.14 3M), dis 57M -> 43M (20M). The regenerated module answers all 21,210 checks of every predicate and stack effect the same as before.
Classes built by a metaclass (`enum`, `abc`, `typing`) spent most of their time outside the class body's own code: - A `__prepare__` namespace with a Python `__setitem__` (enum's `EnumDict`) took a bound method and a fresh, unhashed key string per `STORE_NAME`. The function is now called with the namespace as its first argument and the code's interned name as the key; `LOAD_NAME` probes such a namespace without building a key object. `__set_name__` hooks that are plain functions are called the same way. - Zero-argument `super().name` handed every name it couldn't resolve to a plain function (dunders such as `super().__setitem__` and `super().__new__`, and methods of built-in bases) to the full path, which builds a `super` proxy instance, introspects the frame and runs the generic attribute lookup on the proxy. When nothing observes the call and the receiver passes the basic `supercheck`, `LOAD_SUPER_ATTR` now walks the MRO from the class and receiver already on the stack, sharing the walk with the proxy's own lookup. - Attribute reads on a class whose metaclass is a Python subclass of `type` with the default `__getattribute__` went to the full path. When the metaclass's MRO has no entry for the name, the class's MRO answers exactly as for a plain metaclass, so the leaf paths take them (uncached by site, as the class's version doesn't follow the metaclass). `cls.__dict__` and class-level dicts and lists (`_member_map_`, `_member_names_`) read in the leaf paths too. Creating a three-member `Enum` costs 1.82M -> 1.63M instructions (python3.14 0.37M), an `IntFlag` 2.56M -> 2.27M (0.54M), and a class whose metaclass `__prepare__` returns a dict subclass 216K -> 151K (83K). Import-only instructions, before -> after: typing 67M -> 63M, json 86M -> 84M, dataclasses 186M -> 182M, argparse 80M -> 77M, logging 281M -> 276M, asyncio 447M -> 432M, inspect 183M -> 176M, http.client 228M -> 218M; the ten modules imported together 680M -> 670M (python3.14 447M).
CI builds with `-D warnings` on the current stable toolchain, which now deprecates the `std::u64` module (used by the vendored lzma-sys) and `AtomicU64::fetch_update`, and adds Clippy lints that flagged code on this branch. The vendored crate uses the associated constant, the type version counter its own compare-exchange loop (`try_update` is newer than the MSRV), and the workspace allows the new `assert_is_empty` lint, which fires on tests throughout. The rest is `cargo fmt` and `cargo clippy --fix`.
Compiling tier-2 code without the mid-end optimizer saved a few percent of compile time but cost numeric loops dearly: the A/B bench gate against main showed `jitloop` 3.7x and `spectral_norm` 1.36x slower. Tier 2 optimizes again by default (`WEAVEPY_JIT_OPT=0` turns it off); leaf plans stay unoptimized unless `WEAVEPY_JIT_OPT=1`. Bench gate vs main (macOS, A/B): geometric mean 0.77x, no fixture regressed; 0.52x CPython.
…iles The frame compiler's admission rule counted every `LOAD_ATTR` as native, but its code serves only split-layout fields and natively served instances' fields. A loop reading a class attribute through an instance (`o.K`) compiled anyway and left for the core loop on every iteration, which cost about 1,000 instructions per read against the interpreter's 550. A site now counts as native only once the core loop has recorded a field there.
`core_return`'s in-line return parked every clean-stack activation through `inline_park_clean`, including one that caught an exception and so had its handler state set. That path skips resetting the frame's handler, saved-exception, and frame-object fields, and its debug assertion failed in every debug build that imported the standard library. Such a frame now goes through `inline_park`, as `inline_deliver` already does.
When the literal prefix is skipped in full, `prefix_len - prefix_skip - 1` is -1. CPython computes it with signed sizes; the Rust port overflowed, which panicked in debug builds (release builds wrapped to the right value).
Natively built exceptions no longer store a separate message, so the report's Rust-side rendering fell back to `args[0]` and printed a failed system call's error as `OSError: 9`. `exception_message` now uses `OSError.__str__`'s errno form, shared with `oserror_str`, which fixes `test_signal`'s wakeup-fd write report.
…tion The quiet loop resumed a handler without the per-instruction prologue's GIL countdown and hot-gate check. Code that only calls and catches, such as recursion retried from its `RecursionError` handler, then never gave up the GIL, so `_thread.interrupt_main()` from another thread never landed and `test_threading` timed out. Every catch site now takes the prologue's checks, as the full step's already did.
The parity report now carries the bench gate's A/B results against `main` and CPython 3.14.8, current import and memory costs, per-operation costs for what still trails, and the validation the branch passed.
owenthcarey
force-pushed
the
perf/outpace-cpython
branch
from
October 3, 2026 09:32
aed1889 to
46a97fc
Compare
owenthcarey
marked this pull request as ready for review
October 3, 2026 09:32
The A/B gate builds the merge-base with the workflow's `-D warnings`, so a lint a newer stable toolchain added fails the reference build before anything is timed: Rust 1.99 deprecates `std::u64`, which `main`'s vendored lzma-sys still uses. The reference binary is only timed, so its build now caps lints at warnings.
`inline_method` widened the split block's length with `uextend.i64`, but `uload32` already yields an `i64`. Release builds skip Cranelift's verifier; x86-64 tolerated the instruction, but AArch64 encoded an invalid one, so compiled plans with in-line method loads died with SIGILL on macOS ARM (`test_jit_plan_caches`, `deltablue`). A new regression test runs the plan-cache tests with both JIT verifiers on.
The Windows gate has measured `call_overhead` at 1.2x the merge-base on every PR run, while the same build, environment, and gate pass at 0.65x on separately dispatched runners. `bench_cold_jit.py` now takes its fixture list as an option, and the failure-only Windows diagnostic adds `call_overhead`, so the next failure records its timings and JIT traces from inside the failing job.
In the last failing run, the gate timed the PR binary's `call_overhead` at 89 ms while the cold-JIT diagnostic, in the same job, timed it at 38 ms (the merge-base and CPython agreed between the two at 60 ms). A failure-only step now times direct runs that each change one thing the two differ in: the environment, the caches, and the binary path.
…e fails Direct runs in the failing job time the PR binary's `call_overhead` at 0.65x the merge-base while the gate times it at 1.2x, whatever the environment, caches, or binary path. The failure-only step now also times it with an absolute fixture path, launched from Python, and through `weavepy-bench run` with each binary alone.
… fails The last failing run pinned the slowdown to the fixture path: the PR binary times `call_overhead` at 75 ms from a relative path and 138 ms from an absolute one, however it's launched, while the merge-base is unaffected. The failure-only step now separates the caches involved and records JIT and burst statistics for both paths.
The slow configuration (an absolute fixture path with warm caches) has the same tier-2 compiles as the fast one, and any change to the caches or instrumentation makes it fast again. The failure-only step now runs that configuration with each tier switched off or reconfigured, twice each, to show which one it depends on.
Every JIT tier switch leaves the absolute-path run slow, and with the JIT off both paths time alike, so the difference is in how compiled code runs rather than what compiles. The failure-only step now records the VM's tier-2 counters (native calls, fallbacks, deopts) for both paths.
The slow and fast configurations run the same tier-2 paths, and every in-process statistic tried so far shifts the layout enough to hide the slowdown. The failure-only step now records sampling profiles of both from outside the process and uploads them with the diagnostic artifact.
The slow run's profile spends a fifth of its samples in an unnamed cluster of `python314.dll` functions, which samply couldn't symbolicate on the runner. The failure-only step now uploads the binaries and their PDBs beside the profiles so they can be symbolicated offline.
Each tier-2 keyword call of a pure leaf that collects `**kwargs` allocated a dictionary and freed it on return. When those blocks were alone on their mimalloc page, the page went back to the arena on every free and was fetched fresh on the next call. A Windows profile of the bench gate's `call_overhead` put a quarter of its samples there, which made it 1.2x the merge-base whenever the heap happened to fall that way. The emptied dictionary of a call that nothing kept now serves the next call; clearing restamps it, so the callee still sees a new dictionary. A regression test checks that returned, stored, and captured dictionaries stay intact and that no keys leak between calls.
The Windows bench gate passes now that keyword leaf calls reuse their `**kwargs` dictionary, so the failure-only step that profiled and timed the slowdown has done its job. The cold-JIT diagnostic keeps `call_overhead` in its fixture list.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Makes WeavePy faster than CPython 3.14 on 21 of the 24 bench-gate fixtures (with
dict_opsat parity), and matches or beatsmainon every fixture. Across the suite, the geometric mean is 0.50× CPython's time and 0.75×main's. The full report, with method, per-fixture tables, imports, memory, and what still trails, is in docs/PERFORMANCE-CPYTHON-PARITY.md.What changed
frame_jit): a baseline Cranelift compiler for loops tier 2 leaves behind (method calls, container operations, generator bodies). It admits a loop only when most of its instructions stay native, now including whether each attribute read has hit a field.dict.items()loops, object compares, and deque lanes. Cranelift's mid-end stays on for tier 2.key=lambda t: t[1]).collections.deque,datetime(packed values with a per-class pool), compiled regular expressions and their matches, andmapandfilter.set(), big-int arithmetic in the core loop,isinstanceagainst tuples, cheaper function creation, and keyword leaf calls that reuse their**kwargsdictionary (which also removes an allocator page churn that made Windowscall_overheadlayout-sensitive).-D warnings; the bench gate builds its merge-base reference with lints capped, sincemainpredates those lints.Bench gate (local run, macOS x86-64, vs
mainand CPython 3.14.8)CI's A/B gates pass on Linux, macOS ARM64, and Windows.
mainStill trailing CPython
deltablue(3.0×),generators(1.1×), anddict_ops(0.92 to 1.06× across runs): class attribute reads through an instance,isinstance, and Python-to-Python calls still leave compiled code for general helpers.Validation
mainfails the same way on the development host (test_capi,test_genexps,test_tempfile,test_utf8_mode, and themultiprocessingspawn and forkserver timeouts).cargo test --workspace --all-targets --all-features,cargo test --workspace --doc, Clippy with-D warnings,cargo fmt --check, the tools' tests, and the artifact check pass.OSErrorprinted asOSError: 9in unraisable reports (test_signal);interrupt_main()never landing during call-and-catch recursion (test_threading);