This repository is Open64 with active work around very-high-level WHIRL, DSL metadata, tensor descriptors, and a future Python frontend ingestion path. Use this file as the compact default context. Pull the larger design documents only when the current task needs detail.
- Preserve existing Open64 and WHIRL compatibility unless a task explicitly requires a versioned IR change.
- Binary WHIRL file backward compatibility is a standing requirement. Do not change WHIRL binary image layout, ELF section contracts, opcode encodings, type-kind encodings, or reader/writer behavior unless the task explicitly calls for a versioned IR change with a documented migration path.
- Binary WHIRL uses Open64's mapped-image and ELF framework. The binary writer finalizes table and tree images into their ELF sections, and the reader maps those images back into the compiler address space. Do not describe or implement this as object-stream encoding/decoding or format conversion.
- Treat
WHIRL.pdfas the normative architectural baseline for WHIRL semantics. Before changing DSL node representation, operators, opcode encoding, kid layout, types, mappings, printing, or mapped-image finalization, review the corresponding definitions and design objectives inWHIRL.pdf. - Design DSL WHIRL as an additive extension of the existing WHIRL model. Preserve established rules such as strict tree structure, direct operand kids, operator/result/descriptor semantic completeness, continuous lowering, mapping-table annotations, and standard binary/ASCII inspection. Any intentional departure must be explicitly justified, versioned when necessary, and covered by compatibility tests.
- Maintain enough architectural detail for the DSL extensions to become an
appendix to
WHIRL.pdf. The appendix should distinguish normative semantic contracts from transitional carriers and historical ABI details, and map each extension to its operator definition, node layout, allowed WHIRL levels, type rules, verifier requirements, lowering path, binary format, andir_b2arepresentation. - Do not copy historical physical sizes or ABI assumptions from
WHIRL.pdfwithout checking the current Open64 implementation. Use the document for semantic intent and verify concrete layouts against the current source, target ABI, binary reader, and binary writer. - Treat Python as a source-language frontend and ingestion layer only. The Open64 middle end must consume a binary WHIRL artifact without depending on the Python interpreter.
- Keep WHIRL table construction, symbol/type creation, TensorDescriptorIR attachment, opcode attributes, contracts, compiler metadata, mapped image finalization, and binary IR compatibility in C++.
- Make ingestion emit first-class DSL/common operators or stable DSL markers, not early intrinsic-call placeholders, runtime calls, or target kernels.
- Run a gatekeeper/verifier before lowering so malformed very-high-level WHIRL does not silently reach downstream optimization passes.
- Do not treat the MPL frontend path as a Very High Level WHIRL ingestion
path. MPL is incomplete reference code for porting another IR roughly
equivalent to Middle Level WHIRL; it has no equivalent feature set for Very
High Level WHIRL.
mpl2whirlremains a valid reference for baseline Open64 symbol/type tables, scope and PU ownership, mapped-image finalization, and binary reader/writer mechanics unless a specific incompatibility is demonstrated and documented. - Follow the coding guidelines below for every touched file.
- When new DSL infrastructure overlaps an established Open64 service, first study and reuse that service's intended extension points, controls, data structures, and tests. Preparation and postprocessing for DSL semantics are acceptable; an independent parallel implementation requires a documented incompatibility and explicit design review.
- DSL expression simplification must use the traditional Open64
wn_simp_code.hrule engine whenever the traditional rule has equivalent semantics. Follow the established physical-WN protocol: prepare a stack-local WN view for the simplifier, discard that temporary input after the call, and retain only simplifier results allocated in the normal WN memory-pool space. Tensor descriptor, effect, ownership, and logical-opcode checks belong in preparation; DSL result symbols, metadata, lineage, and mapped-image relationships belong in postprocessing. - Do not duplicate traditional constant folding, identity, reassociation, cancellation, or factorization rules in a separate DSL rule engine. Add a DSL-only rule only when no traditional rule has equivalent semantics, and document and test that distinction.
- Revisit this continuity principle during design reviews for new DSL services. The concrete reuse mechanism may differ by subsystem, but the review must identify the existing Open64 design being preserved or explain why it cannot be reused.
- When WOPT admits unlowered DSL expressions, decode the physical
OPR_DSLboundary representation into a first-class logical CODEREP identity. CODEREP hashing, equality, printing, effects, and simplification must include the DSL operator, version, canonical attributes, and TensorDescriptorIR identity. Continue using WOPT's existingCODEREPinstantiation ofwn_simp_code.h; do not add a parallel WOPT simplifier. - Assign DSL APIs by semantic applicability, not by the first domain that
requested them. An API belongs in
osprey/common/comwhen its contract is domain-neutral or it is naturally reusable by AI/tensor compiler domains, even when FHE funded or first exercised it. Move an API to an FHE-owned module only when its correctness contract depends on cryptographic schemes, ciphertext state, key requirements, bootstrap or approximation policy, encrypted-runtime ABI, or other irreducibly FHE semantics. Caller location, current usage count, and anFHEimplementation milestone are not by themselves sufficient reasons to classify an API as FHE-specific. - Keep
common/compolicy-free. It may define DSL IR records, stable IDs, constructors/interning, structural verification, generic accessors, and logical printers. Fact harvesting from a PU, candidate discovery, legality analysis, cost/profitability modeling, optimization selection, and IR transformation belong in the phase that owns the compilation scope:be/vhofor WHIRL/VHO work,be/optfor CFG/SSA/CODEREP work,be/lnofor canonical-loop work, and IPA for cross-PU work. Do not place an optimization decision incommon/commerely because its selected result is represented by a common IR record. A domain-neutral, AI-reusable atomic mutation mechanism may remain incommon/comwhen the owning phase supplies the complete reviewed request and retains every legality, profitability, and policy decision. - Place persisted FHE source-image, planning-image, approximation, context
state, materialization, encryption, key, and FHE logical-printing contracts
under
osprey/common/fhe. Keep their mapped-image row and binary compatibility obligations intact, but do not place those irreducibly FHE APIs back incommon/com. Generic AI-reusable DSL transactions remain incommon/comunder rules 6 and 7.
- Preserve the two distinct roles of the WOPT component. PREOPT cleans up and canonicalizes WHIRL before major optimization; full WOPT performs optimization on that canonical representation. PREOPT is also the shared canonicalization service for LNO and IPA where those pipelines invoke it; do not describe or design its canonical form as WOPT-only.
- When introducing a new WHIRL or DSL construct, define its canonical form and identify every downstream consumer. Ensure PREOPT produces or verifies that form before WOPT, LNO, or IPA consumes it in the applicable pipeline.
- Optimization candidate selection should rely on canonical IR whenever possible. Do not make every optimization recognize multiple equivalent tree shapes when PREOPT can normalize them once for all downstream phases.
- Keep canonicalization separate from profitability and transformation. PREOPT normalizes representation and exposes optimization opportunities; WOPT, LNO, and IPA identify legal and profitable candidates and perform transformations within their respective compilation scopes.
- A new canonicalization rule must preserve language, tensor descriptor, effect, alias, source-position, and strict floating-point semantics. It must honor the relevant phase and simplifier controls.
- When WOPT, LNO, or IPA requires a new canonical property, update the PREOPT contract, phase ordering, verifier, diagnostics, and tests together. Tests must prove that equivalent input forms converge to the canonical form and that every claimed downstream consumer receives that form in its actual pipeline and compilation scope.
- If a construct cannot be canonicalized before a downstream optimizer, document the reason and the additional candidate-selection complexity explicitly. Treat this as an exception requiring design review, not the default implementation path.
- The compiler driver establishes compilation scope and phase lifetime. A
service invoked by a phase must operate within that established scope; it
must not enlarge its own scope by traversing
PU_Info, restoring another PU's local symbol table, or switchingCurrent_puon its own. - Normal backend compilation is per-PU. VHO, Preopt/WOPT, LNO, and CG operate on the active PU or on an explicitly selected REGION nested within that PU. REGION is a smaller intraprocedural scope, not an interprocedural scope, and its RID/map/pool lifetime remains driver-owned.
- Global symbol, type, string, TCON, and managed DSL tables provide identity, lookup, and boundary evidence. Their process-wide visibility does not grant a per-PU pass authority to analyze or mutate another PU.
- IPL is a per-PU summary-producing phase. It may prepare facts for IPA, but it does not turn an ordinary backend pass into an interprocedural pass.
- Cross-PU analysis or transformation belongs to IPA and occurs only when the
-ipacompilation path establishes call-graph scope. Such work must use IPA-owned call-graph traversal, summaries, and explicit PU-context services such asIPA_NODE_CONTEXT; it must not be smuggled into VHO, WOPT, LNO, CG, or a common/com utility. - Every new analysis or optimization must state its scope: expression, basic block, REGION, PU, file summary, or IPA call graph. Its ownership, invalidation, rollback, diagnostics, and tests must use that same scope.
- A per-PU pass may validate call/formal/result contracts visible at its
boundary, but it must not infer that it may rewrite the opposite side of a
call edge. Coordinated caller/callee refinement requires an explicitly
designed IPA pass enabled by
-ipa. - Test scope must match implementation scope. A PU-local change requires
intraprocedural positive, negative, rollback, and boundary-validation tests;
it does not require a cross-PU transformation test. Multiple-PU fixtures may
still prove independent driver coverage and absence of cross-PU mutation.
Require call-graph propagation or coordinated caller/callee transformation
tests only for code that executes in IPA scope under
-ipa.
- Treat optimization levels as analysis and transformation scope contracts. DSL optimization must follow the same contracts as traditional Open64; a DSL operator, tensor type, REGION, or domain does not grant permission to optimize at a broader scope.
-O0performs no optimization. It may verify legality and perform only the straightforward semantic lowering required to make the program executable. It must not perform fusion, algebraic improvement, layout optimization, profitability-driven rewriting, parallelization, or other performance-oriented transformations.-O1permits basic-block-local optimization only. Its analysis and transformations must not depend on control-flow facts outside the active basic block, REGION-local equivalent, or other explicitly local unit.-O2permits PU-level optimization over the active procedure's control-flow graph. It may use intraprocedural data-flow, alias, SSA, PRE, and related analyses, but it must not infer or transform across a PU boundary.-O3adds optimization around canonical loops, with emphasis on memory behavior and parallelization. Loop transformation, tiling, locality, memory-hierarchy use, vectorization, and parallel execution must consume the canonical loop and dependence contracts established by the earlier phases.-ipaexplicitly expands analysis beyond one PU through the IPA-owned call graph and summaries. Cross-PU transformation remains limited, reviewed, and controlled; enabling-ipadoes not authorize arbitrary whole-program mutation by PU-local phases.- Every new DSL analysis or transformation must declare its minimum optimization level, maximum compilation scope, required canonical form, invalidation behavior, and controlling option. The driver and phase must leave it disabled below that level and must not silently enlarge its scope.
- Validation must exercise the level boundary: prove
-O0preserves the unoptimized semantic form through straightforward lowering, prove the pass runs at its declared level, and prove it neither runs nor consumes out-of-scope facts at lower levels. Cross-PU tests are required only for explicitly enabled-ipawork.
- Do not add a new library dependency to
be.sowithout explicit design and build approval. This includes direct linkage and source-level references that create new defined or undefined symbols in the shared library. - Before approving a dependency, identify every consumer of
be.so, includinglw_inline, backend plugins, and standalone tools, and prove that each consumer remains link-closed on every supported build configuration. - Prefer an existing Open64 service or an already compatible header-only facility when structured parsing or another utility is needed in a backend pass. Do not introduce ad hoc parsing to avoid the dependency review.
- Validation for an approved dependency change must rebuild
be.soand its shared consumers, inspect their defined and undefined symbols, and document platform, static/shared, licensing, version, and binary-distribution impact. - Frontend-only libraries and interfaces, including
DSL_Builder_*, must not leak intobe.so. For example, an FHE backend may parse an authenticated JSON manifest with the repository's compatible header-only RapidJSON facility, but must not add JsonCpp linkage tobe.sowithout approval.
- Do not introduce tab characters in any file touched in this Open64 project.
- Use spaces for indentation and alignment.
- Makefiles may use literal tabs only where required by standard make recipe conventions.
- When editing existing Open64 C/C++ code, match the surrounding indentation style with spaces rather than reformatting unrelated code.
- Builder interfaces that create a symbol must make every reasonable attempt
to set its declaration source position with
Set_ST_Srcpos(). When a builder creates both a defining WN and a result ST, propagate the same complete source position, including file, line, and column, to both objects. - Every newly created
.hand.cxxfile must begin with a concise architectural comment after the copyright notice. State the file's purpose, owning component or compilation scope, important behavior or compatibility boundary, and the repository-relative path of the controlling design or plan document. A public header should identify the contract it exposes; its implementation file should identify what it deliberately does and does not change. - Add short orienting comments before non-obvious algorithm boundaries such as preparation, shared-engine reuse, legality classification, transactional mutation, verification, and postprocessing. Explain the invariant or design reason rather than restating the code. During final review, verify these comments exist in every newly added C/C++ source file and still agree with the implementation and cited design document.
- Every function in a DSL-related
.hor.cxxfile, including public APIs, private/static helpers, callbacks, and commit-only services, must have a function-level contract comment. State the function's purpose, its role in the surrounding caller/callee or phase relationship, important ownership and mutation boundaries, required invariants, and failure or rollback behavior where applicable. Public declarations must describe the caller contract; definitions must describe implementation responsibilities that are not evident from the declaration. Transactional functions must identify preflight, commit, verification, and rollback roles. A brief comment is sufficient for a trivial accessor or predicate, but merely restating the function name, parameters, or return type is not. New DSL files must satisfy this rule for every function before review. When materially modifying an existing DSL function, add or correct its contract comment in the same change; do not require unrelated whole-file comment churn for a narrow fix. - All new DSL-related C/C++ files and APIs, including
dsl_*,opt_dsl_*, and names prefixed withDSL_orWOPT_DSL_, must follow the established Open64be/optnaming style. UseALL_CAPS_WITH_UNDERSCORESfor classes, structs, typedefs, enum types, and public symbolic types;_lowercase_snake_casefor ordinary private data members; andlowercase_snake_casefor parameters and local variables. - Name member functions with an initial capital and underscore-separated
words, while preserving established compiler acronyms in uppercase, for
example
Build_candidates(),Verify_IR(), andCompute_PRE_saves(). Do not introduce CamelCase forms such asBuildCandidates()or mixed-case acronym forms such asDslin new DSL code. - Name trivial field accessors after the property, such as
Cfg()orDescriptor(). UseSet_,Reset_,Is_,Has_, andCan_for mutation and predicate APIs. ReserveGet_for operations that perform a lookup, computation, copy, or output assignment rather than a direct field read. - Give public free functions, globals, enum values, and macros a stable
subsystem prefix such as
DSL_orWOPT_DSL_. Enum values and macros use uppercase underscore-separated names. Extern free functions follow the WOPT API convention: preserve the uppercase subsystem acronym, then use lowercase underscore-separated component and operation words, for exampleDSL_tensor_evolution_create()andWOPT_DSL_populate_tensor_control_snapshot(). Preserve an additional uppercase word only when it is itself an established compiler acronym, such asWNorDIVREM. Avoid new unprefixed global names and prefer inline functions over macros unless an existing Open64 protocol requires a macro. - Use established optimizer abbreviations such as
cr,stmt,bb,cfg,wn,phi,aux_id,kid0, andkid1where their meaning is local and unambiguous. Use descriptive lowercase underscore-separated names for new semantic concepts. Do not use pointer-name prefixes or other Hungarian notation. - Treat compact historical structures, generated interfaces, required Makefile syntax, and tightly scoped template utilities as exceptions, not precedents for new DSL naming. When an owning directory has a stricter established convention, preserve that convention and document any intentional departure during review.
- Preserve Open64's command-line option propagation model. The driver passes the user's complete option set through the compiler pipeline.
- Each compiler phase selects and processes only the options applicable to that phase and silently ignores options owned by other phases.
- A phase must not reject a command merely because unrelated Open64 options
are present. New phases, including
torch2whirl, must follow this convention instead of requiring the driver to remove every unrelated option. - For Python frontend option processing, use
clang2whirlor the existing C frontend as the reference implementation pattern for receiving the complete option set, selecting frontend-owned options, and silently ignoring options owned by other phases.
Python model / DSL
-> torch.export / FX / DSC capture graph
-> open64_dsc.WhirlExportInterpreter
-> open64_dsc._whirl native extension
-> libopen64_whirl_builder
-> binary very-high-level WHIRL
-> opencc -x whirl
The Python package should coordinate capture and graph traversal. The native builder should own all WHIRL-specific construction.
- DSL frontends may attach structured metadata with
OPR_COMMENTnodes using the reserved prefix__WHIRL_DSL__:<domain>:<feature>:<payload>. - Older tools should still see legal WHIRL. DSL-aware passes must validate and consume required markers before WOPT, LNO, CG, and other canonical-WHIRL phases.
- Tensor identity currently stages through an existing
TY_IDXcarrier plus a tensor extension side table. Do not force a new WHIRL operator enum, opcode enum, or binary type-kind change unless the stage explicitly calls for it. - TensorDescriptorIR carries semantic tensor value state: element type, rank, shape, layout, strides, traits, lineage, placement, and related facts.
- Compiler metadata carries source context, diagnostics, pass ownership, lowering hints, and profiling data. Do not mix compiler metadata into tensor type equivalence unless a later schema explicitly promotes it.
- Avoid nested STL containers inside WHIRL-style table records that may later need to be dumped, reloaded, compared, or persisted in a binary IR image.
- Domain operators must remain domain-visible until domain gatekeeper checks complete. Do not hide CNN/Transformer ABI, layout, padding, stride, mask, head-layout, or similar legality checks through premature promotion or lowering.
- Treat physical
OPR_DSLas a private WN escape tag. DSL-aware compiler code, diagnostics, ASCII dumps, and traces must use logical operator APIs and names such asOPR_DSLADD; they must not expose or manually decode the physical escape tag or its internal record index.
- Keep
ir_a2bandir_b2aas compatibility gates. Any staged IR change must preserve behavior for existing WHIRL, or update those tools and matching tests in the same staged change. - The binary WHIRL artifact is the frontend boundary. Avoid special in-memory
bypasses from Python into the compiler pipeline; artifacts should be
inspectable with
ir_b2a -st -srcbefore combined driver integration. - Run gatekeeper verification before canonical lowering. Missing tensor attributes, invalid operand compatibility, unknown contracts, unsupported opcodes, and domain legality failures should be diagnosed before WOPT, LNO, or CG.
- Public DSL opcode enum values, string names, categories, levels, traits, shape rules, effect models, and promotion states are compatibility contracts once they appear in IR files, tests, dumps, or registries.
- TensorDescriptorIR carries semantic tensor value state. Compiler metadata carries source context, diagnostics, pass ownership, lowering hints, profiling, and source names. Compiler metadata must not affect tensor type equivalence.
- Do not add
KIND_TENSORor new binary tensor descriptor sections casually. First-class tensor type or tensor descriptor binary-section changes must land with printer,ir_b2a -st -src, binary reader, binary writer, ASCII reader/printer plan, verifier, and fallback behavior. torch2whirl, Python bindings, and the native frontend bridge must not depend on backend code generation. Backend, runtime, and kernel lowering happen after gatekeeper verification.
Use this workflow when adding a model family or a new domain to openpy:
- Add a small, deterministic, dependency-light source model fixture.
- Capture the real frontend graph and produce a reviewable operator census.
- Classify every captured operation as a common substrate operation, domain expression, region contract, state/effect, compiler metadata, constant, or explicit unsupported case.
- Review and publish stable native names, versions, operands, attributes, descriptor rules, effects, gatekeeper checks, and lowering ownership in the WHIRL infrastructure plan.
- Implement the native builder, mapped-image, reader/writer, logical printer, gatekeeper, and lowering support required by the published contract.
- Migrate the frontend from census/mock handling to opaque native APIs. The frontend must not learn WN layout or private physical DSL encodings.
- Certify the binary artifact, external data,
ir_b2a -st -srctrace, andopenpy -O0path across a process boundary.
Do not allocate DSL opcodes or broaden native WHIRL contracts solely from an expected model architecture. First capture a representative source model and review its semantic operator census. Conversely, do not force a captured model into existing operators when doing so would erase domain semantics. Frontend capture discovers the requirements; common/com review remains authoritative for native contracts and binary WHIRL behavior.
Maintain a separate frontend ingestion plan for each substantial domain or
model family and link it to doc/WHIRL-DSL-INFRASTRUCTURE.md. The frontend plan
owns capture, mapping, artifacts, and diagnostics; the infrastructure plan owns
native representation, compatibility, verification, inspection, and lowering.
- Follow
doc/IMPACTFUL-TEST-STRATEGY.mdfor FHE and other expensive DSL validation. Select tests from the changed contract and its reverse dependencies before launching a broad suite. - Do not use the largest available model as the default edit-loop test. Run focused semantic and transaction regressions first; reserve full-model execution for an explicitly invalidated whole-program boundary, final PR head, milestone closure, or scheduled certification.
- Before starting a test expected to exceed ten minutes, state the affected contract, why smaller evidence is insufficient, and the expected duration.
- Reuse expensive artifacts only when content-addressed producer, trace, and audit receipts prove all relevant inputs unchanged. Documentation, printer, and auditor changes invalidate only their respective downstream evidence.
- Unknown source impact broadens the selected tests. It never justifies skipping validation. A full-model pass supplements rather than replaces focused rollback, malformed-input, and semantic-oracle tests.
- Human review of compiler artifacts is part of the development and
validation process. Tests that produce meaningful WHIRL evidence should
retain the binary
.Bfile,ir_b2a -st -srcoutput, requested phase.ttraces, side payloads, and relevant diagnostics after the test exits. - Clean the designated artifact directory at the start of the next run, not at the end of the current run. A completed run must leave its evidence available until it is replaced by a later run.
- Docker tests must write reviewable artifacts through an explicit host bind mount. Do not leave the only copy in a container filesystem or a container temporary directory that disappears when Docker exits.
- Keep artifact families in clearly named per-test or per-stage directories so one smoke test does not erase another test's evidence.
- Failed runs must not leave partial output with the name of a valid
.Bartifact. Preserve failure logs and diagnostics when useful, while keeping artifact publication atomic. - At the end of validation, report the absolute host paths of retained artifacts so reviewers can inspect them directly.
- Generated review artifacts are normally local build evidence. Do not add them to Git unless the task explicitly requests checked-in golden files or review fixtures.
- When completing a task that produces or validates a compiler trace, include a directly viewable link to the retained trace in the final report. Show a short representative excerpt or summarize the concrete evidence it contains; do not report trace-based validation as complete using only a pass/fail statement.
- Use both
-stand-srcwhenever runningir_b2afor validation or human review. Preserve or mount the original source file at the pathname recorded in the binary WHIRL DST so-srccan interleave source statements with the IR. If the activeir_b2abuild does not yet support-src, treat that as a tooling gap to fix; do not silently omit source cross-reference evidence. - The
ir_b2aoutput must use the input.Bfile's stem, for exampleir_b2a -st -src resnet.B resnet.T. On case-insensitive filesystems whereresnet.Tcollides with a driver-producedresnet.t, preserve the phase trace under a descriptive non-colliding name such asresnet.vho.tbefore producingresnet.T. - Every coding change that adds or modifies an IR transformation must retain
reviewable before-and-after WHIRL evidence. Produce
<case>.before.Band<case>.after.B, reopen each independently withir_b2a -st -srcas<case>.before.Tand<case>.after.T, and retain a unified diff such as<case>.before-after.diff. An in-memory phase trace alone is not a substitute for mapped binary WHIRL evidence. - Capture the before image immediately before the transformation and the after image immediately after it, using the same input, options, target, compilation scope, and source mapping. If the normal driver cannot publish both boundaries, add a focused producer or reviewed checkpoint using the existing WHIRL writer rather than fabricating textual IR.
- The transformation diff must make intentional WN, ST, TY, TensorDescriptorIR,
REGION, and managed-table changes visible while also demonstrating relevant
invariants. Preserve the raw
ir_b2adiff; a focused or normalized excerpt may supplement it but must not replace it. - Treat unexplained diff churn as a review blocker. The validation report and pull-request summary must describe the expected semantic changes, identify important facts that remain unchanged, and link the retained before trace, after trace, and full diff. If a transformation is expected to be a no-op for a fixture, retain and report the empty diff as evidence.
- Stabilize native DSL/common infrastructure before adding Python package code.
- Add or maintain native tests under
osprey/common/com/testsfor tensor types, tensor constants,common.add,common.matmul, andir_b2a -st -srcvisibility. - Keep the first builder API narrow. Python-facing bindings should pass opaque handles and values; C++ should create real WHIRL objects.
- Use existing mapped-image / ELF WHIRL mechanisms for binary artifacts before inventing any new file format.
- Defer combined
opencc -frontend=torch2whirl ...driver integration until the binary WHIRL artifact boundary is stable and inspectable.
When investigating runtime or optimization failures, establish a baseline and add one condition at a time:
- First make
-O0work, then compare-O1,-O2,-O3, and finally-ipa. - If
-O0works but-O1fails, suspect CG. - If
-O1works but-O2fails, suspect WOPT. - If
-O2works but-O3fails, suspect LNO. - If
-O*works but-O* -ipafails, suspect IPA. - After identifying a component, reduce by file, then procedure, then specific optimization. Use binary search where possible.
- Treat
-OPT:wn_simplifyand its abbreviation-OPT:wn_simpas the master WHIRL construction-time simplifier control. When triaging a simplifier regression, preserve separate.Bfiles with simplification enabled and disabled, produce matchingir_b2a -st -srctraces, and compare the resulting WHIRL. New DSL construction-time simplification must honorEnable_WN_Simp; a DSL-specific control may further restrict a stage but must not override the master switch.
Load these only when the task needs the detail:
Open64_Python_FE_Plan.md- compact roadmap for Python ingestion and native WHIRL builder work.doc/WHIRL-DSL-INFRASTRUCTURE.md- DSL marker, tensor extension, and lowering infrastructure details.doc/VHO-DSL-OPTIMIZATION-PLAN.md- optional architecture-independent VHO optimization, parallelization, Preopt/IPA, and compiler-library codesign.doc/HOW-TO-TRIAGE-RUNTIME-FAILURE-OPEN64.md- runtime failure and optimization triage method.doc/TORCH2WHIRL-WHIRL-DSL-API-CONTRACT.md- torch2whirl to common/com DSL builder API inventory and subagent/main-agent API creation protocol.doc/Open64_Domain_Specific_Compiler_IR_Design.md- large staged design; search within it instead of loading it whole.imported_docs/DSC_Master_Design_Doc_v0.9_chapter_7.md- Python DSL ingestion architecture background.osprey/clang2whirl/README.mdandosprey/clang2whirl/NAMING_CONVENTION.md- use only for clang2whirl work.WHIRL.pdf- normative WHIRL architecture and semantic baseline. Consult the relevant operator, node-layout, level, type, mapping, and ASCII-format sections before designing DSL extensions. The planned DSL appendix must remain consistent with this document or explicitly document extensions.
Treat MPL-related sources as reference material only. Do not use them as evidence for the Very High Level WHIRL DSL frontend or Python ingestion design.
Most *.md files under GCC config/ directories are machine-description
inputs, not general project guidance.