Skip to content

07. Environment Variables Reference

github-actions[bot] edited this page Sep 4, 2026 · 7 revisions

Environment Variables Reference

This page provides a complete reference for all environment variables supported by TritonParse.

πŸ’‘ Looking for Python API? See Python API Reference for complete documentation of init(), unified_parse(), TritonParseManager, and other Python interfaces.

πŸ“ Scope: All environment variables documented here affect the tritonparse.structured_logging module, which handles trace generation during Triton kernel compilation and execution.

⚠️ Important: Activation Required in OSS

In OSS (open-source) environments, TritonParse is not automatically enabled. Even when using environment variables, you must call init() in your Python code to activate tracing:

import tritonparse.structured_logging

# Just call init() - environment variables will be automatically applied
tritonparse.structured_logging.init()

How it works:

  • Environment variables are read when the module is imported
  • Calling init() without arguments activates tracing using the environment variable values
  • If you pass arguments to init(), they take precedence over environment variables

Complete example with environment variables:

# Set environment variables
export TRITON_TRACE="./logs/"
export TRITON_TRACE_LAUNCH="1"
export TRITONPARSE_MORE_TENSOR_INFORMATION="1"
import tritonparse.structured_logging

# Activate tracing - environment variables are automatically used
tritonparse.structured_logging.init()

# Your kernel code here...

πŸ“‹ Quick Reference Table

Variable Purpose Default Category
TRITON_TRACE Trace output directory /logs/ Trace Generation
TRITON_TRACE_LAUNCH Enable launch tracing Off Trace Generation
TORCHINDUCTOR_RUN_JIT_POST_COMPILE_HOOK Enable inductor hook Off Trace Generation
TRITONPARSE_MORE_TENSOR_INFORMATION Collect tensor stats Off Tensor Info
TRITONPARSE_SAVE_TENSOR_BLOBS Save tensor blobs Off Tensor Info
TRITONPARSE_TENSOR_SIZE_LIMIT Max tensor blob size 10GB Tensor Info
TRITONPARSE_TENSOR_STORAGE_QUOTA Total storage quota 100GB Tensor Info
TRITONPARSE_DEBUG Debug logging Off Debug
TRITONPARSE_KERNEL_ALLOWLIST Kernel filter patterns All Debug
TRITON_TRACE_LAUNCH_WITHIN_PROFILING Launch tracing during profiler RECORD phase Off Trace Generation
TRITON_TRACE_COMPRESSION Trace file compression format "none" Performance
TRITONPARSE_DERIVED_ARTIFACTS Select derived artifacts to generate "none" Performance
TRITONPARSE_DUMP_SASS SASS dump Off Performance
TRITONPARSE_TENSOR_SAVE_SKIP_RUNS Skip blob saving for first N runs 0 Tensor Info
TRITONPARSE_TENSOR_SAVE_MAX_RUNS Save blobs for at most N runs 0 (unlimited) Tensor Info
TRITON_FULL_PYTHON_SOURCE Full source extraction Off Source
TRITON_MAX_SOURCE_SIZE Max source file size 10MB Source
TRITONPARSE_ANALYSIS Select which analyses to run "all" Analysis
TEST_KEEP_OUTPUT Keep test outputs Off Testing
TORCHINDUCTOR_FX_GRAPH_CACHE PyTorch FX cache On Testing

πŸ“Š Trace Generation Variables

These variables control how TritonParse captures compilation and launch events.

TRITON_TRACE

Description: Directory to store raw trace files.

Property Value
Values Any valid directory path
Default /logs/ (if exists and writable), otherwise disabled
Related Also set via tritonparse.structured_logging.init(trace_folder=...)

Example:

export TRITON_TRACE="./my_logs/"
python your_script.py

TRITON_TRACE_LAUNCH

Description: Enable kernel launch event tracing. When enabled, captures runtime launch parameters including grid dimensions, tensor arguments, and execution metadata.

Property Value
Values "1", "true", "True" to enable
Default Disabled
Related Also set via tritonparse.structured_logging.init(enable_trace_launch=True)

Example:

export TRITON_TRACE_LAUNCH="1"

πŸ’‘ Tip: Launch tracing is essential for reproducer generation and launch diff analysis.


TORCHINDUCTOR_RUN_JIT_POST_COMPILE_HOOK ⭐

Description: Required for tracing kernel launches from TorchInductor (torch.compile). This hook allows TritonParse to intercept and log launch metadata for Inductor-compiled kernels.

Property Value
Values "1", "true", "True" to enable
Default Disabled (auto-enabled if TRITON_TRACE_LAUNCH=1)
When to use When using torch.compile() and need launch tracing

Example:

# Required for torch.compile kernels
export TORCHINDUCTOR_RUN_JIT_POST_COMPILE_HOOK="1"
export TRITON_TRACE_LAUNCH="1"
# Or via Python
import os
os.environ["TORCHINDUCTOR_RUN_JIT_POST_COMPILE_HOOK"] = "1"

import tritonparse.structured_logging
tritonparse.structured_logging.init("./logs/", enable_trace_launch=True)

⚠️ Important: For native Triton kernels (@triton.jit), this variable is not needed. It's only required for torch.compile / TorchInductor generated kernels.


TRITON_TRACE_LAUNCH_WITHIN_PROFILING

Description: Enable launch tracing only during torch.profiler's RECORD phase. When set, TritonParse patches torch.profiler.schedule to toggle launch tracing automatically.

Property Value
Values "1", "true", "True" to enable
Default Disabled
Related Also set via tritonparse.structured_logging.init(enable_trace_launch_within_profiling=True)

Example:

export TRITON_TRACE_LAUNCH_WITHIN_PROFILING="1"

⚠️ Important: Mutually exclusive with TRITON_TRACE_LAUNCH. If both are set, TRITON_TRACE_LAUNCH takes priority and ALL launches will be traced.


πŸ“ˆ Tensor Information Variables

These variables control how TritonParse collects and stores tensor data for debugging and reproducer generation.

TRITONPARSE_MORE_TENSOR_INFORMATION ⭐

Description: Collect detailed tensor statistics including min, max, mean, and standard deviation values. This information is useful for reproducer generation and debugging.

Property Value
Values "1", "true", "True" to enable
Default Disabled
Related Also set via tritonparse.structured_logging.init(enable_more_tensor_information=True)

Collected statistics:

  • min: Minimum value in tensor
  • max: Maximum value in tensor
  • mean: Mean value of tensor elements
  • std: Standard deviation of tensor elements

Example:

export TRITONPARSE_MORE_TENSOR_INFORMATION="1"
# Or via Python
tritonparse.structured_logging.init(
    "./logs/",
    enable_trace_launch=True,
    enable_more_tensor_information=True,
)

πŸ’‘ Use case: When generating reproducers, these statistics allow for better tensor reconstruction that approximates the original data distribution.


TRITONPARSE_SAVE_TENSOR_BLOBS

Description: Save actual tensor data as blob files. Enables highest-fidelity reproducer generation by preserving the exact tensor values.

Property Value
Values "1", "true", "True" to enable
Default Disabled
Storage location <trace_folder>/saved_tensors/
Related Also set via tritonparse.structured_logging.init(enable_tensor_blob_storage=True)

Example:

export TRITONPARSE_SAVE_TENSOR_BLOBS="1"

⚠️ Warning: Can consume significant disk space. Use TRITONPARSE_TENSOR_STORAGE_QUOTA to limit storage.


TRITONPARSE_TENSOR_SIZE_LIMIT

Description: Maximum size for individual tensor blobs. Tensors larger than this limit will be skipped during blob storage.

Property Value
Values Integer (bytes)
Default 10GB (10 * 1024 * 1024 * 1024)

Example:

# Limit to 1GB per tensor
export TRITONPARSE_TENSOR_SIZE_LIMIT=$((1 * 1024 * 1024 * 1024))

TRITONPARSE_TENSOR_STORAGE_QUOTA

Description: Total storage quota for tensor blobs in a single run. Once exceeded, blob storage is disabled for the remainder of the run.

Property Value
Values Integer (bytes)
Default 100GB (100 * 1024 * 1024 * 1024)
Related Also set via tritonparse.structured_logging.init(tensor_storage_quota=...)

Example:

# Limit total storage to 50GB
export TRITONPARSE_TENSOR_STORAGE_QUOTA=$((50 * 1024 * 1024 * 1024))

TRITONPARSE_TENSOR_SAVE_SKIP_RUNS

Description: Skip tensor blob saving for the first N kernel runs (per kernel hash). Useful to skip warm-up runs.

Property Value
Values Integer (number of runs to skip)
Default 0 (no skipping)
Related Also set via tritonparse.structured_logging.init(tensor_save_skip_runs=...)

Example:

# Skip the first 5 runs of each kernel
export TRITONPARSE_TENSOR_SAVE_SKIP_RUNS="5"

TRITONPARSE_TENSOR_SAVE_MAX_RUNS

Description: Save tensor blobs for at most N kernel runs after the skip period. Useful to limit storage usage.

Property Value
Values Integer (0 = unlimited)
Default 0 (unlimited)
Related Also set via tritonparse.structured_logging.init(tensor_save_max_runs=...)

Example:

# Save blobs for at most 10 runs (after skipping)
export TRITONPARSE_TENSOR_SAVE_MAX_RUNS="10"

πŸ”§ Debug Variables

These variables help with debugging TritonParse itself and filtering trace output.

TRITONPARSE_DEBUG

Description: Enable debug logging for TritonParse internal operations.

Property Value
Values "1", "true", "True" to enable
Default Disabled

Example:

export TRITONPARSE_DEBUG="1"
python your_script.py

πŸ’‘ Use case: Helpful when troubleshooting trace generation issues or understanding TritonParse behavior.


TRITONPARSE_KERNEL_ALLOWLIST

Description: Filter which kernels to trace using fnmatch patterns. Only kernels matching at least one pattern will be traced.

Property Value
Values Comma-separated fnmatch patterns
Default All kernels traced

Pattern syntax:

  • * matches any characters
  • ? matches a single character
  • [seq] matches any character in seq

Example:

# Only trace matmul and attention kernels
export TRITONPARSE_KERNEL_ALLOWLIST="matmul*,*attention*,flash_*"

πŸ’‘ Use case: Reduces trace file size and processing time when debugging specific kernels.

πŸ“ Scope: This variable filters events while tracing and is read when tritonparse.structured_logging is imported. It is not implicitly reused by the parser. To filter an existing trace, pass --kernel-allowlist to tritonparse parse or kernel_allowlist= to unified_parse().


πŸ“ Source Extraction Variables

These variables control how Python source code is extracted and stored in traces.

TRITON_FULL_PYTHON_SOURCE

Description: Extract the entire Python source file instead of just the kernel function definition.

Property Value
Values "1", "true", "True" to enable
Default Function-only extraction
Related Also set via tritonparse.structured_logging.init(enable_full_python_source=True), which takes precedence over this variable

Example:

export TRITON_FULL_PYTHON_SOURCE="1"

πŸ’‘ Use case: Useful when you need full context including imports and helper functions for debugging. Required for nested kernels: when the traced @triton.jit entry point is a thin wrapper whose body just calls the real kernel, function-only extraction captures the wrapper and none of the real kernel body.


TRITON_MAX_SOURCE_SIZE

Description: Maximum file size for full Python source extraction. Files larger than this limit will fall back to function-only extraction.

Property Value
Values Integer (bytes)
Default 10MB (10 * 1024 * 1024)

Example:

# Allow up to 20MB source files
export TRITON_MAX_SOURCE_SIZE=$((20 * 1024 * 1024))

⚑ Performance Variables

These variables affect trace file size and compilation performance.

TRITON_TRACE_COMPRESSION

Description: Set the compression format for raw trace files. Each JSON record is compressed individually, allowing incremental writing.

Property Value
Values "none", "gzip", "clp"
Default "none"
File extensions "none" β†’ .ndjson, "gzip" β†’ .bin.ndjson, "clp" β†’ .clp
Related Also set via tritonparse.structured_logging.init(compression=...)

Example:

# Enable gzip compression
export TRITON_TRACE_COMPRESSION="gzip"

# Enable CLP (Compressed Log Processor) format
export TRITON_TRACE_COMPRESSION="clp"

πŸ’‘ Note: Gzip output uses .bin.ndjson extension. Each record is a separate gzip member, so standard gzip tools can decompress the file.

⚠️ Deprecated: The old TRITON_TRACE_GZIP variable is no longer supported. Use TRITON_TRACE_COMPRESSION="gzip" instead.


TRITONPARSE_DERIVED_ARTIFACTS

Description: Select which derived artifacts TritonParse should generate from existing compilation artifacts by invoking external tools. This is the generalized interface for derived capabilities added by the backend adapter layer.

Property Value
Values "none", "all", or comma-separated target stage names such as "sass"
Default "none"
Case sensitivity Case-insensitive
Current example target sass on NVIDIA backends
Requires Any external tool required by the selected derived artifact, for example nvdisasm for sass

How it works:

  • none: disable all derived artifacts
  • all: enable all derived artifacts registered by the active backend
  • Comma-separated names: enable only the requested target stages
  • Empty or whitespace-only input: treated the same as none
  • Unknown target stage names: logged as a warning and ignored

Examples:

# Disable all derived artifacts (default)
export TRITONPARSE_DERIVED_ARTIFACTS="none"

# Enable every derived artifact supported by the active backend
export TRITONPARSE_DERIVED_ARTIFACTS="all"

# Enable only SASS derivation on NVIDIA backends
export TRITONPARSE_DERIVED_ARTIFACTS="sass"

⚠️ Important: all and none should be used alone. If all is mixed with any other names, TritonParse logs a warning and treats the configuration as all. If none is mixed with other names and all is not present, TritonParse logs a warning and treats the configuration as none. If both all and none are present, all takes precedence. ⚠️ Performance note: Derived artifacts run extra external tools during trace extraction. Leave this disabled unless you need the additional artifact content.


TRITONPARSE_DUMP_SASS

Description: Dump NVIDIA SASS (Shader Assembly) from CUBIN files using nvdisasm.

Property Value
Values "1", "true", "True" to enable
Default Disabled
Requires NVIDIA CUDA toolkit with nvdisasm
Related Also set via tritonparse.structured_logging.init(enable_sass_dump=True)
Compatibility behavior Legacy compatibility switch that ensures sass is enabled alongside any configured derived artifacts

Example:

export TRITONPARSE_DUMP_SASS="1"

Recommended modern form:

export TRITONPARSE_DERIVED_ARTIFACTS="sass"

If TRITONPARSE_DERIVED_ARTIFACTS already enables other derived artifacts, setting TRITONPARSE_DUMP_SASS="1" will add the sass artifact.

⚠️ Warning: Significantly slows down compilation. Only enable when needed for low-level analysis.


πŸ”¬ Analysis Variables

These variables control which IR analyses are performed during trace parsing.

TRITONPARSE_ANALYSIS

Description: Select which IR analyses to run during trace parsing. Allows fine-grained control over analysis passes such as loop schedule extraction and procedure checks.

Property Value
Values "all", "none", or comma-separated analysis names (e.g. "loop_schedules,procedure_checks")
Default "all" (all analyses enabled)
Available analyses loop_schedules, procedure_checks, roofline, amd_buffer_ops (AMD only)

Example:

# Run all analyses (default)
export TRITONPARSE_ANALYSIS="all"

# Disable all analyses
export TRITONPARSE_ANALYSIS="none"

# Run only specific analyses
export TRITONPARSE_ANALYSIS="loop_schedules"
export TRITONPARSE_ANALYSIS="loop_schedules,procedure_checks"
export TRITONPARSE_ANALYSIS="roofline"

πŸ’‘ Use case: Use "none" to skip analysis overhead when only raw trace data is needed. Use specific names to run only the analyses relevant to your debugging scenario.


πŸ§ͺ Testing Variables

These variables are primarily used during development and testing.

TEST_KEEP_OUTPUT

Description: Preserve temporary output directories from tests instead of cleaning them up.

Property Value
Values "1", "true", "True" to enable
Default Disabled (cleanup on exit)

Example:

export TEST_KEEP_OUTPUT="1"
make test

TORCHINDUCTOR_FX_GRAPH_CACHE

Description: (PyTorch variable) Control FX graph caching. Disable to ensure fresh compilation for each run.

Property Value
Values "0" to disable
Default Enabled (caching on)

Example:

# Disable cache to ensure fresh compilation
export TORCHINDUCTOR_FX_GRAPH_CACHE="0"
python your_test.py

πŸ’‘ Use case: Essential during testing to ensure kernels are recompiled and traces are generated.


🎯 Common Configurations

Basic Tracing (Native Triton Kernels)

export TRITON_TRACE="./logs/"
python your_triton_script.py

Full Tracing (torch.compile Kernels)

export TRITON_TRACE="./logs/"
export TRITON_TRACE_LAUNCH="1"
export TORCHINDUCTOR_RUN_JIT_POST_COMPILE_HOOK="1"
export TRITONPARSE_MORE_TENSOR_INFORMATION="1"
python your_torch_compile_script.py

Reproducer-Ready Tracing

export TRITON_TRACE="./logs/"
export TRITON_TRACE_LAUNCH="1"
export TORCHINDUCTOR_RUN_JIT_POST_COMPILE_HOOK="1"
export TRITONPARSE_MORE_TENSOR_INFORMATION="1"
export TRITONPARSE_SAVE_TENSOR_BLOBS="1"
export TRITONPARSE_TENSOR_STORAGE_QUOTA=$((10 * 1024 * 1024 * 1024))  # 10GB limit
python your_script.py

Debug Configuration

export TRITONPARSE_DEBUG="1"
export TRITONPARSE_KERNEL_ALLOWLIST="my_kernel*"
export TORCHINDUCTOR_FX_GRAPH_CACHE="0"
python your_script.py

πŸ”— Related Documentation

Clone this wiki locally