Document Version: 1.0
Last Audit: January 2026
Status: Action Required
| Severity | Count | Status |
|---|---|---|
| 🔴 HIGH | 5 | Unfixed |
| 🟠 MEDIUM | 7 | Unfixed |
| 🟡 LOW | 6 | Unfixed |
File: st_petersburg_research.py
Line: 88
Risk: Application crash, undefined behavior
Vulnerable Code:
payoffs = np.power(2.0, heads) # 2^1024+ causes float64 overflowAttack Vector: If geometric distribution produces large values (k > 1024), the power calculation overflows to inf, causing downstream calculations to fail or produce incorrect results.
Fix:
MAX_HEADS = 50 # 2^50 safely within float64 range
heads = np.minimum(trials - 1, MAX_HEADS)
# Alternative: Log-space calculation for extreme precision
# log_payoff = heads * np.log(2)
# payoff = np.exp(log_payoff) # Still overflows but cleanerTesting:
def test_overflow_protection():
np.random.seed(0)
# Force extreme case
trials = np.array([1025])
heads = np.minimum(trials - 1, 50)
payoffs = np.power(2.0, heads)
assert np.isfinite(payoffs).all()File: powerlaw_criticality.py
Line: 211
Risk: NaN propagation, silent failures
Vulnerable Code:
imbalance = (buy_vol_s - sell_vol_s) / total_vol.replace(0, 1)Issue: While replace(0, 1) prevents immediate crash, it produces misleading results (imbalance appears normal when volume is zero).
Fix:
with np.errstate(divide='ignore', invalid='ignore'):
imbalance = np.where(
total_vol > 0,
(buy_vol_s - sell_vol_s) / total_vol,
np.nan # Explicit NaN when no volume data
)File: powerlaw_extremes.py
Lines: 129-132
Risk: Runtime exception
Vulnerable Code:
S = np.std(chunk, ddof=1)
if S == 0:
rs_val = 0 # Should be NaN, not 0
else:
rs_val = R / SIssue: Setting rs_val = 0 when S=0 corrupts the log-log regression since log(0) = -inf.
Fix:
S = np.std(chunk, ddof=1)
if S == 0 or np.isnan(S):
rs_val = np.nan # Will be filtered out of regression
else:
rs_val = R / S
# Filter NaNs before regression
valid_mask = ~np.isnan(rs_values) & (np.array(scales) > 0)
valid_scales = np.array(scales)[valid_mask]
valid_rs = np.array(rs_values)[valid_mask]Files: All three Python files
Lines: Various result collection loops
Risk: Out-of-memory crash for large datasets
Vulnerable Code:
results = []
for i in range(start_idx, len(df)):
# ... calculations ...
results.append(metric) # Grows without boundFix (Option 1 - Pre-allocation):
n_results = len(df) - start_idx
results = [None] * n_results
for idx, i in enumerate(range(start_idx, len(df))):
results[idx] = metricFix (Option 2 - Streaming):
def run_analysis_streaming():
"""Generator-based analysis for memory-constrained environments."""
for i in range(start_idx, len(df)):
metric = calculate_metric(i)
yield metric
# Write directly to disk
with open('results.jsonl', 'w') as f:
for metric in run_analysis_streaming():
f.write(json.dumps(vars(metric)) + '\n')Files: All three Python files
Lines: RANDOM_SEED = 42|99|123
Risk: Predictable randomness, reproducibility issues in production
Vulnerable Code:
RANDOM_SEED = 42
np.random.seed(RANDOM_SEED)Issue: Fixed seeds make all simulations deterministic and predictable. While useful for testing, this is inappropriate for research requiring statistical validity.
Fix:
import os
import secrets
# Allow override via environment variable
_seed_str = os.environ.get('RANDOM_SEED', '')
if _seed_str:
RANDOM_SEED = int(_seed_str)
elif os.environ.get('RESEARCH_MODE') == 'deterministic':
RANDOM_SEED = 42 # Explicit deterministic mode
else:
RANDOM_SEED = secrets.randbelow(2**32) # True randomness
logger.info(f"Using random seed: {RANDOM_SEED}")File: st_petersburg_research.py
Line: 137
Risk: Incorrect metrics
Vulnerable Code:
ruin_step = np.argmax(path==0) if 0 in path else 'None'Issue: np.argmax returns 0 for both "ruin at step 0" and "no zeros found".
Fix:
def get_ruin_step(path: np.ndarray) -> Optional[int]:
"""Returns first index where wealth <= 0, or None if never ruined."""
ruin_mask = path <= 0
if np.any(ruin_mask):
return int(np.argmax(ruin_mask))
return NoneFiles: All configuration blocks
Risk: Silent failures, invalid states
Fix:
def validate_config():
"""Validate all configuration parameters at startup."""
# Criticality module
assert LOOKBACK_ALPHA > MIN_TAIL_SIZE, \
f"LOOKBACK_ALPHA ({LOOKBACK_ALPHA}) must exceed MIN_TAIL_SIZE ({MIN_TAIL_SIZE})"
assert 0 < TAIL_PERCENTILE < 1, \
f"TAIL_PERCENTILE must be in (0,1), got {TAIL_PERCENTILE}"
assert ALPHA_CHAOTIC < ALPHA_CRITICAL < ALPHA_STABLE, \
"Alpha thresholds must be strictly ordered: CHAOTIC < CRITICAL < STABLE"
assert COMPRESSION_THRESHOLD > 0, \
"COMPRESSION_THRESHOLD must be positive"
# Everest module
assert BOOTSTRAP_ITERATIONS >= 10, \
"BOOTSTRAP_ITERATIONS must be at least 10 for statistical validity"
assert 0 < EXTREME_QUANTILE < 1, \
"EXTREME_QUANTILE must be in (0,1)"
# St. Petersburg module
assert SIM_NUM_GAMES > 0, "SIM_NUM_GAMES must be positive"
assert SIM_START_BANKROLL > 0, "SIM_START_BANKROLL must be positive"
# Call at module load
validate_config()Files: All plt.savefig() calls
Risk: Write failures on read-only systems, file overwriting
Vulnerable Code:
plt.savefig('powerlaw_criticality_results.png')Fix:
import os
from datetime import datetime
OUTPUT_DIR = os.environ.get('OUTPUT_DIR', os.getcwd())
os.makedirs(OUTPUT_DIR, exist_ok=True)
def get_output_path(base_name: str, extension: str = 'png') -> str:
"""Generate timestamped output path."""
timestamp = datetime.now().strftime('%Y%m%d_%H%M%S')
filename = f"{base_name}_{timestamp}.{extension}"
return os.path.join(OUTPUT_DIR, filename)
# Usage
plt.savefig(get_output_path('powerlaw_criticality_results'))File: powerlaw_criticality.py
Lines: 219-220
Risk: Data quality issues masked
Vulnerable Code:
if np.isnan(alpha):
return MarketState.NORMAL # Defaults to normal without warningFix:
class MarketState(Enum):
UNKNOWN = "UNKNOWN" # Add explicit unknown state
NORMAL = "NORMAL"
COMPRESSED = "COMPRESSED"
UNSTABLE = "UNSTABLE"
CRITICAL = "CRITICAL"
CHAOTIC = "CHAOTIC"
def determine_regime(alpha, compression, imbalance, bar_idx=None):
if np.isnan(alpha):
logger.warning(f"Insufficient data for alpha at bar {bar_idx}")
return MarketState.UNKNOWN
# ... rest of logicFile: powerlaw_extremes.py
Lines: 190-195
Risk: Silent iteration skipping
Vulnerable Code:
if threshold <= 0: continue # Silent skip
if len(valid_sample) < 10: continue # Silent skipFix:
skipped_iterations = 0
for iter_num in range(BOOTSTRAP_ITERATIONS):
threshold = sorted_data[k]
if threshold <= 0:
skipped_iterations += 1
logger.debug(f"Bootstrap iter {iter_num}: skipped (threshold <= 0)")
continue
valid_sample = sample[sample > threshold]
if len(valid_sample) < 10:
skipped_iterations += 1
logger.debug(f"Bootstrap iter {iter_num}: skipped ({len(valid_sample)} samples)")
continue
# ... calculation ...
if skipped_iterations > BOOTSTRAP_ITERATIONS * 0.5:
logger.warning(f"Bootstrap: {skipped_iterations}/{BOOTSTRAP_ITERATIONS} iterations skipped")File: st_petersburg_research.py
Lines: 42-49
Risk: Dead code, confusion
Vulnerable Code:
STRATEGIES = [
{'name': 'Fixed Cost', 'type': 'fixed', 'amt': 10.0},
{'name': 'Kelly Fraction', 'type': 'percent', 'pct': 0.01},
{'name': 'Martingale', 'type': 'martingale', 'base': 1.0}
]
# Never used!Fix (Option 1 - Remove):
# STRATEGIES - Removed: Not implemented in current version
# See README for planned featuresFix (Option 2 - Implement):
def simulate_with_strategy(payoffs, bankroll, strategy):
"""Run simulation with configurable betting strategy."""
if strategy['type'] == 'fixed':
return simulate_wealth_path(payoffs, bankroll, strategy['amt'])
elif strategy['type'] == 'percent':
return simulate_kelly_path(payoffs, bankroll, strategy['pct'])
elif strategy['type'] == 'martingale':
return simulate_martingale_path(payoffs, bankroll, strategy['base'])
else:
raise ValueError(f"Unknown strategy type: {strategy['type']}")File: powerlaw_extremes.py
Line: 75
Risk: Data distortion
Vulnerable Code:
steps = np.clip(steps, -0.05, 0.05) # Arbitrary clippingIssue: Clipping extreme values fundamentally changes the distribution's tail properties — the very thing we're studying!
Fix:
# Configurable clipping with warning
LEVY_CLIP_MIN = float(os.environ.get('LEVY_CLIP_MIN', '-0.10'))
LEVY_CLIP_MAX = float(os.environ.get('LEVY_CLIP_MAX', '0.10'))
original_max = np.max(np.abs(steps))
steps = np.clip(steps, LEVY_CLIP_MIN, LEVY_CLIP_MAX)
clipped_count = np.sum((steps == LEVY_CLIP_MIN) | (steps == LEVY_CLIP_MAX))
if clipped_count > 0:
logger.warning(f"Clipped {clipped_count} extreme values (original max: {original_max:.4f})")Files: All logging configuration
Risk: Disk space exhaustion in long runs
Fix:
from logging.handlers import RotatingFileHandler
LOG_FILE = os.environ.get('LOG_FILE', 'research.log')
LOG_MAX_BYTES = 5 * 1024 * 1024 # 5MB
LOG_BACKUP_COUNT = 3
file_handler = RotatingFileHandler(
LOG_FILE,
maxBytes=LOG_MAX_BYTES,
backupCount=LOG_BACKUP_COUNT
)
file_handler.setFormatter(logging.Formatter(
'%(asctime)s [%(levelname)s] %(name)s: %(message)s'
))
logger.addHandler(file_handler)Risk: Compatibility issues with future package versions
Fix: Create requirements.txt:
numpy>=1.20,<3.0
pandas>=1.3,<3.0
scipy>=1.7,<2.0
matplotlib>=3.4,<4.0
File: powerlaw_criticality.py
Line: 315
Risk: Modifies DataFrame in place
Vulnerable Code:
stats_df.set_index('timestamp', inplace=True)Fix:
# Create copy to avoid modifying original
plot_df = stats_df.copy()
plot_df.set_index('timestamp', inplace=True)File: st_petersburg_research.py
Line: 155
Risk: Misinterpretation of negative values
Issue: symlog scale handles negative values, but wealth should never go negative (should be clamped at 0 for ruin).
Fix:
# Ensure no negative values before plotting
path = np.maximum(path, 0)
# Or use log scale with appropriate minimum
plt.yscale('log')
plt.ylim(bottom=max(1, path[path > 0].min() * 0.9))File: powerlaw_extremes.py
Line: 90, 96
Risk: Parameter appears used but always overridden
Vulnerable Code:
def calculate_hurst_rs(series, min_chunk_size=8):
# ...
min_scale = min_chunk_size # Uses parameter
# But hardcoded 8 appears elsewhereFix: Ensure consistent usage or document:
def calculate_hurst_rs(series: np.ndarray, min_chunk_size: int = 8) -> float:
"""
Calculate Hurst exponent via R/S analysis.
Parameters
----------
series : np.ndarray
Time series data
min_chunk_size : int, default 8
Minimum chunk size for R/S calculation.
Must be >= 8 for statistical validity.
"""
min_chunk_size = max(8, min_chunk_size) # Enforce minimum
# ... restRisk: Infinite loops in edge cases
Fix:
import signal
class TimeoutError(Exception):
pass
def timeout_handler(signum, frame):
raise TimeoutError("Analysis exceeded time limit")
# Set 5 minute timeout
MAX_RUNTIME_SECONDS = int(os.environ.get('MAX_RUNTIME_SECONDS', 300))
signal.signal(signal.SIGALRM, timeout_handler)
signal.alarm(MAX_RUNTIME_SECONDS)
try:
run_analysis()
finally:
signal.alarm(0) # Cancel timeout- [ ] V-001: Cap heads in St. Petersburg
- [ ] V-002: Guard division in imbalance calculation
- [ ] V-003: Handle S=0 in R/S analysis
- [ ] V-004: Pre-allocate or stream results
- [ ] V-005: Make seeds configurable
- [ ] V-006: Fix ruin detection logic
- [ ] V-007: Add config validation
- [ ] V-008: Configurable output paths
- [ ] V-009: Add UNKNOWN market state
- [ ] V-010: Log skipped bootstrap iterations
- [ ] V-011: Remove or implement STRATEGIES
- [ ] V-012: Configurable Lévy clipping
- [ ] V-013: Add rotating log handler
- [ ] V-014: Create requirements.txt
- [ ] V-015: Don't modify DataFrames in place
- [ ] V-016: Handle negative wealth in plots
- [ ] V-017: Document min_chunk_size parameter
- [ ] V-018: Add timeout protectionThe following are not vulnerabilities but may warrant attention:
- No authentication — These are research scripts, not services
- No input sanitization — No user input is processed
- No network access — Purely local computation
- No file path injection — Hardcoded paths only
If these scripts are ever converted to:
- A web service: Add authentication, rate limiting, input validation
- A CLI tool: Add argument parsing with validation
- A library: Add type hints, docstrings, unit tests
Document generated: January 2026