Skip to content

Security: Astralchemist/Expert-Advisor-v0.5.1

Security

SECURITY.md

🔒 Security Vulnerabilities & Fixes

Document Version: 1.0
Last Audit: January 2026
Status: Action Required


📋 Summary

Severity Count Status
🔴 HIGH 5 Unfixed
🟠 MEDIUM 7 Unfixed
🟡 LOW 6 Unfixed

🔴 HIGH SEVERITY VULNERABILITIES

V-001: Floating Point Overflow in Payoff Calculation

File: st_petersburg_research.py
Line: 88
Risk: Application crash, undefined behavior

Vulnerable Code:

payoffs = np.power(2.0, heads)  # 2^1024+ causes float64 overflow

Attack Vector: If geometric distribution produces large values (k > 1024), the power calculation overflows to inf, causing downstream calculations to fail or produce incorrect results.

Fix:

MAX_HEADS = 50  # 2^50 safely within float64 range
heads = np.minimum(trials - 1, MAX_HEADS)

# Alternative: Log-space calculation for extreme precision
# log_payoff = heads * np.log(2)
# payoff = np.exp(log_payoff)  # Still overflows but cleaner

Testing:

def test_overflow_protection():
    np.random.seed(0)
    # Force extreme case
    trials = np.array([1025])
    heads = np.minimum(trials - 1, 50)
    payoffs = np.power(2.0, heads)
    assert np.isfinite(payoffs).all()

V-002: Division by Zero in Volatility Imbalance

File: powerlaw_criticality.py
Line: 211
Risk: NaN propagation, silent failures

Vulnerable Code:

imbalance = (buy_vol_s - sell_vol_s) / total_vol.replace(0, 1)

Issue: While replace(0, 1) prevents immediate crash, it produces misleading results (imbalance appears normal when volume is zero).

Fix:

with np.errstate(divide='ignore', invalid='ignore'):
    imbalance = np.where(
        total_vol > 0,
        (buy_vol_s - sell_vol_s) / total_vol,
        np.nan  # Explicit NaN when no volume data
    )

V-003: Division by Zero in R/S Analysis

File: powerlaw_extremes.py
Lines: 129-132
Risk: Runtime exception

Vulnerable Code:

S = np.std(chunk, ddof=1)
if S == 0:
    rs_val = 0  # Should be NaN, not 0
else:
    rs_val = R / S

Issue: Setting rs_val = 0 when S=0 corrupts the log-log regression since log(0) = -inf.

Fix:

S = np.std(chunk, ddof=1)
if S == 0 or np.isnan(S):
    rs_val = np.nan  # Will be filtered out of regression
else:
    rs_val = R / S

# Filter NaNs before regression
valid_mask = ~np.isnan(rs_values) & (np.array(scales) > 0)
valid_scales = np.array(scales)[valid_mask]
valid_rs = np.array(rs_values)[valid_mask]

V-004: Unbounded Memory Growth

Files: All three Python files
Lines: Various result collection loops
Risk: Out-of-memory crash for large datasets

Vulnerable Code:

results = []
for i in range(start_idx, len(df)):
    # ... calculations ...
    results.append(metric)  # Grows without bound

Fix (Option 1 - Pre-allocation):

n_results = len(df) - start_idx
results = [None] * n_results

for idx, i in enumerate(range(start_idx, len(df))):
    results[idx] = metric

Fix (Option 2 - Streaming):

def run_analysis_streaming():
    """Generator-based analysis for memory-constrained environments."""
    for i in range(start_idx, len(df)):
        metric = calculate_metric(i)
        yield metric

# Write directly to disk
with open('results.jsonl', 'w') as f:
    for metric in run_analysis_streaming():
        f.write(json.dumps(vars(metric)) + '\n')

V-005: Hardcoded Random Seeds

Files: All three Python files
Lines: RANDOM_SEED = 42|99|123
Risk: Predictable randomness, reproducibility issues in production

Vulnerable Code:

RANDOM_SEED = 42
np.random.seed(RANDOM_SEED)

Issue: Fixed seeds make all simulations deterministic and predictable. While useful for testing, this is inappropriate for research requiring statistical validity.

Fix:

import os
import secrets

# Allow override via environment variable
_seed_str = os.environ.get('RANDOM_SEED', '')
if _seed_str:
    RANDOM_SEED = int(_seed_str)
elif os.environ.get('RESEARCH_MODE') == 'deterministic':
    RANDOM_SEED = 42  # Explicit deterministic mode
else:
    RANDOM_SEED = secrets.randbelow(2**32)  # True randomness

logger.info(f"Using random seed: {RANDOM_SEED}")

🟠 MEDIUM SEVERITY VULNERABILITIES

V-006: Ambiguous Ruin Step Detection

File: st_petersburg_research.py
Line: 137
Risk: Incorrect metrics

Vulnerable Code:

ruin_step = np.argmax(path==0) if 0 in path else 'None'

Issue: np.argmax returns 0 for both "ruin at step 0" and "no zeros found".

Fix:

def get_ruin_step(path: np.ndarray) -> Optional[int]:
    """Returns first index where wealth <= 0, or None if never ruined."""
    ruin_mask = path <= 0
    if np.any(ruin_mask):
        return int(np.argmax(ruin_mask))
    return None

V-007: No Configuration Validation

Files: All configuration blocks
Risk: Silent failures, invalid states

Fix:

def validate_config():
    """Validate all configuration parameters at startup."""
    # Criticality module
    assert LOOKBACK_ALPHA > MIN_TAIL_SIZE, \
        f"LOOKBACK_ALPHA ({LOOKBACK_ALPHA}) must exceed MIN_TAIL_SIZE ({MIN_TAIL_SIZE})"
    assert 0 < TAIL_PERCENTILE < 1, \
        f"TAIL_PERCENTILE must be in (0,1), got {TAIL_PERCENTILE}"
    assert ALPHA_CHAOTIC < ALPHA_CRITICAL < ALPHA_STABLE, \
        "Alpha thresholds must be strictly ordered: CHAOTIC < CRITICAL < STABLE"
    assert COMPRESSION_THRESHOLD > 0, \
        "COMPRESSION_THRESHOLD must be positive"
    
    # Everest module
    assert BOOTSTRAP_ITERATIONS >= 10, \
        "BOOTSTRAP_ITERATIONS must be at least 10 for statistical validity"
    assert 0 < EXTREME_QUANTILE < 1, \
        "EXTREME_QUANTILE must be in (0,1)"
    
    # St. Petersburg module
    assert SIM_NUM_GAMES > 0, "SIM_NUM_GAMES must be positive"
    assert SIM_START_BANKROLL > 0, "SIM_START_BANKROLL must be positive"
    
# Call at module load
validate_config()

V-008: Hardcoded Output Paths

Files: All plt.savefig() calls
Risk: Write failures on read-only systems, file overwriting

Vulnerable Code:

plt.savefig('powerlaw_criticality_results.png')

Fix:

import os
from datetime import datetime

OUTPUT_DIR = os.environ.get('OUTPUT_DIR', os.getcwd())
os.makedirs(OUTPUT_DIR, exist_ok=True)

def get_output_path(base_name: str, extension: str = 'png') -> str:
    """Generate timestamped output path."""
    timestamp = datetime.now().strftime('%Y%m%d_%H%M%S')
    filename = f"{base_name}_{timestamp}.{extension}"
    return os.path.join(OUTPUT_DIR, filename)

# Usage
plt.savefig(get_output_path('powerlaw_criticality_results'))

V-009: Silent NaN Propagation in Regime Detection

File: powerlaw_criticality.py
Lines: 219-220
Risk: Data quality issues masked

Vulnerable Code:

if np.isnan(alpha):
    return MarketState.NORMAL  # Defaults to normal without warning

Fix:

class MarketState(Enum):
    UNKNOWN = "UNKNOWN"  # Add explicit unknown state
    NORMAL = "NORMAL"
    COMPRESSED = "COMPRESSED"
    UNSTABLE = "UNSTABLE"
    CRITICAL = "CRITICAL"
    CHAOTIC = "CHAOTIC"

def determine_regime(alpha, compression, imbalance, bar_idx=None):
    if np.isnan(alpha):
        logger.warning(f"Insufficient data for alpha at bar {bar_idx}")
        return MarketState.UNKNOWN
    # ... rest of logic

V-010: Bootstrap Threshold Edge Case

File: powerlaw_extremes.py
Lines: 190-195
Risk: Silent iteration skipping

Vulnerable Code:

if threshold <= 0: continue  # Silent skip
if len(valid_sample) < 10: continue  # Silent skip

Fix:

skipped_iterations = 0

for iter_num in range(BOOTSTRAP_ITERATIONS):
    threshold = sorted_data[k]
    if threshold <= 0:
        skipped_iterations += 1
        logger.debug(f"Bootstrap iter {iter_num}: skipped (threshold <= 0)")
        continue
    
    valid_sample = sample[sample > threshold]
    if len(valid_sample) < 10:
        skipped_iterations += 1
        logger.debug(f"Bootstrap iter {iter_num}: skipped ({len(valid_sample)} samples)")
        continue
    
    # ... calculation ...

if skipped_iterations > BOOTSTRAP_ITERATIONS * 0.5:
    logger.warning(f"Bootstrap: {skipped_iterations}/{BOOTSTRAP_ITERATIONS} iterations skipped")

V-011: Unused Strategy Configuration

File: st_petersburg_research.py
Lines: 42-49
Risk: Dead code, confusion

Vulnerable Code:

STRATEGIES = [
    {'name': 'Fixed Cost', 'type': 'fixed', 'amt': 10.0},
    {'name': 'Kelly Fraction', 'type': 'percent', 'pct': 0.01},
    {'name': 'Martingale', 'type': 'martingale', 'base': 1.0}
]
# Never used!

Fix (Option 1 - Remove):

# STRATEGIES - Removed: Not implemented in current version
# See README for planned features

Fix (Option 2 - Implement):

def simulate_with_strategy(payoffs, bankroll, strategy):
    """Run simulation with configurable betting strategy."""
    if strategy['type'] == 'fixed':
        return simulate_wealth_path(payoffs, bankroll, strategy['amt'])
    elif strategy['type'] == 'percent':
        return simulate_kelly_path(payoffs, bankroll, strategy['pct'])
    elif strategy['type'] == 'martingale':
        return simulate_martingale_path(payoffs, bankroll, strategy['base'])
    else:
        raise ValueError(f"Unknown strategy type: {strategy['type']}")

V-012: Lévy Distribution Clipping

File: powerlaw_extremes.py
Line: 75
Risk: Data distortion

Vulnerable Code:

steps = np.clip(steps, -0.05, 0.05)  # Arbitrary clipping

Issue: Clipping extreme values fundamentally changes the distribution's tail properties — the very thing we're studying!

Fix:

# Configurable clipping with warning
LEVY_CLIP_MIN = float(os.environ.get('LEVY_CLIP_MIN', '-0.10'))
LEVY_CLIP_MAX = float(os.environ.get('LEVY_CLIP_MAX', '0.10'))

original_max = np.max(np.abs(steps))
steps = np.clip(steps, LEVY_CLIP_MIN, LEVY_CLIP_MAX)

clipped_count = np.sum((steps == LEVY_CLIP_MIN) | (steps == LEVY_CLIP_MAX))
if clipped_count > 0:
    logger.warning(f"Clipped {clipped_count} extreme values (original max: {original_max:.4f})")

🟡 LOW SEVERITY VULNERABILITIES

V-013: No Logging Rotation

Files: All logging configuration
Risk: Disk space exhaustion in long runs

Fix:

from logging.handlers import RotatingFileHandler

LOG_FILE = os.environ.get('LOG_FILE', 'research.log')
LOG_MAX_BYTES = 5 * 1024 * 1024  # 5MB
LOG_BACKUP_COUNT = 3

file_handler = RotatingFileHandler(
    LOG_FILE, 
    maxBytes=LOG_MAX_BYTES, 
    backupCount=LOG_BACKUP_COUNT
)
file_handler.setFormatter(logging.Formatter(
    '%(asctime)s [%(levelname)s] %(name)s: %(message)s'
))
logger.addHandler(file_handler)

V-014: Missing Dependency Pins

Risk: Compatibility issues with future package versions

Fix: Create requirements.txt:

numpy>=1.20,<3.0
pandas>=1.3,<3.0
scipy>=1.7,<2.0
matplotlib>=3.4,<4.0

V-015: Index Collision Potential

File: powerlaw_criticality.py
Line: 315
Risk: Modifies DataFrame in place

Vulnerable Code:

stats_df.set_index('timestamp', inplace=True)

Fix:

# Create copy to avoid modifying original
plot_df = stats_df.copy()
plot_df.set_index('timestamp', inplace=True)

V-016: Symlog Scale Ambiguity

File: st_petersburg_research.py
Line: 155
Risk: Misinterpretation of negative values

Issue: symlog scale handles negative values, but wealth should never go negative (should be clamped at 0 for ruin).

Fix:

# Ensure no negative values before plotting
path = np.maximum(path, 0)

# Or use log scale with appropriate minimum
plt.yscale('log')
plt.ylim(bottom=max(1, path[path > 0].min() * 0.9))

V-017: Unused min_chunk_size Return Path

File: powerlaw_extremes.py
Line: 90, 96
Risk: Parameter appears used but always overridden

Vulnerable Code:

def calculate_hurst_rs(series, min_chunk_size=8):
    # ...
    min_scale = min_chunk_size  # Uses parameter
    # But hardcoded 8 appears elsewhere

Fix: Ensure consistent usage or document:

def calculate_hurst_rs(series: np.ndarray, min_chunk_size: int = 8) -> float:
    """
    Calculate Hurst exponent via R/S analysis.
    
    Parameters
    ----------
    series : np.ndarray
        Time series data
    min_chunk_size : int, default 8
        Minimum chunk size for R/S calculation. 
        Must be >= 8 for statistical validity.
    """
    min_chunk_size = max(8, min_chunk_size)  # Enforce minimum
    # ... rest

V-018: No Iteration/Time Limit

Risk: Infinite loops in edge cases

Fix:

import signal

class TimeoutError(Exception):
    pass

def timeout_handler(signum, frame):
    raise TimeoutError("Analysis exceeded time limit")

# Set 5 minute timeout
MAX_RUNTIME_SECONDS = int(os.environ.get('MAX_RUNTIME_SECONDS', 300))
signal.signal(signal.SIGALRM, timeout_handler)
signal.alarm(MAX_RUNTIME_SECONDS)

try:
    run_analysis()
finally:
    signal.alarm(0)  # Cancel timeout

✅ Remediation Checklist

- [ ] V-001: Cap heads in St. Petersburg
- [ ] V-002: Guard division in imbalance calculation
- [ ] V-003: Handle S=0 in R/S analysis
- [ ] V-004: Pre-allocate or stream results
- [ ] V-005: Make seeds configurable
- [ ] V-006: Fix ruin detection logic
- [ ] V-007: Add config validation
- [ ] V-008: Configurable output paths
- [ ] V-009: Add UNKNOWN market state
- [ ] V-010: Log skipped bootstrap iterations
- [ ] V-011: Remove or implement STRATEGIES
- [ ] V-012: Configurable Lévy clipping
- [ ] V-013: Add rotating log handler
- [ ] V-014: Create requirements.txt
- [ ] V-015: Don't modify DataFrames in place
- [ ] V-016: Handle negative wealth in plots
- [ ] V-017: Document min_chunk_size parameter
- [ ] V-018: Add timeout protection

📝 Notes

Not Security Vulnerabilities

The following are not vulnerabilities but may warrant attention:

  1. No authentication — These are research scripts, not services
  2. No input sanitization — No user input is processed
  3. No network access — Purely local computation
  4. No file path injection — Hardcoded paths only

Future Considerations

If these scripts are ever converted to:

  • A web service: Add authentication, rate limiting, input validation
  • A CLI tool: Add argument parsing with validation
  • A library: Add type hints, docstrings, unit tests

Document generated: January 2026

There aren't any published security advisories