A reproducible PyTorch workbench for studying rank-augmented linear attention, hybrid local/global attention, safe kernel formulas, and rank diagnostics on controlled vision and associative-recall tasks.
This repository is the standalone RALA Formula Lab. It is intentionally
separate from the broader AI_MODELS workspace. Use Sample mode in the
sidebar for an instant cached view; it reads committed results and does not
download data or train models.
This is an independent research prototype inspired by Breaking the Low-Rank Dilemma of Linear Attention. It is not the authors' official implementation and does not claim to reproduce the paper's ImageNet results.
Can rank-aware global memory and output modulation preserve useful feature diversity without giving up the efficiency advantages of linear attention?
The lab makes that question inspectable through:
- RALA, vanilla linear attention, Softmax, and a hybrid local/global variant.
- An AST-validated formula language for
kappa(x)andphi(x). - Per-layer memory, global-output, and final-output rank diagnostics.
- CIFAR-10 and synthetic associative-recall experiments.
- JSON-first experiment records and deterministic report generation.
- A Streamlit interface for configuring, running, and exporting experiments.
The strongest controlled artifact compares three variants with the same seed, model size, sample count, patch size, and 30-epoch budget.
| Variant | Best validation | Best epoch | Final validation | Recorded inference |
|---|---|---|---|---|
| Hybrid | 52.70% | 29 | 51.95% | 1333.04 ms |
| Softmax | 53.75% | 28 | 52.50% | 636.46 ms |
| Linear | 51.80% | 30 | 51.80% | 661.72 ms |
Configuration: CIFAR-10, seed 7, D=64, four heads, two layers, 10,000 training samples, patch size 4, learning rate 1e-3.
The hybrid run is 1.05 percentage points below the best Softmax checkpoint and 0.90 points above linear attention. These are single-seed observations—not a superiority claim. The recorded timings are provenance only because the current artifacts do not establish a controlled hardware benchmark.
Newer experiments add two important findings:
- Large D=512 CIFAR-10 configurations reached 46.09% best validation accuracy; they changed several variables and are not a clean ablation.
- The latest associative-recall run ended at 72.68% train versus 14.70% validation accuracy. That failure is retained as evidence of memorization without useful held-out generalization.
See the generated result report, machine-readable summary, and artifact guide for every committed run.
flowchart LR
A[Streamlit UI or benchmark script] --> B[Validated experiment config]
B --> C[Safe kappa / phi formula compiler]
C --> D{Attention variant}
D --> E[Softmax]
D --> F[Linear / RALA]
D --> G[Hybrid local + global]
E --> H[Trainer and diagnostics]
F --> H
G --> H
H --> I[Versioned JSON artifact]
I --> J[Summary tables and figures]
The implemented RALA path follows the paper's rank-augmentation idea:
Q_g = mean(Q)
alpha_j = N * softmax(Q_g kappa(K_j)^T)
B = sum_j alpha_j kappa(K_j)^T V_j
Y_i = phi(X_i) * (kappa(Q_i) B / normalizer)
The denominator includes epsilon protection. Formula input is parsed with Python's AST and permits only the documented tensor operations; imports, attributes, indexing, lambdas, comprehensions, and unknown functions are rejected.
rala-formula-lab/
├── app.py # Streamlit research interface
├── rala_lab/ # Attention, models, data, metrics, training
├── tests/ # Formula-safety and attention diagnostics
├── scripts/
│ ├── build_results.py # Validate JSON and rebuild tables/graphs
│ ├── run_benchmark.py # Programmatic benchmark entry point
│ ├── stress_test.py # Structural forward/backward smoke test
│ └── generate_runbook.py # Rebuild the DOCX experiment runbook
├── results/
│ ├── runs/ # 19 immutable experiment JSON files
│ ├── figures/ # Generated PNG and SVG figures
│ ├── LATEST_RESULTS.md # Generated evidence summary
│ └── summary.csv # Flattened machine-readable metrics
└── docs/ # Report, runbook, formulation, presentation
git clone https://github.com/0xSuleman/rala-formula-lab.git
cd rala-formula-lab
python -m venv .venvActivate the environment and install the project:
# Linux/macOS
source .venv/bin/activate
# Windows PowerShell
.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"Run the dashboard:
streamlit run app.pyFor Streamlit Community Cloud, create a new app from this repository, branch
main, and entrypoint app.py. The root requirements.txt and
.streamlit/config.toml are included for the hosted environment. See the
deployment guide and
official Streamlit documentation.
Run the verification suite and rebuild result artifacts:
pytest
python scripts/build_results.py --check
python scripts/build_results.pyDataset downloads are cached in data/. New UI runs are written to the
ignored results/generated/ directory so an experiment is reviewed before it
becomes a committed artifact.
Each exported JSON records configuration, seed, epoch history, per-layer rank statistics, final diagnostics, warnings, and recorded inference time. The portfolio tables and graphs are generated from those JSON files rather than manually copied values.
Before making a comparative claim:
- keep dataset split, seed set, parameter budget, batch size, and training budget matched;
- run multiple seeds and report mean, spread, and individual outcomes;
- separate accuracy studies from controlled latency/memory benchmarks;
- retain failed and negative runs; and
- distinguish rank correlation from causal improvements.
| Artifact | Purpose |
|---|---|
| Experiment runbook | Planned ablations and run protocol |
| Model formulation | Interactive architecture explanation |
| Presentation | Browser-based project walkthrough |
| Experiment worksheet | Editable research notebook template |
| Capacity analysis | Structural scale smoke test and its limits |
| Reference material | Upstream papers and citations (linked, not copied) |
- The committed result set uses seed 7; there are no confidence intervals.
- Several exploratory runs changed data, architecture, and optimization together, so they cannot isolate causal effects.
- CIFAR-10 subset results are research diagnostics, not state-of-the-art claims.
- Associative recall currently exposes a generalization failure.
- Existing timing values are not a hardware-controlled benchmark.
The next decisive milestone is a parameter-matched, five-seed ablation of alpha weighting, the output gate, the salience gate, and global memory against Softmax and vanilla linear attention.
- Qihang Fan, Huaibo Huang, and Ran He, Breaking the Low-Rank Dilemma of Linear Attention, CVPR 2025.
- RG-LRU/local-attention hybrid design is informed by Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.
The cited authors own their respective work. This repository contains an independent educational implementation and experimental tooling by Suleman Ahmed. See CITATION.cff for software citation metadata.
Code in this repository is available under the MIT License. External papers are cited by link only and are not redistributed by this repository.