Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
64 commits
Select commit Hold shift + click to select a range
a5aa639
Add LXAI Rioplatense matched instruction suite
krahd Aug 25, 2026
a13550f
Add single-entry LXAI Rioplatense experiment runner
krahd Aug 25, 2026
5ddd779
Prevent Kivy from consuming LXAI CLI arguments
krahd Aug 25, 2026
671c0e0
Refine LXAI protocol and preserve formatting diagnostics
krahd Aug 25, 2026
bd9f7d9
Capture exact experiment provenance
krahd Aug 25, 2026
3e67c6f
Freeze LXAI Rioplatense experiment protocol
krahd Aug 25, 2026
865285e
Record headline results from LXAI main run
krahd Aug 25, 2026
5838d1b
Store LXAI Rioplatense main experiment results
krahd Aug 25, 2026
205ab1a
Add decision-complexity follow-up suite
krahd Aug 25, 2026
18e0d29
Add protocol for LXAI decision-complexity experiment
krahd Aug 25, 2026
5dce837
Add explicit state-reasoning runner for LXAI experiment 2
krahd Aug 25, 2026
012f7a7
Clarify state-use prompt and pilot exclusion for experiment 2
krahd Aug 25, 2026
0a255a8
Use direct Ollama transport for LXAI experiments
krahd Aug 25, 2026
80aac46
Add direct Ollama runner for LXAI experiment 1
krahd Aug 25, 2026
476cb25
Use direct Ollama transport for LXAI decision experiment
krahd Aug 25, 2026
b3d662e
Document LXAI Ollama transport audit
krahd Aug 25, 2026
3e56d7d
Remove invalid LXAI Experiment 1 results
krahd Aug 25, 2026
2cf6244
Remove invalid LXAI Experiment 1 results
krahd Aug 25, 2026
0e89e2d
Remove invalid LXAI Experiment 1 results
krahd Aug 25, 2026
02ff332
Remove invalid LXAI Experiment 1 results
krahd Aug 25, 2026
671261a
Remove invalid Experiment 1 interpretation
krahd Aug 25, 2026
8dc3287
Remove invalid Experiment 1 baseline from LXAI protocol
krahd Aug 25, 2026
142a956
Clarify LXAI counter-state diagnostics before main run
krahd Aug 25, 2026
45a4b7b
Freeze LXAI decision experiment after branch-order diagnostics
krahd Aug 25, 2026
86b0692
Add LXAI experiment continuation handoff
krahd Aug 25, 2026
567d2f8
Record safe post-run branch sync sequence
krahd Aug 25, 2026
a63e421
Record completed LXAI main-run topline
krahd Aug 25, 2026
3e0d7ea
Store LXAI decision-complexity main results
krahd Aug 25, 2026
5110b65
Add reproducible LXAI decision-result diagnostics
krahd Aug 25, 2026
547c267
Record LXAI main-run results
krahd Aug 25, 2026
56a7514
Fix numeric action parsing in LXAI diagnostics
krahd Aug 25, 2026
750493c
Add LXAI policy-level diagnostics
krahd Aug 25, 2026
8042102
Add LXAI policy-level diagnostics to main results
krahd Aug 25, 2026
ed570ad
Update BatLLM LXAI handoff after completed diagnostics
krahd Aug 25, 2026
a28faa3
Add comprehensive LXAI frozen-run audit
krahd Aug 25, 2026
2eb04be
Add LXAI explicit-deliberation control
krahd Aug 25, 2026
50e3933
Add LXAI branch-order control
krahd Aug 25, 2026
2192e96
Add analysis for LXAI post-freeze controls
krahd Aug 25, 2026
2c1ff5f
Add LXAI post-freeze control results
krahd Aug 25, 2026
ca251a2
Add deliberative branch-order interaction control
krahd Aug 25, 2026
e908189
Add interface by branch-order factorial analysis
krahd Aug 25, 2026
f0b4aa3
Complete LXAI interface branch-order factorial
krahd Aug 25, 2026
4112432
Add clean symmetric clause-order factorial for LXAI full paper
krahd Aug 25, 2026
36e5861
Add analysis for symmetric clause-order factorial
krahd Aug 25, 2026
613c21c
Audit clean clause-order factorial before execution
krahd Aug 25, 2026
c659fad
Add LXAI symmetric clause-order factorial
krahd Aug 25, 2026
ea6920c
Add LXAI deliberation budget control
krahd Aug 25, 2026
2de04b0
Add LXAI symmetric factorial mechanism audit
krahd Aug 25, 2026
bc14dc4
Add LXAI deliberation budget analysis
krahd Aug 25, 2026
5d251b2
Strengthen LXAI budget-control paired diagnostics
krahd Aug 25, 2026
8714b45
Add LXAI format-only interface control
krahd Aug 25, 2026
a5b68b0
Update LXAI experiment handoff for budget control
krahd Aug 25, 2026
30112cf
Add LXAI budget and mechanism controls
krahd Aug 25, 2026
d55d37a
Add LXAI format-only control analysis
krahd Aug 25, 2026
b709e37
Update LXAI experiment handoff after budget control
krahd Aug 25, 2026
f6ef1d0
Add LXAI format-only causal control
krahd Aug 25, 2026
41a490d
Close LXAI experiment handoff after final control
krahd Aug 25, 2026
0a2550c
Add final LXAI paper claim audit
krahd Aug 25, 2026
e7feff6
Audit final LXAI budget and length-distribution claims
krahd Aug 25, 2026
086703f
Fix LXAI quartile audit convention
krahd Aug 25, 2026
e5d2907
Use higher empirical quartiles in LXAI audit
krahd Aug 25, 2026
7fa6f8f
Extend LXAI audit for post-repair diagnostics
krahd Aug 25, 2026
df20b8f
Audit final LXAI policy scope and repaired-cell plot
krahd Aug 25, 2026
82f7685
Audit LXAI sentinel-contaminated deliberative diagnostics
krahd Aug 25, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
114 changes: 114 additions & 0 deletions research/lxai2026/HANDOFF-2026-08-25.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,114 @@
# LXAI 2026 experiment handoff

**Branch:** `lxai-rioplatense-experiment`
**State:** all authorised LXAI experiments complete; science frozen; no further model execution for this submission.
**Paper owner:** `krahd/research/academic-writing/my_papers_2026/2026 - LXAI NeurIPS 2026/`.
**Do not alter the frozen main dataset.**

## Read first

1. `controls/FORMAT-ONLY-CONTROL.md`
2. `controls/DELIBERATION-BUDGET-CONTROL.md`
3. `controls/SYMMETRIC-FACTORIAL-MECHANISM-AUDIT.md`
4. `controls/SYMMETRIC-CLAUSE-FACTORIAL.md`
5. paper-side `HANDOFF-2026-08-25.md`
6. paper-side `FULL-PAPER-FINAL-AUDIT-POST-CONTROLS-2026-08-25.md`

## Frozen main dataset — immutable

`research/lxai2026/results/20260825T094526Z/`

576 calls; zero provider errors; 217 strict correct; 569/576 valid. Pooled strict correctness: English 40.6%, tuteo 35.4%, Rioplatense voseo 37.0%. Spanish combined: 73/96 state-invariant groups and 0/96 policy-complete.

Legacy key `es_standard` means **tuteo**. Never call it “standardised Spanish”. Deleted Modelito data and all pre-freeze pilot/debug calls remain excluded.

## Clean clause-order result

Run: `controls/symmetric_clause_factorial/20260825T165129Z/`
Analysis: `controls/SYMMETRIC-CLAUSE-FACTORIAL.md`
Mechanism audit: `controls/SYMMETRIC-FACTORIAL-MECHANISM-AUDIT.md`

The two order prompts contain the same two clause strings verbatim. Truth is first in exactly half the cells.

| Interface | Overall | true clause first | true clause second |
|---|---:|---:|---:|
| direct | 135/256 = 52.7% | 116/128 = 90.6% | 19/128 = 14.8% |
| format-only | 130/256 = 50.8% | 99/128 = 77.3% | 31/128 = 24.2% |
| explicit deliberation, 256 | 238/256 = 93.0% | 120/128 = 93.8% | 118/128 = 92.2% |

Interpretation: direct output is strongly position-dominated; a `FINAL:` wrapper with intermediate text forbidden changes some position behaviour but does not recover accuracy; the large rescue is associated with the explicit intermediate-deliberation condition.

Do not claim visible scratchpad text is uniquely causal: the deliberative system prompt also changes the instruction to deliberate.

## Format-only control — COMPLETE

Canonical: `controls/FORMAT-ONLY-CONTROL.md/.json`
Raw run: `controls/format_only/20260825T194809Z/`

- correct 130/256 = 50.8%;
- valid 256/256;
- first-position selection 171/256 = 66.8%;
- output changes across exact order swap 83/128 = 64.8%;
- correct both orders 27/128 = 21.1%;
- follows first in both 58/128 = 45.3%;
- policy-complete 3/64 = 4.7%;
- state-invariant 40/64 = 62.5%.

This closes the most obvious final-output-wrapper alternative.

## 1024-token budget control — COMPLETE

Canonical: `controls/DELIBERATION-BUDGET-CONTROL.md/.json`
Raw: `controls/deliberation_budget/20260825T185106Z/`

Original 256 deliberative arm:

- 238/256 correct;
- 241/256 valid;
- 15 final-line extraction failures.

Paired 1024 rerun:

- 253/256 = 98.8% correct;
- 256/256 valid;
- 0 extraction failures;
- all end with `done_reason=stop`;
- 15/15 old failures become valid and correct;
- 14/15 old failed strings are exact prefixes of the longer completion;
- 240/241 old successful full responses are byte-for-byte unchanged;
- 241/241 old successful final commands remain unchanged.

Language at 1024:

- pooled tuteo 127/128 = 99.2%;
- pooled voseo 126/128 = 98.4%;
- Qwen30 tuteo 64/64; voseo 63/64.

Qwen30 mean generated tokens: tuteo 119.9, voseo 165.6. Mistral: 63.8 vs 64.1.

Interpretation: the old 15 extraction failures were fixed-output-budget truncations. The earlier raw 7.8-point deliberative voseo deficit was primarily manufactured by a model-specific language-conditioned generation-length difference interacting with the cap.

Do not claim general novelty: Goyal & Ray, *Mind the Cap* (arXiv:2608.04160, 2026), is direct prior art for multilingual output-budget effects.

## Experiment stop rule

**No more model runs for the LXAI submission.**

Do not add repeats, thinking checkpoints, schema localisation, extra languages/dialects, more models, or another reasoning condition. These are future archival-extension work unless a genuinely fatal implementation error is discovered.

## Paper handoff

The paper repo has already been rewritten after these final controls:

- title: *Output Interface, Clause Order, and Inference Budget in State-Conditioned Action Selection across Tuteo and Rioplatense Voseo*;
- `PAPER-FULL-DRAFT.md` synchronised;
- `submission/paper.tex` rewritten;
- `submission/references.bib` includes `Mind the Cap`;
- `submission/checklist.tex` updated;
- `FULL-PAPER-FINAL-AUDIT-POST-CONTROLS-2026-08-25.md` records final claims/limits.

Next work is only official NeurIPS-template compilation, 6–8-page fit, PDF/anonymity/font preflight, live OpenReview verification, and submission.

## Interpretation constraints

No population-level model inference; no tuteo/voseo equivalence claim; no broad Rioplatense robustness claim; no novelty claim for position bias, deliberation, constrained-output costs, multilingual tool calling, executable evaluation, or output-budget multilingual gaps; no identifying repository/project metadata in double-blind submission artefacts.
146 changes: 146 additions & 0 deletions research/lxai2026/MAIN-RUN-RESULTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,146 @@
# LXAI 2026 main-run results

**Dataset:** `results/20260825T094526Z/`
**State:** frozen direct-transport main run complete; 576/576 calls; zero provider errors.
**Raw files:** `metadata.json`, `results.csv`, `results.jsonl`, `summary.json`.
**Derived diagnostics:** `results/20260825T094526Z/DECISION-ANALYSIS.md`, `decision_analysis.json`.

## Strict command correctness

| Model | English | Standard Spanish | Rioplatense | Std − Rio |
|---|---:|---:|---:|---:|
| Mistral Small 3.2 24B | 56.2% | 50.0% | 41.7% | +8.3 pp |
| Qwen3 30B-A3B | 43.8% | 45.8% | 50.0% | −4.2 pp |
| Qwen3 4B | 50.0% | 35.4% | 35.4% | 0.0 pp |
| Llama 3.2 | 12.5% | 10.4% | 20.8% | −10.4 pp |

Across the 192 matched standard-Spanish/Rioplatense case/model pairs:

- standard Spanish: 68/192 = 35.4%;
- Rioplatense: 71/192 = 37.0%;
- standard-only correct: 9;
- Rioplatense-only correct: 12;
- standard minus Rioplatense: −1.56 percentage points;
- exact McNemar p = 0.6636.

No prospectively defined equivalence margin exists. The nonsignificant paired result is therefore not an equivalence result.

Per-model exact McNemar results for strict correctness:

- Mistral 24B: standard-only 6, Rioplatense-only 2, p = 0.2891;
- Qwen3 30B-A3B: standard-only 2, Rioplatense-only 4, p = 0.6875;
- Qwen3 4B: no discordant Spanish pairs; p undefined;
- Llama 3.2: standard-only 1, Rioplatense-only 6, p = 0.1250.

The direction is heterogeneous across the four fixed model conditions. Do not infer a model-population effect from the pooled descriptive result.

## Strict correctness by decision-structure tier

| Tier | Standard Spanish | Rioplatense | Std − Rio | Exact McNemar p |
|---|---:|---:|---:|---:|
| D1 | 40.6% | 40.6% | 0.0 pp | 1.000 |
| D2 | 43.8% | 43.8% | 0.0 pp | n/a (no discordance) |
| D3 | 37.5% | 31.2% | +6.25 pp | 0.688 |
| D4 | 40.6% | 46.9% | −6.25 pp | 0.500 |
| D5 | 28.1% | 37.5% | −9.38 pp | 0.375 |
| D6 | 21.9% | 21.9% | 0.0 pp | 1.000 |

There is no monotonic Rioplatense penalty as D1–D6 changes. D1 and D2 are identical across Spanish conditions; the sign changes at D3/D4; Rioplatense is descriptively higher at D5; D6 is identical. D1–D6 are designed structural strata, not a validated interval-scale complexity measure.

## Other outcome metrics

### Action-family correctness

Across 192 Spanish pairs:

- standard Spanish: 40.1%;
- Rioplatense: 43.2%;
- standard minus Rioplatense: −3.13 pp;
- standard-only 9, Rioplatense-only 15;
- exact McNemar p = 0.3075.

### Executable-consequence correctness

Across 192 Spanish pairs:

- standard Spanish: 43.8%;
- Rioplatense: 42.2%;
- standard minus Rioplatense: +1.56 pp;
- standard-only 11, Rioplatense-only 8;
- exact McNemar p = 0.6476.

The sign of the small pooled Spanish difference therefore depends on the outcome definition: strict command correctness slightly favours Rioplatense, while executable-consequence correctness slightly favours standard Spanish. Neither difference is statistically decisive. This is evidence against a simple stable pooled regional-variety direction in this bounded suite.

Executable correctness by tier:

| Tier | Standard Spanish | Rioplatense | Std − Rio | Exact McNemar p |
|---|---:|---:|---:|---:|
| D1 | 56.2% | 53.1% | +3.12 pp | 1.000 |
| D2 | 50.0% | 50.0% | 0.0 pp | n/a |
| D3 | 40.6% | 34.4% | +6.25 pp | 0.688 |
| D4 | 56.2% | 46.9% | +9.38 pp | 0.250 |
| D5 | 34.4% | 43.8% | −9.38 pp | 0.375 |
| D6 | 25.0% | 25.0% | 0.0 pp | 1.000 |

### Validity and failure classes

- Correct strict commands: 217/576 = 37.7%.
- Valid but wrong action (`action_error`): 352.
- Invalid/unrecoverable: 7.
- Format-only: 0.
- Format + action error: 0.
- Provider errors: 0.

All Mistral/Qwen responses were valid. The seven invalid outputs were confined to Llama 3.2. The dominant failure is therefore action selection rather than provider or formatting failure.

## Policy-level state-use diagnostics

The prospectively declared diagnostics were derived from the frozen CSV by `analyze_decision_results.py`; the derived files are committed beside the raw results.

| Condition | Policy-complete | State-invariant | First-action selected | Oracle first-action | Excess |
|---|---:|---:|---:|---:|---:|
| English | 3/48 (6.2%) | 37/48 (77.1%) | 63.0% | 43.8% | +19.3 pp |
| Standard Spanish | 0/48 (0.0%) | 37/48 (77.1%) | 57.8% | 43.8% | +14.1 pp |
| Rioplatense | 0/48 (0.0%) | 36/48 (75.0%) | 59.9% | 43.8% | +16.1 pp |

A policy group is `model × language condition × policy`, evaluated over its four counter-states. In the two Spanish conditions combined, none of 96 groups answered all four counter-states correctly, and 73/96 (76.0%) emitted one invariant command despite the oracle requiring more than one command across the four states. First-mentioned actions were selected substantially more often than their oracle frequency in both Spanish conditions.

This is the strongest explanatory result of the study: the models frequently fail to condition action selection on the supplied state. The state-use failure is large in both Spanish varieties and is much larger than the small pooled standard-vs-Rioplatense difference.

### Spanish policy diagnostics by model

| Model | Condition | Policy-complete | State-invariant | First-action excess |
|---|---|---:|---:|---:|
| Llama 3.2 | Standard | 0/12 | 10/12 (83.3%) | −29.2 pp |
| Llama 3.2 | Rioplatense | 0/12 | 8/12 (66.7%) | −16.7 pp |
| Mistral 24B | Standard | 0/12 | 7/12 (58.3%) | +14.6 pp |
| Mistral 24B | Rioplatense | 0/12 | 9/12 (75.0%) | +25.0 pp |
| Qwen3 30B-A3B | Standard | 0/12 | 11/12 (91.7%) | +50.0 pp |
| Qwen3 30B-A3B | Rioplatense | 0/12 | 9/12 (75.0%) | +37.5 pp |
| Qwen3 4B | Standard | 0/12 | 9/12 (75.0%) | +20.8 pp |
| Qwen3 4B | Rioplatense | 0/12 | 10/12 (83.3%) | +18.8 pp |

The first-action diagnostic is explanatory rather than a universal bias measure: Llama often selected other recurrent actions, yielding negative excess, while the Qwen models showed especially strong first-action excess. The common cross-model feature is state-invariant output, not one single lexical heuristic.

### Spanish state invariance by tier

| Tier | Standard | Rioplatense |
|---|---:|---:|
| D1 | 6/8 (75.0%) | 5/8 (62.5%) |
| D2 | 7/8 (87.5%) | 8/8 (100.0%) |
| D3 | 5/8 (62.5%) | 4/8 (50.0%) |
| D4 | 7/8 (87.5%) | 6/8 (75.0%) |
| D5 | 5/8 (62.5%) | 6/8 (75.0%) |
| D6 | 7/8 (87.5%) | 7/8 (87.5%) |

State-invariant behaviour is already prevalent at D1 and remains high across the designed strata. This prevents interpreting a raw D1→D6 decline as a clean effect of increasing decision complexity.

## Interpretation supported by the full analysis

The main run does not show the hypothesised pooled disadvantage for voseo-marked Rioplatense instructions, nor a stable pooled regional-variety direction. More importantly, the counter-state analysis shows that model outputs are usually insensitive to state changes that require different actions.

The strongest bounded interpretation is:

> In this controlled executable suite, regional Spanish marking is not the dominant source of failure. Across both standardised and voseo-marked Rioplatense Spanish, models frequently fail to condition their selected action on the supplied state; state-invariant policy execution and model-specific action-selection behaviour dominate the small, unstable regional-variety difference.

Do not convert this into an equivalence claim, a statement that models “understand” Rioplatense, a general robustness claim, or a claim about internal mechanisms. The experiment is behavioural and bounded to these constructions, policies and four fixed local model conditions.
Loading
Loading