Skip to content

Add LXAI Rioplatense executable-action experiment - #53

Open
krahd wants to merge 64 commits into
mainfrom
lxai-rioplatense-experiment
Open

krahd wants to merge 64 commits into
mainfrom
lxai-rioplatense-experiment

Conversation

@krahd

@krahd krahd commented Aug 25, 2026

Copy link
Copy Markdown
Owner

Adds a single-entry headless runner for the LXAI 2026 Rioplatense executable-action study plus a separately auditable matched-language suite.

Design:

  • 30 semantic cases across direct, parameterised, and state-conditioned action tasks;
  • English, broadly standardised Spanish, and Rioplatense Spanish conditions;
  • stateless Modelito/Ollama calls;
  • BatLLM production parsing and deterministic replay semantics;
  • command, action-family, parameter, and executable-consequence scoring;
  • paired standard-Spanish vs Rioplatense exact McNemar summaries;
  • shuffled condition order within each model to avoid model-loading thrash;
  • immediate JSONL persistence plus CSV, metadata, and summary JSON outputs;
  • dry-run and limited-call smoke-test modes.

Validation performed outside the repository runtime: generated runner passed python -m py_compile; suite JSON parsed successfully with 30 unique case IDs. Live Ollama execution was not run from the agent environment.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a13550fdac

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

metadata = {
"experiment": "LXAI 2026 Rioplatense executable-action fidelity",
"created_at_utc": datetime.now(timezone.utc).isoformat(),
"models": models, "conditions": args.conditions, "condition_labels": LABELS,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Record source and resolved model provenance

When the same experiment is run from another checkout or after an Ollama tag is updated, this metadata records only the requested model name and cannot identify either the scoring implementation or the model weights that produced the rows. This makes otherwise identical-looking result directories non-reproducible and can misattribute behavioural differences; capture the Git revision, relevant dependency versions, and each resolved model's immutable digest or equivalent provenance before issuing calls.

Useful? React with 👍 / 👎.

Comment on lines +344 to +346
stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
output = args.output_dir or HERE / "results" / stamp
output.mkdir(parents=True, exist_ok=True)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Create collision-resistant result directories

When two default-path runs start within the same second, both resolve to the same directory and open results.jsonl with truncating mode, so parallel experiments can overwrite or interleave each other's rows and summaries. Include sub-second precision or a unique suffix and create the directory exclusively before writing.

Useful? React with 👍 / 👎.

@@ -0,0 +1,407 @@
"""Run the LXAI 2026 Rioplatense executable-action experiment.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Record the LXAI runner in STATUS.md

This adds a durable research capability and a new important repository path, but STATUS.md still describes only the URUCON research runtime and lists only research/urucon2026/. Update the current capability and important-path sections, including the required synchronised timestamp, so the canonical snapshot remains accurate.

AGENTS.md reference: AGENTS.md:L26-L34

Useful? React with 👍 / 👎.

krahd added 27 commits August 25, 2026 02:01
krahd added 30 commits August 25, 2026 09:24

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant