Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions skills/configure/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -177,8 +177,8 @@ user runs the dry-run and returns its output before final approval.
| Installation and authentication | `../fireworks-training/references/getting-started.md` |
| Method and data selection | `../fireworks-training/references/choose-method.md` |
| Preference data and evaluators | `../fireworks-training/references/preference-data-and-evaluators.md` |
| Managed RFT | `../fireworks-training/references/managed-rft-operations.md` |
| RFT tracing | `../fireworks-training/references/rft-agent-tracing.md` |
| Managed RFT (deprecated; existing jobs only) | `../fireworks-training/references/managed-rft-operations.md` |
| RFT tracing (deprecated; existing jobs only) | `../fireworks-training/references/rft-agent-tracing.md` |
| Training API | `../fireworks-training/references/training-api.md` |
| Training API losses | `../fireworks-training/references/training-api-losses.md` |
| Secure training | `../fireworks-training/references/secure-training-operations.md` |
Expand Down
17 changes: 7 additions & 10 deletions skills/fireworks-training/references/choose-method.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ JSONL, one object per line; OpenAI-style `messages`. **Min 3, max 3M** (aim for
{"messages":[{"role":"system","content":"You are a helpful assistant."},{"role":"user","content":"Capital of France?"},{"role":"assistant","content":"Paris."}]}
```

Docs: https://docs.fireworks.ai/fine-tuning/fine-tuning-models.md. For managed RFT weighting and launch controls, read `managed-rft-operations.md` and the installed CLI help.
Docs: https://docs.fireworks.ai/fine-tuning/fine-tuning-models.md.

## DPO format

Expand All @@ -57,9 +57,11 @@ Preference pairs, **one-turn only** (preferred/non-preferred must be the last as

If the user has prompts but no preference pairs, do not reject the task or silently invent labels. Use `references/preference-data-and-evaluators.md` to plan, cost, generate, review, and preserve pair provenance before upload.

## RFT — reinforcement fine-tuning
## RL — reinforcement learning

Provide three things (not necessarily labeled outputs): a **dataset** of prompts; an **evaluator or inline reward** that scores an output 0.0→1.0; and the **agent** being trained. Managed RFT uses a registered evaluator. Training API RFT uses reward code and may read `ground_truth` or any other field declared by that reward. Start with **200–500 diverse prompts**. Docs: https://docs.fireworks.ai/fine-tuning/reinforcement-fine-tuning-models.md. Evaluator authoring: `preference-data-and-evaluators.md`.
Provide three things (not necessarily labeled outputs): a **dataset** of prompts; a **reward** that scores an output 0.0→1.0; and the **agent** being trained. Reward code may read `ground_truth` or any other field it declares. Start with **200–500 diverse prompts**. Docs: https://docs.fireworks.ai/fine-tuning/training-api/cookbook/rl.md. Rollout and scheduling detail: `rl-async.md`; multi-turn agents: `rl-agentic.md`.

Managed RFT (`firectl rftj`, registered evaluators) is deprecated and accepts no new jobs. Route new work to the Training API; use `managed-rft-operations.md` only to monitor or recover a job that already exists.

## Classification (a common SFT task)

Expand Down Expand Up @@ -119,8 +121,7 @@ Catch format errors locally before `firectl dataset create`. A malformed row oth
```python
import json, sys

method = "sft" # "sft" | "dpo" | "managed-rft" | "sdk-rft"
managed_evaluator_required_fields = [] # from the reviewed evaluator contract
method = "sft" # "sft" | "dpo" | "sdk-rft"
sdk_reward_required_fields = [] # e.g. ["ground_truth"]
allowed_roles = {"system", "user", "assistant", "tool"}

Expand Down Expand Up @@ -175,10 +176,6 @@ for i, line in enumerate(open(sys.argv[1]), 1):
assert len(dpo_turns) == 1 and dpo_turns[0]["role"] == "user", f"line {i}: DPO input must contain exactly one user turn"
validate_preference_output(o.get("preferred_output"), i, "preferred_output")
validate_preference_output(o.get("non_preferred_output"), i, "non_preferred_output")
elif method == "managed-rft":
validate_messages(o.get("messages"), i, final_assistant=False)
for field in managed_evaluator_required_fields:
assert field in o, f"line {i}: missing {field!r} required by evaluator"
elif method == "sdk-rft":
validate_messages(o.get("messages"), i, final_assistant=False)
for field in sdk_reward_required_fields:
Expand All @@ -195,7 +192,7 @@ first = json.loads(next(l for l in open(sys.argv[1]) if l.strip()))
detected = "dpo" if ("preferred_output" in first or "chosen" in first) else "sft/rft"
if method == "dpo" and detected != "dpo":
warnings.append("requested method=dpo but rows look SFT/RFT-shaped (no preferred_output/chosen) -> wrong method or wrong file")
if method in ("sft", "managed-rft", "sdk-rft") and detected == "dpo":
if method in ("sft", "sdk-rft") and detected == "dpo":
warnings.append(f"requested method={method} but rows look DPO-shaped (preferred_output/chosen present) -> wrong method")

print(f"OK: {n} valid {method} rows")
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ A fine-tuned LoRA **cannot run on serverless** — it needs an **on-demand (dedi
| Perf | Matches base | Slightly higher TTFT; lower max throughput |
| Best for | Single model in prod | Experiments / many variants |

**One adapter → live merge** (simplest). Always pass a deployment shape: a bare `firectl deployment create <model>` drops into an **interactive shape picker**, and choosing "Create without using shape" fails with `accelerator_type must be specified for non-embeddings engines`. The interactive prompt also breaks non-interactive / agent / CI use, so pass the shape explicitly and add `--wait`. Find a deployable shape first:
**One adapter → live merge** (simplest). Always pass a deployment shape: a bare `firectl deployment create <model>` drops into an **interactive shape picker**, and choosing "Create without using shape" fails with `accelerator_type must be specified for non-embeddings engines`. The interactive prompt also breaks non-interactive / agent / CI use, so pass the shape explicitly and add `--wait`. Either find a deployable shape, or pass `--deployment-shape default` to let the server pick one:
```bash
firectl deployment-shape-version match --model "accounts/<ACCOUNT_ID>/models/<FINE_TUNED_MODEL_ID>"
```
Expand Down
12 changes: 9 additions & 3 deletions skills/fireworks-training/references/managed-rft-operations.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,14 @@
# Managed RFT: launch, monitor, and validate
# Managed RFT: launch, monitor, and validate (deprecated)

*Source of truth: live [RFT overview](https://docs.fireworks.ai/fine-tuning/reinforcement-fine-tuning-models.md), [Models matrix](https://docs.fireworks.ai/fine-tuning/models.md), [RFT parameters](https://docs.fireworks.ai/fine-tuning/rft-parameters-reference.md), and [Eval Protocol](https://evalprotocol.io/introduction). Defer flags and defaults to installed CLI `--help`.*
> **Managed RFT is deprecated. Do not route a new run here.** Reinforcement learning
> has moved to the Training API: read `references/rl-async.md` and
> [Cookbook: Reinforcement Learning](https://docs.fireworks.ai/fine-tuning/training-api/cookbook/rl.md).
> This reference is retained only for jobs that already exist — monitoring, resuming,
> and recovering them.

Use this reference for managed RFT preflight, launch, job states, monitoring, and recovery. Training API or cookbook RL belongs in `references/training-api.md` and `references/rl-async.md`.
*Source of truth: live [RFT overview](https://docs.fireworks.ai/fine-tuning/reinforcement-fine-tuning-models.md), [Models matrix](https://docs.fireworks.ai/fine-tuning/models.md), [RL parameters](https://docs.fireworks.ai/fine-tuning/rft-parameters-reference.md), and [Eval Protocol](https://evalprotocol.io/introduction). Defer flags and defaults to installed CLI `--help`.*

Use this reference for managed RFT job states, monitoring, and recovery on existing jobs. Training API or cookbook RL belongs in `references/training-api.md` and `references/rl-async.md`.

## Preflight

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ Generate preference pairs and evaluators transparently in the user's workspace:
| Ideal labeled answers | SFT. Do not manufacture preference pairs. |
| Human or model-ranked pairs | DPO or ORPO. Normalize to the managed preference schema. |
| Prompts only, plus a clear preference criterion | Generate pairs, review a sample, then run DPO or ORPO. |
| Prompts plus objective correctness | Managed RFT with a registered evaluator, or Training API RFT with an inline reward. |
| Prompts plus objective correctness | Training API RL with an inline reward. |
| Open-ended quality criteria | Write and calibrate an LLM-judge rubric before training. |

Never silently turn prompts into preference data. Pair generation adds inference cost and embeds the generator or judge's bias into the training set.
Expand Down Expand Up @@ -87,9 +87,9 @@ Dependencies and network/credential requirements:

Show the spec to the user and resolve ambiguity before implementing the evaluator.

### Managed RFT evaluator
### Managed RFT evaluator (deprecated)

Managed RFT uses a registered evaluator with a reviewed entry point. Use Eval Protocol's current code-first flow, read `managed-rft-operations.md`, and defer exact APIs to the live [RFT overview](https://docs.fireworks.ai/fine-tuning/reinforcement-fine-tuning-models.md).
Managed RFT is deprecated and accepts no new jobs; for new work write a reward inside a Training API rollout function instead (see `rl-async.md`). The rest of this section applies only to evaluators already registered against an existing job. Managed RFT uses a registered evaluator with a reviewed entry point. Use Eval Protocol's current code-first flow, read `managed-rft-operations.md`, and defer exact APIs to the live [RFT overview](https://docs.fireworks.ai/fine-tuning/reinforcement-fine-tuning-models.md).

1. Write the Eval Protocol reward in the workspace.
2. Add deterministic unit examples for full credit, partial credit, zero, malformed output, and edge cases.
Expand Down
8 changes: 6 additions & 2 deletions skills/fireworks-training/references/rft-agent-tracing.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,12 @@
# Managed RFT remote tracing
# Managed RFT remote tracing (deprecated)

> **Managed RFT is deprecated.** For new multi-turn agent work use `references/rl-agentic.md`
> and [Cookbook: Agentic Reinforcement Learning](https://docs.fireworks.ai/fine-tuning/training-api/cookbook/agentic-rl.md).
> This reference covers remote environments attached to jobs that already exist.

*Source of truth: live [Remote Environment Setup](https://docs.fireworks.ai/fine-tuning/connect-environments.md) and [Eval Protocol](https://evalprotocol.io/introduction).*

Use this reference when implementing a managed RFT remote environment, wiring Fireworks tracing, or debugging a reward-to-rollout join. For custom Training API agent trajectories, use `references/rl-agentic.md`.
Use this reference when debugging a reward-to-rollout join on an existing managed RFT remote environment. For custom Training API agent trajectories, use `references/rl-agentic.md`.

## Why tracing matters

Expand Down
2 changes: 1 addition & 1 deletion skills/fireworks-training/references/sdk-shapes.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ trainer/deployment provisioning path.

## Deployment shape

**Do not create deployments without a [shape](https://docs.fireworks.ai/faq-new/deployment-infrastructure/what-is-a-deployment-shape).** Shapeless deployments are the most common cause of failed deployment creations, and the shapeless path may be deprecated in the future.
**Do not create deployments without a [shape](https://docs.fireworks.ai/faq-new/deployment-infrastructure/what-is-a-deployment-shape).** Shapeless deployments are the most common cause of failed deployment creations, and the [shapeless path](https://docs.fireworks.ai/guides/ondemand-deployments#explicitly-creating-a-deployment-without-a-shape-advanced-users-only) will be deprecated. Find a deployable shape, or pass `default` to let the server pick one.

Do not set `cfg.deployment.deployment_shape` manually. The SDK resolves it from
the requested [shape](https://docs.fireworks.ai/faq-new/deployment-infrastructure/what-is-a-deployment-shape) or the selected training profile, and recipes read
Expand Down
20 changes: 10 additions & 10 deletions skills/fireworks-training/references/training-api.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ Before selecting a Training API path, confirm that the target account was enable

## Managed training vs Training API

Use **managed training** for standard SFT/DPO/ORPO/RFT jobs. Reach for the **Training API** when you need a custom **loss/reward**, **RL with rollouts** (inference-in-the-loop), forward-pass internals (for example MoE routing for R3), distillation, or multi-turn/agentic trajectories.
Use **managed training** for standard SFT/DPO/ORPO jobs. Reach for the **Training API** when you need a custom **loss/reward**, **RL with rollouts** (inference-in-the-loop), forward-pass internals (for example MoE routing for R3), distillation, or multi-turn/agentic trajectories.

## Training API infrastructure

Expand All @@ -25,18 +25,18 @@ Use **managed training** for standard SFT/DPO/ORPO/RFT jobs. Reach for the **Tra

Read the live [serverless](https://docs.fireworks.ai/fine-tuning/training-api/serverless.md) and [dedicated](https://docs.fireworks.ai/fine-tuning/training-api/dedicated.md) pages before choosing.

## Two agent-drivable ways to run RFT/RL
## How to run RL

There are two RFT paths, and they differ in **where the reward lives**. This matters a lot when a coding agent is driving:
RL runs on the Training API. Fork `training.recipes.rl_loop` / `async_rl_loop` and supply an
inline `reward_fn(completion, row) -> float`. It may read `ground_truth`, another declared
reference field, tool outcomes, environment state, or a judge result. There is no evaluator
resource — same shape as Tinker's reward-in-the-loop — and the SDK provisions the trainer plus
rollout deployment. This is fully agent-drivable.

| Path | Reward | Agent-drivable? |
|---|---|---|
| **Managed RFT** — `firectl reinforcement-fine-tuning-job create --evaluator <id>` | A **registered evaluator resource** (server-side, built in an e2b sandbox, eval v3) | **Yes once the evaluator exists.** Register via **eval-protocol** (`pytest` auto-registers) or the **UI**. Evaluator authoring may require an admin role; a scoped key can still launch with an evaluator it can access. `firectl evaluator create` (V1) is **deprecated**. |
| **Training-API RL** — fork `training.recipes.rl_loop` / `async_rl_loop` | An **inline `reward_fn(completion, row) -> float`** in the forked recipe. It may read `ground_truth`, another declared reference field, tool outcomes, environment state, or a judge result. | **Yes.** No evaluator resource. Same shape as Tinker's reward-in-the-loop. The SDK provisions the trainer + rollout deployment. |

**Prefer the managed path for standard RFT** (same as the managed UI): `firectl reinforcement-fine-tuning-job create --dataset <ds> --evaluator <id>` — it resolves the training shape for you and is proven live (qwen3-4b, 2026-07-15). Reuse an existing evaluator or author one via eval-protocol. **Use the inline-reward recipe (below) for users with Training API access** who need a custom loop/reward, rollouts, or agentic trajectories. Both paths are agent-drivable; they differ in reward location, access, billing, and capability.
**Managed RFT (`firectl reinforcement-fine-tuning-job create --evaluator <id>`) is deprecated**
and accepts no new jobs. The evaluator material below applies only to jobs that already exist.

### Managed RFT: authoring the eval3 evaluator
### Managed RFT: authoring the eval3 evaluator (deprecated)

`firectl reinforcement-fine-tuning-job create` needs an **eval3 evaluator with an `entry_point`** — legacy evaluators are rejected (`InvalidArgument: managed RFT requires an eval3 evaluator`), and `firectl evaluator create` is deprecated. The code-first way to make one is **eval-protocol** (no UI). Check current evaluator authorization in the live docs and handle the observed role gate below:

Expand Down
4 changes: 2 additions & 2 deletions skills/research/references/case-studies.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ Runnable end-to-end notebooks in `training/case-studies/`. Each README has an
| `sft_prompt_router` | SFT / classification | End-to-end fine-tuning on a gradeable classification task | `prompt_router_dedicated.ipynb`, `prompt_router_serverless.ipynb` | managed SDK, serverless |
| `sft_cord_receipts` | Vision SFT | One right output shape (JSON, tags, codes) from examples; invoice/OCR/form extraction | `cord_receipt_sft_sdk.ipynb` | managed SDK |
| `dpo_style` | DPO | Accurate but wrong tone; easier to rank two answers than write the ideal one | `dpo_helpsteer3_sdk.ipynb` | managed SDK |
| `reasoning_rl` | GRPO / managed RFT | Objectively checkable answers; grader exists but no gold worked solutions | `rft_grpo_math.ipynb` | managed RFT |
| `reasoning_rl` | GRPO | Objectively checkable answers; grader exists but no gold worked solutions | `rft_grpo_math.ipynb` (legacy managed RFT) | New runs: Training API `training/examples/rl/deepmath/` |
| `embedding_support_search` | Contrastive embedding | RAG returns adjacent but wrong article; policy structure not in base model | `airbnb_policy_embedding.ipynb` | Training API `embedding_loop` |
| `agentic_rl_text2sql` | GRPO / serverless RL | Tool-calling agent (SQL, APIs); multi-turn rollouts with verifiable rewards | `sql_agent_rl_loop.ipynb` | serverless Training API |
| `multilora_fleet` | LoRA SFT / multi-LoRA serving | Many tenants or locales sharing one base model; per-tenant adapters served from a single deployment | `multilora_fleet.ipynb` | managed SDK |
Expand All @@ -22,7 +22,7 @@ Cookbook table: [`training/README.md`](https://github.com/fw-ai/cookbook/blob/ma
| `sft_prompt_router` | SFT | managed SDK |
| `sft_cord_receipts` | SFT | managed SDK |
| `dpo_style` | DPO | managed SDK |
| `reasoning_rl` | RFT (GRPO) | managed RFT |
| `reasoning_rl` | RL (GRPO) | Training API `training/examples/rl/deepmath/`; `rft_grpo_math.ipynb` is legacy managed RFT |
| `embedding_support_search` | embedding fine-tune | Training API dedicated |
| `agentic_rl_text2sql` | RL (GRPO) | serverless Training API |
| `multilora_fleet` | SFT (LoRA) | managed SDK |
Expand Down
2 changes: 1 addition & 1 deletion training/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -150,7 +150,7 @@ When `training_shape_id` is not set, the SDK selects validated runtime
defaults. Explicit `training_shape_id`, `reference_training_shape_id`, and
deployment-shape overrides still take precedence.

**Do not create deployments without a [shape](https://docs.fireworks.ai/faq-new/deployment-infrastructure/what-is-a-deployment-shape).** Shapeless deployments are the most common cause of failed deployment creations, and the shapeless path may be deprecated in the future. The cookbook resolves the deployment shape from the training shape profile automatically; find deployable shapes for a model with `firectl deployment-shape-version match --model <MODEL>`. Fields such as replica count override a shape when set; the shape owns accelerator selection.
**Do not create deployments without a [shape](https://docs.fireworks.ai/faq-new/deployment-infrastructure/what-is-a-deployment-shape).** Shapeless deployments are the most common cause of failed deployment creations, and the shapeless path will be deprecated. Find deployable shapes for a model with `firectl deployment-shape-version match --model <MODEL>` (or pass `default` to let the server pick one). The cookbook resolves the deployment shape from the training shape profile automatically. Fields such as replica count override a shape when set; the shape owns accelerator selection.

To launch trainers with replicated HSDP, set the run-level replica count on
`TrainerConfig`; it is not part of the validated training shape:
Expand Down
Loading
Loading