The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.
dragon_reinforce() trains a model with Group Relative
Policy Optimization (GRPO), the method behind the recent reasoning
models, in its plain on-policy form. For every prompt the model writes
several answers, each is scored by rewards you define, and the model is
nudged towards the answers that beat their group’s average. A KL penalty
against the model it started from keeps it from drifting.
Reinforcement learning on a small model works when the reward is verifiable: something a function can check without opinion.
It works poorly as a substitute for preference data on goals such as
“be more helpful” or “sound friendlier”. A reward that is really a
model’s opinion is noisy, and a small policy will find the noise. For
those goals use dragon_prefer() with judge-ranked pairs
instead; see the post-training loop vignette.
Rows carry a prompt and, optionally, a reference answer for rewards
that compare against one. Every other column travels along as
fields for custom rewards.
library(dragonfarm)
a <- sample(10:999, 400, replace = TRUE)
b <- sample(10:999, 400, replace = TRUE)
math <- data.frame(
question = sprintf("What is %d + %d? Reply with just the number.", a, b),
answer = as.character(a + b)
)
prompts <- dragon_map_prompts(dragon_dataset(math), prompt = "question", reference = "answer")
dragon_preview(prompts, n = 1)dragon_reward() builds one reward; a list of them is
summed by weight.
rewards <- list(
dragon_reward("numeric"), # last number equals the reference
dragon_reward("length", max_chars = 12, weight = 0.3) # keep it terse
)The built-ins: "exact", "contains",
"numeric", "regex" (with
pattern), "json" (with optional
keys, partial credit per key), "length" (with
min_chars and max_chars), and
"keyword" (with words and mode).
For anything else, "custom" points at a Python file:
writeLines(c(
"def reward(prompt, completion, reference, row):",
" # 1.0 when the reply mentions the product named in the row",
" return 1.0 if row.get('product', '') and row['product'] in completion else 0.0"
), "product_reward.py")
mention <- dragon_reward("custom", file = "product_reward.py", name = "mentions_product")The file is copied into the run, so the run stays self-contained and travels with the cloud bundle.
Start from a fine-tuned run when you have one; the RL stage folds its adapters in before adding its own.
rl <- dragon_reinforce(
prompts, sft, # or a model id such as "Qwen/Qwen2.5-0.5B-Instruct"
rewards = rewards,
group_size = 6, # answers sampled per prompt
beta = 0.04, # KL penalty towards the starting model
temperature = 1.0, # sampling temperature; must be above 0
max_new_tokens = 16,
args = dragon_train_args(learning_rate = 1e-5, epochs = 1, batch_size = 4, grad_accum = 1, save_steps = 20),
wait = TRUE
)A step samples batch_size * group_size completions,
scores them, and takes one optimizer step. Sampling dominates the cost,
so short max_new_tokens and modest group_size
keep steps fast. Progress rows carry reward,
reward_std, kl, and
completion_len, plus one column per reward:
dragon_progress(rl)[, c("step", "reward", "reward_numeric", "reward_length", "kl", "completion_len")]The Monitor panel plots mean reward on a second axis next to the loss.
dragon_evaluate(rl) # mean held-out reward, per reward
dragon_generate(rl, "What is 417 + 285? Reply with just the number.", temperature = 0)
dragon_compare(sft, rl)reward in the comparison table is the mean total reward
on held-out prompts. Compare it with the same rewards computed on the
starting model to see what the stage bought:
dragon_evaluate() on the starting run does not know these
rewards, so the quickest check is dragon_generate() on both
with base = TRUE on the RL run and a few prompts.
beta trades speed of change for
stability. Raise it if replies degrade in ways the rewards do not see;
lower it if nothing moves.temperature needs to be high enough
that the group varies. 0.8 to 1.2 is typical.Reinforcement learning is a stage like the others. A typical order is fine-tune, then preference optimization for style, then RL for a verifiable target, and each starts from the previous run. In a pipeline:
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.