The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.

Quickstart: fine-tune a small model from R

This walks through one complete fine-tune: a support-ticket dataset, a 135M parameter model that trains on a CPU in minutes, and a before-and-after comparison. Swap in your own file and a bigger model afterwards.

1. Check the machine

library(dragonfarm)
dragon_check()

The first call builds a Python environment with torch and transformers. That is a 2 to 3 GB download and takes a few minutes. Later calls take a second. The report tells you which device training will use. A CPU is fine for the 135M and 360M models. Anything larger wants a GPU.

2. Load and map the data

A dataset is a table with one row per example. The bundled example has subject, body, product, and reply columns.

ds <- dragon_dataset(dragon_example_data())
ds

dragon_map() says which columns form the user turn and the assistant turn. Each argument is a column name or a template that combines columns.

ds <- dragon_map(ds,
  prompt = "{subject}\n\n{body}",
  response = "reply",
  system = "You are a support agent for a smart-home company. Be concrete and brief."
)
dragon_preview(ds, n = 1)

The system argument here is a constant. It could also be a column.

3. Train

run <- dragon_train(
  ds,
  model = "HuggingFaceTB/SmolLM2-135M-Instruct",
  lora = dragon_lora(r = 8, alpha = 16),
  args = dragon_train_args(epochs = 2, batch_size = 4, grad_accum = 2, max_seq_len = 512),
  wait = TRUE
)

With wait = TRUE you get a progress bar and the function returns when training ends. Without it, the function returns at once and you poll:

run <- dragon_train(ds, "HuggingFaceTB/SmolLM2-135M-Instruct")
dragon_status(run)$state
tail(dragon_progress(run))
dragon_logs(run, 10)
dragon_wait(run)

Runs are directories under dragonfarm_runs/. They survive the R session:

dragon_runs()
run <- dragon_run(dragon_runs()$dir[1])

4. Evaluate and try it

Training holds out 5 percent of rows and reports loss and perplexity on them, plus a few generated replies next to the reference replies.

ev <- dragon_evaluate(run)
ev
ev$samples

Compare the tuned model with the base model on a fresh prompt:

prompt <- "Charged twice for Sentry doorbell\n\nMy card shows two charges for one order."
dragon_generate(run, prompt, temperature = 0)
dragon_generate(run, prompt, temperature = 0, base = TRUE)

5. Ship it

The adapter alone is small and loads with peft. For a standalone model that needs neither peft nor dragonfarm, merge:

merged <- dragon_merge(run, "models/support-135m")
dragon_generate(merged, prompt)

To reproduce the run later, or share it, ask for the code:

cat(dragon_code(run))

Choosing settings

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.