The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.
This walks through one complete fine-tune: a support-ticket dataset, a 135M parameter model that trains on a CPU in minutes, and a before-and-after comparison. Swap in your own file and a bigger model afterwards.
The first call builds a Python environment with torch and transformers. That is a 2 to 3 GB download and takes a few minutes. Later calls take a second. The report tells you which device training will use. A CPU is fine for the 135M and 360M models. Anything larger wants a GPU.
A dataset is a table with one row per example. The bundled example
has subject, body, product, and
reply columns.
dragon_map() says which columns form the user turn and
the assistant turn. Each argument is a column name or a template that
combines columns.
ds <- dragon_map(ds,
prompt = "{subject}\n\n{body}",
response = "reply",
system = "You are a support agent for a smart-home company. Be concrete and brief."
)
dragon_preview(ds, n = 1)The system argument here is a constant. It could also be a column.
run <- dragon_train(
ds,
model = "HuggingFaceTB/SmolLM2-135M-Instruct",
lora = dragon_lora(r = 8, alpha = 16),
args = dragon_train_args(epochs = 2, batch_size = 4, grad_accum = 2, max_seq_len = 512),
wait = TRUE
)With wait = TRUE you get a progress bar and the function
returns when training ends. Without it, the function returns at once and
you poll:
run <- dragon_train(ds, "HuggingFaceTB/SmolLM2-135M-Instruct")
dragon_status(run)$state
tail(dragon_progress(run))
dragon_logs(run, 10)
dragon_wait(run)Runs are directories under dragonfarm_runs/. They
survive the R session:
Training holds out 5 percent of rows and reports loss and perplexity on them, plus a few generated replies next to the reference replies.
Compare the tuned model with the base model on a fresh prompt:
The adapter alone is small and loads with peft. For a
standalone model that needs neither peft nor dragonfarm, merge:
To reproduce the run later, or share it, ask for the code:
HuggingFaceTB/SmolLM2-360M-Instruct or
Qwen/Qwen2.5-0.5B-Instruct. Move up only if quality is not
enough.2e-4 for LoRA. Halve it
if the loss curve is jagged.batch_size, raise
grad_accum to compensate, and turn on
gradient_checkpointing.These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.