The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.
Description field (per CRAN reviewer feedback: software
names must be single-quoted in
Title/Description).Description: Friedman
(2001) doi:10.1214/aos/1013203451, the canonical gradient
boosting machine reference (per CRAN reviewer request).objective = "multiclass")fastgbm() now covers a fourth task type: multiclass
classification, implemented as one binary (one-vs-rest)
fastgbm sub-model per class, all sharing the same
hyperparameters and reusing the already-tested, already-correct compiled
logistic objective rather than adding a new compiled kernel.
predict(type = "prob") renormalizes the K independent
binary probabilities to sum to 1 per row;
predict(type = "class") is their argmax.objective is inferred automatically: a factor/character
response with 3+ levels defaults to "multiclass" (2-level
factors still default to "binary", unchanged).metrics() reports accuracy and multiclass log loss;
importance() combines split gain across the K
sub-models.y ~ .)
interfaces, and supports
validation/early_stopping (applied per
sub-model).inst/benchmarks/run-benchmark-scaling.R’s
P_GRID now includes 1000 (was
{10, 50, 100}). Confirms and sharpens the earlier
colsample finding: at n = 1000, the
fastgbm/ranger training-time ratio reaches 12.04x (slower) at
p = 1000, vs. 4.45x at p = 100. Unlike lower
p, growing n does not rescue
fastgbm at p = 1000 within the tested range –
it is still ~2.8x slower than ranger even at
n = 20000 (colsample tuning matters most
exactly in this regime). The thread-count sweep, now run at the new
largest grid point (n = 20000, p = 1000), shows
better multi-threaded speedup than at p = 100
(5.8x vs. 3.0x at 16 threads) since a wider
colsample-sampled feature set gives the level-wise split
search more to parallelize per level.inst/benchmarks/run-benchmark-multithreaded.R: a
direct head-to-head of
fastgbm/ranger/xgboost at matched
thread counts ({1, 4, 12}, gbm shown only as a
fixed single-threaded reference), at the largest n/p grid point from the
scaling benchmark. Finding: the single-threaded win over
ranger reported below did not survive multi-threading –
ranger (independent trees split across threads) and
xgboost (histogram build parallelized across rows and
features) both scaled substantially better than fastgbm’s
prior per-node split-search parallelism, to the point
fastgbm became the slowest of the three at high thread
counts despite winning at threads = 1.parallelFor()
dispatch (one call per node, parallelizing across that node’s own
sampled features) ran out of useful work several levels into a tree – a
single node’s row/feature count shrinks fast with depth, but the fixed
cost of a parallelFor() call does not, so deep, numerous,
small nodes got little benefit from more threads.build_tree() now grows trees breadth-first, one
level at a time. Every node active at the current depth has its
split-search tasks pooled into a single combined batch
(LevelSplitFinder, replacing the previous per-node
SplitFinder) before any of them run, so the parallelized
unit of work is an entire tree level, not one node – keeping
per-dispatch work large even deep in the tree, since each level roughly
doubles its node count while roughly halving each node’s row count (the
total per-level work stays close to constant).colsample subset), verified rather than assumed safe: the
existing leaf-value correctness test and a new
multi-threaded-vs-single-threaded identical-tree check both pass
(tests/testthat/test-tree-building.R), plus the full
pre-existing suite (244 tests, unchanged pass count).n = 20000, p = 100:
fastgbm’s own multi-threaded speedup improved from
1.9x/3.3x (4/12 threads) to 2.8x/5.2x. This
narrows, but does not close, the gap to ranger’s
4.6x/9.7x and xgboost’s 3.2x/6.0x –
see README’s “Multi-threaded head-to-head” and “Level-wise tree growth”
sections.parallelFor()’s
grainSize from 1 to roughly one chunk per
thread, on the theory that fewer/larger scheduled tasks would cut
scheduling overhead – was tried, measured, and found to make things
worse (e.g. threads=4 at n=20000, p=100: ~1.36s
-> ~1.56s), then reverted. RcppParallel’s own work-stealing scheduler
at grainSize = 1 already balanced the per-node case better
than a fixed static chunking did.build_tree()’s grow()
recomputed its total gradient/Hessian sum
(G/H) with a fresh full pass over its own
rows, even though those exact values were already produced as a
byproduct of the parent’s winning split
(GL/HL/GR/HR,
captured in NodeSplit’s new fields during the
SplitFinder histogram scan). G/H
are now passed down to each child instead, removing a second full
O(rows) pass per node on top of the split search itself; only the root
node (no parent split to inherit from) still computes them directly,
once.tests/testthat/test-tree-building.R) independently
recomputes each leaf’s -G/(H+lambda) in R from the actual
rows reaching it (subsample/ colsample fixed
at 1 so that set is unambiguous) and checks it against the
compiled value to 1e-10, plus a
multi-threaded-vs-single-threaded identical-tree check; both pass, along
with the full pre-existing suite (225 tests, unchanged).fastgbm training time dropped further, from ~147ms to
~143-146ms, turning the prior statistical tie with ranger
into a narrow but real win (median training time 0.099s
vs. ranger’s 0.104s over 10 repeated splits, win/loss/tie
6/3/1; see
inst/benchmarks/run-benchmark-regression.R).inst/benchmarks/run-benchmark-scaling.R: sweeps
n in {1000, 5000, 20000} and p in
{10, 50, 100} (Friedman1 plus pure-noise columns) comparing
fastgbm vs. ranger training time,
single-threaded, plus a thread-count sweep (1/2/4/all-cores) at the
largest grid point.fastgbm’s relative speed
vs. ranger improves as n grows (1.30x slower
at n=1000, p=10 to 0.57x, i.e. faster, at
n=20000, p=10) but degrades as p grows (1.30x
-> 4.41x slower at n=1000 as p goes 10
-> 100). Root cause identified, not just observed:
fastgbm’s default colsample = 0.8 searches far
more features per split than ranger’s default
mtry (p/3 for regression) once p
is large; matching it (colsample = 1/3) turns a loss into a
win at n=5000, p=100 (0.98s -> 0.50s,
vs. ranger’s 0.57s). Multi-threading gives real but
sub-linear speedup (1.8x at 4 threads, 2.3x at 16, on a 16-core
machine), as expected for within-tree (not across-tree)
parallelism.src/fastgbm.cppSplitFinder made two full passes over the node’s rows - one
only to find the highest occupied bin, a second to accumulate
gradient/hessian sums into an array sized to that bin. Since each
feature’s bin range is already known globally from the initial binning
pass (max_bin_by_feature, computed once in
fastgbm_fit_cpp), the histogram buffer is now sized up
front and a single pass both accumulates the sums and finds the tightest
split-loop bound (the highest bin actually present among these specific
rows).thread_local buffers across every (node, feature)
call on a given thread instead - safe because parallelFor()
gives each worker thread disjoint index ranges, never concurrent access
to the same thread’s buffer.RcppParallel::parallelFor() was dispatched for
large-enough nodes even when threads = 1L was explicitly
requested, paying real thread-pool scheduling overhead for work that
only ever ran on one thread. A single explicit thread request now always
takes the direct serial call, regardless of node/feature size.n = 2000, p = 10, 200 trees, single-threaded,
bench::mark() median): fastgbm training time
dropped from ~159ms to ~147ms, closing what was previously a clear gap
to ranger (~143-145ms) to a statistical tie (win/loss/tie
5/5 across 10 repeated splits in
inst/benchmarks/run-benchmark-regression.R, and
fastgbm’s own median training time is marginally faster
than ranger’s: 0.1005s vs. 0.1015s). See README’s
“Split-search optimizations” section and
inst/benchmarks/regression-exectime-bench.png/.csv.test-parallel.R, test-gradients.R).objective = "regression" / "binary")fastgbm() now covers three task types, not just
survival: objective = "regression" (squared error) and
objective = "binary" (logistic classification, 0/1 response
or two-level factor) join the existing
"cox"/"aft"/"pexp" survival
objectives. All three task types share the same tree-growing engine,
histogram binning, missing-value routing, and validation-based
early-stopping machinery – the compiled backend
(src/fastgbm.cpp) only needed two new gradient/Hessian
pairs (compute_gaussian_grad_hess,
compute_logistic_grad_hess) and matching loss functions for
early stopping, verified against finite differences in
tests/testthat/test-gradients.R.objective is now optional: it defaults to
"cox" when time/status or a
survival::Surv response is supplied, otherwise it is
inferred from y (a 0/1 vector or two-level factor defaults
to "binary", anything else numeric to
"regression").y ~ .) now supports
non-Surv responses for regression/binary, not just
Surv(time, status) ~ ..predict(), metrics() (RMSE for regression,
log loss for binary), importance(), and pdp()
all work unchanged across the new objectives.time C++ argument slot for the
plain response, exactly like pexp already reuses it for
exposure), and the full pre-existing survival test suite still
passes.colsample) from
once-per-tree (xgboost’s colsample_bytree semantics) to
once-per-node (ranger’s mtry semantics,
resampled at every split): src/fastgbm.cpp’s
build_tree() now takes
p/colsample/rng instead of a
precomputed feature_ids list, and resamples inside
grow() at each node. Verified to preserve thread-count
determinism (all parallel-determinism tests still pass). Improves both
training speed (smaller effective per-node search) and C-index versus
per-tree sampling on the benchmark datasets."colnames" attribute on the input matrix (base
R matrices store column names under dimnames(x)[[2]], not a
top-level "colnames" attribute), so
importance() always showed generic V1, V2, ...
labels regardless of the matrix’s actual colnames().fastgbm()’s default hyperparameters to match
what all this session’s diagnostics actually validated – they had
drifted out of sync with the benchmark’s own settings:
learning_rate 0.05 to 0.1, max_depth 6 to 5,
min_node_size 20 to 10, ntrees 500 to 200.
Documented explicitly in ?fastgbm that training all
ntrees rounds without
validation/early_stopping reliably overfits on
small-to-medium survival data, per the diagnostics.inst/benchmarks/run-benchmark.R now benchmarks all
three implemented objectives (fastgbm_cox,
fastgbm_aft, fastgbm_pexp) as separate arms
against gbm/xgboost/ranger, not
just Cox.mtry-scaled-to-sqrt(p)/p default helped on
higher-p datasets but hurt on low-p ones; a
bagged-ensemble-of-fastgbm-fits wrapper (bootstrap
resampling + averaging, directly borrowing ranger’s core
variance-reduction mechanism) closed the C-index gap to
ranger on some datasets (e.g. beat it on pbc)
but not others. See Roadmap in
paper/fastgbm-benchmark.qmd.objective = "pexp")pexp models the log hazard rate
jointly over covariates and time: training data is expanded
into person-time rows via the standard “Poisson trick”
(fastgbm_expand_person_time(), R/pexp.R), with
log(interval midpoint) appended as an extra feature, and
the ensemble is fit with an ordinary Poisson-with-offset
gradient/Hessian
(compute_pexp_grad_hess()/compute_pexp_loss(),
src/fastgbm.cpp) – no derivative-of-the-ensemble
approximation is needed, unlike a Royston-Parmar- style spline baseline,
because the hazard is directly exp(pred) rather than a
spline whose local slope must be estimated.pexp_bins = 10L quantiles of the
observed event times (fastgbm_pexp_cutpoints()).predict(type = "survival") evaluates the
hazard-over-time surface directly at the requested times
(piecewise-linear cumulative hazard, extrapolated past the last cutpoint
by holding the final interval’s hazard rate constant).
predict(type = "link"/"response") uses the cumulative
hazard at the model’s full fitted time horizon as a fixed risk score (no
single “linear predictor” exists the way Cox/AFT have one, since the
model is a function of time).tests/testthat/test-gradients.R), a
Poisson-GLM-on-expanded-data equivalence check against the coefficient
coxph recovers on the same synthetic proportional-hazards
data (tests/testthat/test-pexp.R), and end-to-end tests for
missing-value handling, early stopping, thread-count determinism,
serialization, and the formula interface (same coverage bar as
Cox/AFT).pexp,
applies to all objectives): the compiled backend read feature names from
a nonexistent "colnames" attribute on the input matrix (R
matrices store column names under dimnames(x)[[2]], not a
top-level "colnames" attribute), so
importance() always showed generic V1, V2, ...
labels regardless of the matrix’s actual colnames(). Fixed
in src/fastgbm.cpp.paper/fastgbm-benchmark.qmd);
pexp was implemented first because it required no such
approximation.inst/benchmarks/error-analysis/diagnose.R:
learning-curve, discordant-pair, and regularization diagnostics run
against ranger on all 6 benchmark datasets. Found that
test-set C-index consistently peaked between 10 and 100 boosting rounds
and then degraded with further training on every dataset
(e.g. pbc: train C-index 0.999 at 300 trees, test C-index
peaked at 0.768 at 50 trees, down to 0.753 at the previous fixed default
of 200) – classic overfitting from having no way to stop early, and a
secondary finding that ranking errors concentrate disproportionately in
a few high-leverage subjects.validation/early_stopping parameters
previously accepted but ignored, per spec section 14): added
compute_cox_loss()/compute_aft_loss() to
src/fastgbm.cpp, extended fastgbm_fit_cpp() to
bin a validation set with the training cuts, track validation loss per
round, and stop after early_stopping rounds without
improvement. fit$trees is truncated to the
best-validation-loss iteration by default; the full run is preserved in
fit$n_trees_grown, fit$validation_history,
fit$best_iteration, and fit$stopping_reason.
Verified deterministic across thread counts
(tests/testthat/test-early-stopping.R).colsample’s default from 1 to
0.8, based on the diagnostic regularization probe showing
small, consistently non-negative C-index gains.colsample default: C-index improved on every dataset
(e.g. colon_cancer 0.605 to 0.625), and training time
dropped 4-10x since far fewer trees are actually grown –
fastgbm is now faster to train than xgboost on
all 6 datasets (previously 4 of 6). ranger’s discrimination
lead narrowed but did not close.colsample; changed
subsample’s default from 1 to 0.8
as well. Rebenchmarked again: fastgbm now beats both
gbm and xgboost on C-index on 4 of 6 datasets
(heart_failure 0.681/0.682 to 0.703, breast
0.648/0.648 to 0.677, crc_mondaca2020 0.558/0.553 to 0.568,
and effectively ties gbm on pbc); the
ranger gap narrowed further (breast: 5.9
points of C-index to 2.0) but did not close on any dataset.fastgbm → fastgbm. All
identifiers, exported functions, and the compiled backend
(src/fastgbm.cpp) were renamed to match.reg:squarederror and
binary:logistic objectives entirely: fastgbm
is now survival-only. Objective strings simplified from
"survival:cox"/ "survival:aft" to
"cox"/"aft" (default "cox"). The
plain-y-vector matrix/formula interface is gone; the
response is always time/status or a
survival::Surv object.rmse(), logloss(),
accuracy(), auroc(), and the pure-R
squared-error/logistic reference implementation
(tests/reference/) - no longer applicable now that the
package only supports survival objectives.RcppParallel as a compiled dependency and
parallelized the per-feature split search (SplitFinder
worker in src/fastgbm.cpp): each feature’s best split is
computed independently and reduced via a fixed-order argmax, so training
is bit-identical regardless of thread count. The existing
threads parameter (0 = automatic) is now wired through to
RcppParallel::parallelFor(); below an internal
node/feature-count threshold, the same worker code runs serially instead
(dispatch overhead isn’t worth it for small nodes). See
tests/testthat/test-parallel.R.pbc,
heart_failure, breast,
colon_cancer, crc_mondaca2020,
framingham) and the non-survival benchmark arms
(pima_diabetes, diabetes_prediction,
maternal_bangladesh) were removed from
inst/benchmarks/run-benchmark.R.man/*.Rd files are now fully
roxygen2-generated (previously hand-written) - run
devtools::document() after any R-level API change, never
edit NAMESPACE by hand.grad = pdf/(sigma*sf) instead
of -pdf/(sigma*sf)), which pushed boosting updates in the
wrong direction for every censored AFT observation. Found via
finite-difference tests against the negative log-likelihood.fastgbm_survival_baseline()
(the Breslow baseline cumulative hazard used by
predict(..., type = "survival") for Cox models) computed a
risk set that only grew and never shrank over time, so the risk-set
denominator was effectively constant at the total sample size for every
event time instead of the correct shrinking at-risk set. This distorted
every Cox survival-probability prediction. Found by comparing against
survival::basehaz() on synthetic data; now matches
exactly.metrics() computed the
survival C-index using a risk-direction score for AFT models that was
actually a predicted-time score (higher = longer survival, the opposite
of “higher = more risk”), silently producing near-inverted concordance
for AFT models. Fixed by negating the AFT linear predictor before
scoring.importance() crashed (or
silently returned garbage data-frame shapes) because the compiled
backend never attached feature names to feature_importance;
it now falls back to the model’s feature_names.fastgbm() did not validate
that nrow(x) matched length(y) (or
length(time)/length(status)), leading to
out-of-bounds memory reads and silently garbage predictions on
mismatched input; this now raises a clear error.tests/testthat/test-gradients.R
(finite-difference checks for all four objectives, via a new internal
fastgbm_grad_hess() test hook onto the compiled kernels),
test-survival-validity.R (Breslow baseline vs
survival::basehaz, Cox ranking vs coxph),
test-missing.R, and test-safeguards.R.binary:logistic reference
implementation (reference_logistic_fit()) alongside the
existing squared-error one.These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.