The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.
This guide collects the common estimation and post-estimation
patterns. It assumes that Python has been configured as described in
getting-started. Every model below is specified in R and
compiled once into native Biogeme expressions before numerical work
begins.
The chunks are intentionally not evaluated during package documentation builds: they are complete, runnable examples, but estimation and simulation require a configured Python environment and may create native output files.
library(rbiogeme)
data <- data.frame(
choice = c(1, 2, 3, 1, 2, 3, 1, 2),
time_1 = c(10, 8, 9, 12, 7, 10, 11, 8),
time_2 = c(8, 10, 7, 9, 11, 8, 10, 12),
time_3 = c(12, 9, 11, 8, 10, 12, 9, 11),
cost_1 = c(5, 7, 6, 5, 8, 6, 5, 7),
cost_2 = c(7, 5, 8, 6, 7, 5, 8, 6),
cost_3 = c(9, 8, 7, 10, 8, 9, 7, 8),
weight = c(1, 1, 2, 1, 1, 2, 1, 1),
person = c(1, 1, 2, 2, 3, 3, 4, 4)
)
database <- biogeme_database("workflows", data)
b_time <- biogeme_beta("b_time", start = 0)
b_cost <- biogeme_beta("b_cost", start = 0)
asc_2 <- biogeme_beta("asc_2", start = 0)
asc_3 <- biogeme_beta("asc_3", start = 0)
utilities <- list(
`1` = b_time * variable("time_1") + b_cost * variable("cost_1"),
`2` = asc_2 + b_time * variable("time_2") + b_cost * variable("cost_2"),
`3` = asc_3 + b_time * variable("time_3") + b_cost * variable("cost_3")
)
output_directory <- tempfile("rbiogeme-workflows-")
dir.create(output_directory)
clean_control <- biogeme_control(
output_directory = output_directory,
generate_html = FALSE,
generate_yaml = FALSE,
save_iterations = FALSE
)Availability is an expression for each alternative. A value of
1 means available and 0 means unavailable; a
symbolic indicator can vary by observation. If the utility names are
labels instead of integer codes, supply alternative_codes
explicitly.
logit <- logit_model(
database = database,
choice = "choice",
utilities = utilities,
availability = list(
`1` = 1,
`2` = variable("cost_2") < 9,
`3` = variable("time_3") < 12
),
weight = variable("weight")
)
fit <- estimate(logit, model_name = "workflow_logit", control = clean_control)
summary(fit)
# Native probability columns for the fitted alternatives.
predicted <- predict(fit)
head(predicted)The lower-level probability constructors are useful when the probability is part of a larger likelihood or simulation formula:
choice_probability <- logit_probability(
utilities = utilities,
availability = list(`1` = 1, `2` = 1, `3` = 1),
alternative = variable("choice")
)
choice_log_probability <- logit_log_probability(
utilities = utilities,
availability = list(`1` = 1, `2` = 1, `3` = 1),
alternative = variable("choice")
)
# The same expression can be used in a generic model.
generic_logit <- biogeme_model(
database = database,
formula = choice_log_probability,
probability = choice_probability,
weight = variable("weight")
)Nested-logit structures are declarative. A non-trivial nest contains its alternative codes and a nest parameter; alternatives not listed in such a nest are treated as trivial nests.
mu_motor <- biogeme_beta("mu_motor", start = 1, lower = 1)
nests <- nested_nests(
choice_set = c(1, 2, 3),
nests = list(
nested_nest(mu_motor, alternatives = c(2, 3), name = "motor")
)
)
nested <- nested_logit_model(
database = database,
choice = "choice",
utilities = utilities,
nests = nests
)
nested_fit <- estimate(nested, model_name = "workflow_nested", control = clean_control)
nested_logit_correlation(nested, beta_values = coef(nested_fit))In a cross-nested model, an alternative can receive a symbolic or
numeric allocation in more than one nest. With
sparse = TRUE, structurally omitted allocations are
allowed.
mu_a <- biogeme_beta("mu_a", start = 1, lower = 1)
mu_b <- biogeme_beta("mu_b", start = 1, lower = 1)
cnl_nests <- cross_nested_nests(
choice_set = c(1, 2, 3),
nests = list(
cross_nested_nest(mu_a, c(`1` = 1, `2` = 0.5, `3` = 0)),
cross_nested_nest(mu_b, c(`1` = 0, `2` = 0.5, `3` = 1))
)
)
cnl <- cross_nested_logit_model(
database = database,
choice = "choice",
utilities = utilities,
nests = cnl_nests
)
cross_nested_sparsity_report(cnl_nests)
cnl_fit <- estimate(cnl, model_name = "workflow_cnl", control = clean_control)
cross_nested_logit_correlation(cnl, beta_values = coef(cnl_fit))Declare the panel identifier once. Rows for one individual must be contiguous; the constructor checks this rather than silently changing the panel structure. For a random-coefficient panel model, place the trajectory and integration nodes in the generic expression and provide draw metadata.
panel_database <- biogeme_panel_database(
"workflow_panel",
data,
panel_id = "person"
)
panel_probability <- logit_probability(
utilities = utilities,
alternative = variable("choice")
)
trajectory <- panel_likelihood_trajectory(panel_probability)
b_time_random <- biogeme_beta("b_time_random", start = 0)
random_utility <- list(
`1` = (b_time_random + biogeme_beta("sd_time", start = 1, lower = 0) *
draw("time_draw", "NORMAL")) * variable("time_1"),
`2` = asc_2 + (b_time_random + biogeme_beta("sd_time", start = 1, lower = 0) *
draw("time_draw", "NORMAL")) * variable("time_2"),
`3` = asc_3 + (b_time_random + biogeme_beta("sd_time", start = 1, lower = 0) *
draw("time_draw", "NORMAL")) * variable("time_3")
)
random_probability <- logit_probability(
utilities = random_utility,
alternative = variable("choice")
)
panel_model <- biogeme_model(
database = panel_database,
formula = log(monte_carlo(panel_likelihood_trajectory(random_probability))),
draws = biogeme_draws(
name = "time_draw",
draw_type = "NORMAL_ANTI",
number_of_draws = 128L,
seed = 1223L
)
)
panel_fit <- estimate(panel_model, model_name = "workflow_panel", control = clean_control)panel_likelihood_trajectory() multiplies the
observation-level probabilities for each individual in native Biogeme.
monte_carlo() then averages over the declared draw design.
Neither operation is evaluated by R. For numerical quadrature over a
standard normal variable, use
integrate_normal(expression, name, number_of_quadrature_points)
instead.
Simulation is a separate operation from estimation. Give the model a
named simulations list, or pass one directly to
simulate().
simulation_model <- biogeme_model(
database = database,
formula = choice_log_probability,
simulations = list(
probability = choice_probability,
expected_time = variable("time_1") * choice_probability,
cost_share = variable("cost_1") / (1 + variable("cost_1"))
)
)
simulated <- simulate(simulation_model, beta = fit, control = clean_control)
as.data.frame(simulated)
# Expressions can also be supplied for one simulation call.
simulate(
simulation_model,
expressions = list(probability = choice_probability),
beta = fit,
control = clean_control
)The package also exposes native derivative checks, confidence intervals, and cross-validation:
check_derivatives(
model = generic_logit,
model_name = "workflow_derivatives",
control = clean_control,
verbose = TRUE
)
biogeme_confidence_intervals(
model = simulation_model,
beta_values = list(coef(fit), coef(fit)),
expressions = list(probability = choice_probability),
interval_size = 0.90,
control = clean_control
)
validate(
model = logit,
fit = fit,
folds = 5L,
groups = "person",
seed = 1234,
control = clean_control
)validate_model() checks the compiled native
specification before estimation. check_derivatives()
compares native analytical derivatives with native finite differences.
biogeme_confidence_intervals() delegates repeated parameter
draws and quantile calculations to native Biogeme.
validate() delegates fold assignment, estimation, and
scoring to native Biogeme.
There are two deliberately different result-file workflows:
yaml_file <- file.path(output_directory, "workflow.yaml")
# Always estimate afresh and optionally write the named YAML result.
fresh_fit <- estimate(
logit,
model_name = "workflow_file",
yaml_file_name = yaml_file,
control = clean_control
)
# Loading is explicit. force = FALSE permits reading the named file;
# force = TRUE ignores it and estimates again.
loaded_fit <- estimate_or_load(
logit,
yaml_file_name = yaml_file,
force = FALSE,
model_name = "workflow_file",
controls = clean_control
)
# The ordinary R representation can also be saved and read explicitly.
save_results(fresh_fit, yaml_file)
read_results(yaml_file)Use a fresh temporary directory in equivalence tests. This makes it clear which files belong to the current run and prevents a stale YAML or iteration file from changing the result.
biogeme_control() is a named R list. Named arguments
such as seed, number_of_draws,
optimization_algorithm,
variance_covariance_type, and
bootstrap_samples are mapped to their native controls.
Additional named arguments are passed through, which keeps the interface
usable when a native operation exposes a specialized control not yet
given a dedicated R argument.
control <- biogeme_control(
output_directory = output_directory,
seed = 1234,
numerically_safe = TRUE,
optimization_algorithm = "automatic",
number_of_draws = 128L,
generate_html = FALSE,
generate_yaml = FALSE,
save_iterations = FALSE
)Use the model-specific help pages for the complete argument contract of each operation. The next guide covers Bayesian, Monte Carlo, MDCEV, catalog, assisted-specification, hybrid-choice, and sampled-alternative workflows.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.