The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.
Pre-submission audit fixes, ahead of the first CRAN release.
ev_bootstrap() and ev_rank() called
set.seed() on the global random stream and left it altered.
A user who passed seed to make one call reproducible found
every later random draw in their session silently shifted. The seed now
applies for the duration of the call only and the user’s
.Random.seed is restored on exit. Seeded calls remain
reproducible and unseeded calls still vary.
ev_score() and ev_cluster() now warn
when the interval has degenerated. A slice where every item passes gives
a zero standard error, so the Wald interval collapses to a point and
appears to claim certainty: 50 of 50 correct reported
[1.0000, 1.0000] where Wilson gives about
[0.93, 1.00]. Too few clusters pushes a pass rate’s
interval outside [0, 1]. Both are properties of the normal
approximation rather than errors, so they warn rather than abort, and
the binary case quotes the Wilson interval for comparison.
README gains a Limitations section covering the interval method and where it degrades, the cluster-count requirement, the random-sampling assumption behind prediction-powered inference, and the difference between judge agreement and judge accuracy.
First release.
Statistical inference for language model evaluations, following Miller (2024) doi:10.48550/arXiv.2411.00640 for the core standard error and experiment design results, and the prediction-powered inference literature for the model judge functions.
as_eval() builds the long-format evaluation object
every other function consumes, from a data frame or a bare vector of
scores.ev_score() reports a mean score with a central limit
theorem standard error and confidence interval, switching to
cluster-robust inference automatically when a cluster column is
present.ev_cluster() gives the cluster-robust standard error
with the design effect and intra-cluster correlation that explain it,
and ev_icc() returns the intra-cluster correlation on its
own.ev_resample() separates between-item variance from
response sampling noise when several responses are drawn per question,
and projects the standard error achievable at any number of responses
per item.ev_bootstrap() provides a cluster bootstrap for
statistics that are not means.ev_paired() compares two models on the same questions
and reports the variance saved by pairing.ev_unpaired() handles disjoint question sets with a
Welch interval.ev_variance_reduction() applies a control variate from
a reference model with a known score on the full question bank.ev_multi() adjusts across a benchmark suite, pools by
inverse variance, and tests heterogeneity.ev_power() and ev_mde() size a comparison
and report its minimum detectable effect, both accounting for
clustering.ev_judge_agreement() reports accuracy, Cohen’s kappa,
and a McNemar test for systematic judge bias.ev_judge_debias() implements the power-tuned
prediction-powered estimator, combining many judge scores with a small
human sample.ev_judge_power() sizes the human labelling budget.ev_rank() gives bootstrap rank intervals and the
probability each model is best.ev_elo() fits Bradley-Terry ratings with standard
errors on the Elo scale.ev_table() and ev_plot() collect results
into a table and draw them with error bars.These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.