as_eval                 Build an Evaluation Object
ev_bootstrap            Cluster Bootstrap for an Arbitrary Statistic
ev_cluster              Cluster-Robust Standard Error for an Evaluation
ev_elo                  Bradley-Terry Ratings from Pairwise Preferences
ev_icc                  Intra-Cluster Correlation
ev_judge_agreement      Agreement Between a Model Judge and a Human
                        Gold Standard
ev_judge_debias         Debias a Model Judge with a Small Human Sample
ev_judge_power          How Many Human Labels a Debiased Evaluation
                        Needs
ev_mde                  Smallest Difference an Evaluation Can Detect
ev_multi                Multiplicity Adjustment Across a Benchmark
                        Suite
ev_paired               Paired Comparison of Two Models
ev_plot                 Plot Results with Error Bars
ev_power                Number of Questions Needed to Detect a
                        Difference
ev_rank                 Bootstrap Rank Intervals for a Leaderboard
ev_resample             Variance Decomposition for Repeated Sampling
ev_score                Evaluation Score with a Standard Error
ev_table                Collect Results into a Table
ev_unpaired             Unpaired Comparison of Two Models
ev_variance_reduction   Variance Reduction with a Reference Model
