An open library for your next question. Public pilot
Executable Science
Log inCreate account
← Explore the library
CURATED COLLECTION · 191 RESOURCES

AI evaluation

Choose inputs, comparisons and metrics before trusting a score.

Choose a useful starting point.

Two scores are comparable only when the task, input split, normalization, aggregation and failure handling are stated. Begin by freezing those choices and a simple baseline. The evaluation-harness brief asks whether two tools agree on the same saved predictions; it does not claim either tool has been run here.

For agent evaluations, preserve the environment, action trace, budget and unsuccessful attempts alongside the scoring rule. The linked protocols help you plan this record. They do not make a benchmark suitable for every agent or establish that a qualified runtime supports its dependencies.

Questions to work through.

Explore the sources.

All 191 resources →
Datasets

Iris

Flower measurements for a compact, interpretable multiclass baseline and leakage audit.

License evidence recorded
Methods

1.13. Feature selection

A guide to variance filters, univariate tests, recursive elimination, model-based selection and selection inside pipelines.

Rights need review
Methods

1.16. Probability calibration

A probability-calibration guide covering reliability curves and sigmoid, isotonic and temperature-scaling approaches.

Rights need review
Methods

2.7. Novelty and Outlier Detection

A guide distinguishing outlier detection from novelty detection and comparing assumptions of several anomaly-detection methods.

Rights need review
Datasets

2D elastodynamic metamaterials

Pixelated metamaterial unit-cell designs and band-gap locations and widths for investigating simulation surrogate reliability.

License evidence recorded
Datasets

3W dataset

Oil-production process signals for studying event detection across operating scenarios.

License evidence recorded
Methods

5.2. Permutation feature importance

A model-inspection guide explaining permutation importance and its limitations when predictors are strongly correlated.

Rights need review