An open library for your next question. Public pilot
Executable Science
Log inCreate account
← All research briefs
RESEARCH BRIEF · 6 SOURCES

Explain a score difference before blaming the model

Do evaluation harnesses agree on the same frozen predictions once normalization, aggregation and failure handling are made explicit?

Make this plan your own ↓

Read, edit and export without an account. This is preparation; no run or result is claimed.

What to compare

Hold predictions fixed while comparing scorers. Match denominators and case inclusion before comparing aggregate scores. Model calls are a later, separately budgeted step; distinguish task dependence from an IID binomial model.

Useful outputs

  • A small original prediction/answer fixture with exact expected scores and explicit failure rows.
  • A mapping of prompt/template, scorer, normalization, aggregation, truncation and cache conventions.
  • A source-version-pinned comparison report that separates fixture agreement from model performance.

Inputs and prerequisites

  • Read the pinned APIs and prepare adapters against a qualified dependency environment.
  • Benchmark data, model weights, provider terms and API spending are separate from harness code licenses.
  • The current qualified runtime does not include these external harnesses.

What this work would not establish

  • This is an unrun protocol proposal; it reports no harness discrepancy or model ranking.
  • The referenced uncertainty tools do not automatically validate every benchmark design.

Start with these sources.

Software

Inspect AI

Evaluation framework for composing tasks, model interactions, scoring and inspectable logs.

License evidence recorded
Software

Language Model Evaluation Harness

Task and model interfaces for reproducible language-model evaluation with inspectable prompting and scoring choices.

License evidence recorded
Software

LightEval

Language-model evaluation tooling for tracking prompt construction, scoring choices and evaluation configuration.

License evidence recorded
Software

TorchMetrics

Metrics for tensor-based learning systems, useful for testing reductions, state accumulation and distributed aggregation.

License evidence recorded
Software

Uncertainty Toolbox

Predictive-uncertainty analysis tools for examining calibration, sharpness and scoring on controlled regression predictions.

License evidence recorded
YOUR WORKING PLAN

Make the question your own.

Working copy · In this tab

Edit freely. Export a copy before reloading or leaving. Your text stays in this tab until you choose to copy, download or save it privately.