Datasets
LAB-Bench
Upstream project repository
Benchmark repository for evaluating language models on biology and laboratory research knowledge tasks.
Investigate Can a small non-procedural, rights-reviewed subset support an auditable scoring and contamination-risk review?
Listing reviewed · source terms apply↗
Software
Language Model Evaluation Harness
Upstream project repository
Task and model interfaces for reproducible language-model evaluation with inspectable prompting and scoring choices.
Investigate Can a tiny authored multiple-choice task expose the effect of length normalization on the selected answer?
Listing reviewed · source terms apply↗
Datasets
Leaf
UCI Machine Learning Repository
Leaf image descriptors for comparing small botanical classification baselines.
Investigate Which leaf classes are unstable under repeated fixed splits and train-only feature scaling?
Listing reviewed · source terms apply↗
External articles
Learning with Noisy Labels [~Re]visited
ReScience C
A replication report on learning from labels that contain errors, with source metadata identifying a Python deep-learning implementation.
Investigate How does a prespecified label-corruption schedule affect held-out accuracy and calibration across fixed seeds?
Listing reviewed · source terms apply↗
Datasets
LED Display Domain
UCI Machine Learning Repository
A synthetic display-recognition problem with controllable observation noise.
Investigate Does a classifier's error track the documented LED noise mechanism on a frozen evaluation design?
Listing reviewed · source terms apply↗
Software
LightEval
Upstream project repository
Language-model evaluation tooling for tracking prompt construction, scoring choices and evaluation configuration.
Investigate Can two declared scoring configurations be compared without changing task examples or model revision?
Listing reviewed · source terms apply↗
Software
LiteLLM
Upstream project repository
Model-provider interface and proxy tooling for inspecting request normalization, usage accounting and backend differences.
Investigate Can a local mock provider show whether retries and usage fields remain traceable without issuing any real model calls?
Listing reviewed · source terms apply↗
Software
LiveCodeBench
Upstream project repository
Code-model benchmark repository with time-aware task collections for studying evaluation provenance and data contamination.
Investigate Can task timestamps and release identifiers support a reproducible evaluation cut without silently mixing newer problems?
Listing reviewed · source terms apply↗
Software
llama.cpp
Upstream project repository
C/C++ language-model inference implementation for studying quantization and execution tradeoffs across supported hardware.
Investigate What output deviation and resource changes accompany one fixed quantization choice on an independently licensed tiny model?
Listing reviewed · source terms apply↗
Datasets
Madelon
UCI Machine Learning Repository
A constructed feature-selection problem for testing signal recovery among distracting dimensions.
Investigate Does feature selection improve Madelon classification when selection is nested inside training folds?
Listing reviewed · source terms apply↗
Datasets
MAGIC Gamma Telescope
UCI Machine Learning Repository
Simulated telescope event features for studying signal-background discrimination.
Investigate How do classwise calibration and low-false-positive performance compare for two bounded gamma-event baselines?
Listing reviewed · source terms apply↗
Datasets
Metro Interstate Traffic Volume
UCI Machine Learning Repository
Aggregate road-traffic counts and weather context for evaluating demand forecasts over time.
Investigate Does a seasonal traffic baseline outperform a weather-assisted model on a final chronological holdout?
Listing reviewed · source terms apply↗
Recorded license evidence is scoped to each source; it is not a blanket permission or a verification result.