An open library for your next question. Public pilot
Executable Science
Log inCreate account
← All research briefs
RESEARCH BRIEF · 6 SOURCES

Make an agent benchmark attempt auditable

Can another operator reconstruct one bounded agent-evaluation attempt from its task, environment, actions, failures and scoring rule?

Make this plan your own ↓

Read, edit and export without an account. This is preparation; no run or result is claimed.

What to compare

Reconstruct a single attempt before comparing aggregate rates. Keep retries and unsuccessful attempts in the denominator specified by the protocol. Author actions and generated tests cannot manufacture a platform-controlled verdict.

Useful outputs

  • One selected task and immutable data/code/environment identifiers with scoped rights records.
  • An explicit allowed-tools/network policy, model/prompt configuration, attempt budget and immutable scoring rule.
  • An attempt log preserving timeout, infrastructure failure, invalid output and scientific/task outcome separately.

Inputs and prerequisites

  • Benchmark payloads, issue text, reference code and model rights need their own checks.
  • Agent-generated code and benchmark environments require a qualified isolated execution path; ordinary host execution or a local container is insufficient.
  • Resolve SciCode's recorded route/dependency caveat and any moving data default before pinning an input.

What this work would not establish

  • None of these benchmarks or agents was executed by this collection.
  • No success rate, independent replication, safe sandbox qualification or superiority claim is made.

Start with these sources.

Software

SWE-bench

Software issue-resolution benchmark and evaluation tooling for studying patch assessment and reproducibility.

License evidence recorded
Software

mini-SWE-agent

Compact software engineering agent implementation for examining minimal agent loops and inspectable action traces.

License evidence recorded
Software

AgentBench

Multi-environment agent benchmark repository for investigating task adapters and evaluation consistency.

License evidence recorded
Software

SciCode

Scientific coding benchmark repository linking problem formulations and evaluation logic for researcher-oriented model assessment.

License evidence recorded
Software

EvalPlus

Code-generation evaluation tooling and expanded testing resources for studying test coverage and result reproducibility.

License evidence recorded
Software

HumanEval

Code-generation evaluation repository for examining functional correctness and pass-at-k estimation on programming tasks.

License evidence recorded
YOUR WORKING PLAN

Make the question your own.

Working copy · In this tab

Edit freely. Export a copy before reloading or leaving. Your text stays in this tab until you choose to copy, download or save it privately.