Make an agent benchmark attempt auditable
Can another operator reconstruct one bounded agent-evaluation attempt from its task, environment, actions, failures and scoring rule?
Make this plan your own ↓Read, edit and export without an account. This is preparation; no run or result is claimed.
What to compare
Reconstruct a single attempt before comparing aggregate rates. Keep retries and unsuccessful attempts in the denominator specified by the protocol. Author actions and generated tests cannot manufacture a platform-controlled verdict.
Useful outputs
- One selected task and immutable data/code/environment identifiers with scoped rights records.
- An explicit allowed-tools/network policy, model/prompt configuration, attempt budget and immutable scoring rule.
- An attempt log preserving timeout, infrastructure failure, invalid output and scientific/task outcome separately.
Inputs and prerequisites
- Benchmark payloads, issue text, reference code and model rights need their own checks.
- Agent-generated code and benchmark environments require a qualified isolated execution path; ordinary host execution or a local container is insufficient.
- Resolve SciCode's recorded route/dependency caveat and any moving data default before pinning an input.
What this work would not establish
- None of these benchmarks or agents was executed by this collection.
- No success rate, independent replication, safe sandbox qualification or superiority claim is made.
Start with these sources.
SWE-bench
Software issue-resolution benchmark and evaluation tooling for studying patch assessment and reproducibility.
mini-SWE-agent
Compact software engineering agent implementation for examining minimal agent loops and inspectable action traces.
AgentBench
Multi-environment agent benchmark repository for investigating task adapters and evaluation consistency.
SciCode
Scientific coding benchmark repository linking problem formulations and evaluation logic for researcher-oriented model assessment.
EvalPlus
Code-generation evaluation tooling and expanded testing resources for studying test coverage and result reproducibility.
HumanEval
Code-generation evaluation repository for examining functional correctness and pass-at-k estimation on programming tasks.
Make the question your own.
Edit freely. Export a copy before reloading or leaving. Your text stays in this tab until you choose to copy, download or save it privately.