# 🤖 fixed source-selection ledger at 2026-07-26 PDT
id	decision	family	work	status_at_cutoff	primary_source	reason	uncertainty
time_horizons	include	long-horizon evaluation	Measuring AI Ability to Complete Long Software Tasks	NeurIPS 2025 accepted	https://neurips.cc/virtual/2025/poster/119302	cross-model task-duration scaling with human task times	task suite is software-heavy and the 50% horizon is a fitted construct
tau_bench	include	planning and tool use	tau-bench	ICLR 2025 accepted	https://openreview.net/forum?id=roNSXZpUDN	stateful user-agent-tool interaction and consistency metric	retail and airline domains only
ultrahorizon	include	long-horizon planning	UltraHorizon	ICML 2026 regular	https://openreview.net/forum?id=qRNtMWrTvo	ultra-long simulated tasks expose compounding failures	accepted ICML version coexists with rejected ICLR submission
agentcompany	include	workplace agents	TheAgentCompany	NeurIPS 2025 datasets and benchmarks poster	https://openreview.net/forum?id=LZnKNApvhG	consequential multi-app workplace tasks	organization simulation is not a live firm
swe_agent	include	software engineering	SWE-agent	NeurIPS 2024 accepted	https://proceedings.neurips.cc/paper_files/paper/2024/hash/5a7c947568c1b1328ccc5230172e1e7c-Abstract-Conference.html	landmark agent-computer interface result	older anchor and original SWE-bench limitations apply
swe_bench_pro	include	software engineering	SWE-Bench Pro	ICML 2026 regular	https://openreview.net/forum?id=uEVTdoAbnK	long-horizon multi-file professional repositories	accepted ICML version coexists with rejected ICLR submission
swe_lancer	include	software engineering economics	SWE-Lancer	ICML 2025 accepted	https://proceedings.mlr.press/v267/miserendino25a.html	real freelance tasks with dollar values and failure evidence	one marketplace and historical model snapshot
utboost	include	evaluation reliability	UTBoost	ACL 2025 main	https://aclanthology.org/2025.acl-long.189/	test inadequacy and patch-validity audit of coding-agent claims	focuses SWE-bench and generated tests
osworld	include	computer use	OSWorld	NeurIPS 2024 datasets and benchmarks	https://proceedings.neurips.cc/paper_files/paper/2024/hash/5d413e48f84dc61244b6be550f1cd8f5-Abstract-Datasets_and_Benchmarks_Track.html	cross-application real-computer benchmark	older systems and a controlled VM setting
workarena	include	web and enterprise work	WorkArena	ICML 2024 accepted	https://proceedings.mlr.press/v235/drouin24a.html	common knowledge-work tasks in ServiceNow	one enterprise platform and older anchor
webrl	include	learning and adaptation	WebRL	ICLR 2025 accepted	https://openreview.net/forum?id=oVKEAFjEqv	self-evolving curriculum for open web agents	training gains may not transfer across websites or model families
browsecomp	include	browsing research	BrowseComp	OpenAI first-party report and arXiv preprint 2025	https://openai.com/index/browsecomp/	frontier difficult browsing benchmark that shaped current evidence	unreviewed first-party report with potential benchmark exposure
paperbench	include	research agents	PaperBench	ICML 2025 accepted	https://proceedings.mlr.press/v267/starace25a.html	end-to-end replication scored by paper-specific rubrics	only 20 papers and expensive expert grading
rebench	include	research agents	RE-Bench	ICML 2025 accepted	https://proceedings.mlr.press/v267/wijk25a.html	time-budgeted comparison against human experts	seven environments and non-comparable model scaffolds
scienceagentbench	include	scientific agents	ScienceAgentBench	ICLR 2025 accepted	https://openreview.net/forum?id=6z4YKr0GK6	task-level data-driven discovery evaluation	science workflow coverage is selective
mle_bench	include	scientific and ML engineering	MLE-bench	ICLR 2025 oral	https://iclr.cc/virtual/2025/oral/31914	75 real Kaggle competitions with contamination analysis	competition medals are a proxy for research ability
memoryagentbench	include	memory and adaptation	MemoryAgentBench	ICLR 2026 accepted	https://openreview.net/forum?id=DT7JyQC3MR	incremental multi-turn memory taxonomy and controlled evaluation	new benchmark with limited independent replication
multiagentbench	include	multi-agent coordination	MultiAgentBench	ACL 2025 main	https://aclanthology.org/2025.acl-long.421/	collaboration and competition with coordination metrics	simulated protocols may not predict open-world teams
agentdojo	include	security and reliability	AgentDojo	NeurIPS 2024 datasets and benchmarks	https://proceedings.neurips.cc/paper_files/paper/2024/hash/97091a5177d8dc64b1da8bf3e1f6fb54-Abstract-Datasets_and_Benchmarks_Track.html	dynamic prompt-injection evaluation over utility and security	attacks and defenses age quickly
agentharm	include	safety	AgentHarm	ICLR 2025 accepted	https://proceedings.iclr.cc/paper_files/paper/2025/hash/c493d23af93118975cdbc32cbe7323f5-Abstract-Conference.html	harmful autonomous task benchmark and jailbreak sensitivity	harm taxonomy and sandbox tasks are necessarily bounded
st_webagentbench	include	safety and trustworthy web use	ST-WebAgentBench	ICLR 2026 accepted	https://openreview.net/forum?id=MuCDzH0ctf	joint task completion and policy compliance	landing-page task-count versions conflict
ai_control	include	control	AI Control	ICML 2024 accepted	https://proceedings.mlr.press/v235/greenblatt24a.html	landmark adversarial-control experiment with trusted and untrusted models	older narrow coding proxy
androidworld	exclude	computer use	AndroidWorld	ICLR 2025 accepted	https://openreview.net/forum?id=il5yUQsrjC	mobile-only coverage was redundant after OSWorld and WorkArena	meaningful omission if Android deployment is central
toolsandbox	exclude	planning and tool use	ToolSandbox	NAACL 2025 findings	https://aclanthology.org/2025.findings-naacl.65/	lower venue tier and overlap with tau-bench	may contain distinct state-dependency evidence
core_bench	exclude	software engineering	CORE-Bench	arXiv preprint 2026	https://arxiv.org/abs/2606.11864	ambiguous search match and no verified premier acceptance	possible recent benchmark missed by the fixed cutoff search
magentic_one	exclude	multi-agent systems	Magentic-One	Microsoft technical report 2024	https://www.microsoft.com/en-us/research/publication/magentic-one-a-generalist-multi-agent-system-for-solving-complex-tasks/	system report lacked a uniquely necessary controlled result after MultiAgentBench	first-party evidence may still inform systems engineering
swe_bench_audit	exclude	evaluation reliability	What's in a Benchmark? The Case of SWE-Bench	ACM publication 2026	https://dl.acm.org/doi/10.1145/3786583.3786904	indirect benchmark-audit context rather than a flagship agent-capability paper	its findings could sharpen the leakage section in later revision
controlarena	exclude	control	ControlArena	AISI library and technical material 2025	https://control-arena.aisi.org.uk/	no premier paper was verified and AI Control supplies the accepted anchor	a later accepted evaluation may supersede this decision
riosworld	exclude	computer-use risk	RiOSWorld	NeurIPS 2025 accepted	https://proceedings.neurips.cc/paper_files/paper/2025/hash/0c79d6ed1788653643a1ac67b6ea32a7-Abstract-Conference.html	fixed breadth cap favored AgentDojo plus ST-WebAgentBench for security and policy evidence	important risk-benchmark omission
asb	exclude	agent security	Agent Security Bench	ICLR 2025 accepted	https://openreview.net/forum?id=V4y0CpX4hK	overlapped AgentDojo and AgentHarm under the fixed breadth cap	broader attack taxonomy may reveal missed failure modes
longmemeval	exclude	memory	LongMemEval	ICLR 2025 accepted	https://iclr.cc/virtual/2025/poster/28290	chat-assistant memory is less action-coupled than MemoryAgentBench	stronger longitudinal evidence may merit later inclusion
webcoach	exclude	memory and learning	WebCoach	ICLR 2026 accepted	https://iclr.cc/virtual/2026/10021277	WebRL plus MemoryAgentBench more directly isolate training and memory evidence	cross-session guidance is a reachable frontier omitted for breadth
membench	exclude	memory	MemBench	ACL 2025 findings	https://aclanthology.org/2025.findings-acl.989/	lower venue tier and overlap with MemoryAgentBench	could broaden memory-type coverage
osworld_2	exclude	computer use	OSWorld 2.0	arXiv preprint 2026	https://arxiv.org/abs/2606.29537	very recent unreviewed update not necessary for the fixed accepted-evidence core	likely changes absolute frontier numbers
ultrahorizon_iclr_submission	exclude	duplicate record	UltraHorizon ICLR submission	ICLR 2026 rejected	https://openreview.net/forum?id=FTZfVHWAIq	duplicate rejected submission; accepted ICML record selected	title and text version differ
swe_bench_pro_iclr_submission	exclude	duplicate record	SWE-Bench Pro ICLR submission	ICLR 2026 rejected	https://openreview.net/forum?id=9R2iUHhVfr	duplicate rejected submission; accepted ICML record selected	version differences require care
