A registry of agent benchmarks (including many OSS agent trajectories)
A registry of agent benchmarks (including many OSS agent trajectories)
If you're interested in exploring what LLM-based agent systems these days actually do to solve certain benchmarks such as SWEBench or WebArena, we created a small leaderboard with our team, that allows to view a lot of public and OSS agent results including all the runtime traces (the step-by-step reasoning behind the scenes). Looking at traces is actually quite interesting, as they reveal a lot about the inner working and shortcomings of current agent system, e.g. see https://explorer.invariantlabs.ai/u/invariant/webarena--SteP... for an example trace.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Agent Smith Is OSS
Thredded forums are minimalistic and OSS
Dynmgrm – Operate DynamoDB with GORM (Golang OSS)
We Put Chromium on a Unikernel (OSS Apache 2.0)
My First OSS as a Teen
Benchmarks of UUID bintext codecs in Go
Crowdfunding forecasts and benchmarks
Capricorn – A browser for aliyun-oss
Capricorn – A browser for aliyun-oss
Thanks to SMC, firefighters with iPhones can track stairs climbed (OSS)