Tracecore: Benchmark AI Agents on Deterministic Coding Tasks
Tracecore: Benchmark AI Agents on Deterministic Coding Tasks
I'm sharing Tracecore, an OS tool I'm building to evaluate AI agents' ability to handle deterministic software tasks like log triage, config remediation, and incident recovery. It started as a way to test whether agents could reliably perform structured operations without free-form guessing, inspired by frustrations with brittle automation in ops workflows. What sets it apart: Unlike benchmarks like SWE-Bench (which tests code generation on open-ended GitHub issues) or general agent evaluation suites (that mix diverse reasoning, coding, and interaction tasks), Tracecore focuses on deterministic episodes where agents must use constrained actions (e.g., file operations, ops triage) to achieve exact outcomes, with strict validation. It includes 15+ tasks across suites like operations and games, and supports running agents via adapters for frameworks like OpenClaw and Autogen, or custom scripts. You can try it out by installing through pip/uv, or by cloning the repo and installing the optional dev dependencies, and running the dashboard, the cli wizard or the cli commands. It outputs structured results with success/failure, steps used, traces for analysis, diffs, bundles and more. I've been iterating on this over the past few weeks, adding new tasks and improving the harness. Previous discussions on AI eval tools were helpful in shaping the design. Feedback welcome, especially on expanding task suites or integration ideas.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Cua-Bench – a benchmark for AI agents in GUI environments
WebGL Sprites Benchmark
NAB – The Numenta Anomaly Benchmark
NAB – The Numenta Anomaly Benchmark
AgentMafia – A Social Deduction Benchmark
Manage coding norms across your AI agents
A New Implementation of the Seven GUIs Benchmark
Webbench, a WASM Based Benchmark
An open benchmark for AI agents that test APIs
LLM Deceptiveness and Gullibility Benchmark