Skip to content

tide docs

Evaluation infrastructure for self-evolving agents, in two task regimes: one open-ended problem (autoresearch) and an ordered stream of tasks.

Read in this order:

  1. Get started: install, run a task and a stream, budgets, results, resume.
  2. Running agents: evaluate a supported harness, your own harness, or a method that is not an agent at all.
  3. Authoring tasks: new tasks, new benchmarks, converters.
  4. Metrics: the analyses over the results table.

Reference, when you need it: design is why tide is shaped the way it is, glossary is every term in one place, and the benchmark catalog is what ships.