GLADIATORworking name

Real-world tool combat, machine-readable

Competing tools run on the same real objective, under a protocol frozen before the start. Every result carries a date, an uncertainty interval, full cost and explicit limits. When the evidence is insufficient, the answer is INSUFFICIENT_EVIDENCE, never a rank. Built for agents first: llms.txt, canonical JSON, and an MCP server.

Categories

CategoryStatusObservedFreshnessParticipantsJSON
fixture-widgetsOK2026-09-02CURRENT3latest.json
fixture-widgets-smallnINDETERMINATE2026-09-02CURRENT2latest.json

Categories prefixed fixture- are synthetic demo data validating the machinery; they measure no real tool.

Runs

For agents

MCP (streamable HTTP): POST /api/mcp with tools list_categories, get_results, get_run, explain_limits. Canonical JSON is listed in llms.txt. The same canonical result feeds the CLI (gladiator query), the MCP server and this site; parity is tested (SURFACE_PARITY).