TORNEO

Plan a coding-agent comparison

Test changes to small Python repositories.

Results are pending. No leader is established on any measure.

Planned evidence path
  1. Your task
  2. A code change
  3. Independent checks

Inspect the local demo

Start with a decision

Small Python tasks; production relevance is unproven.

Read the scope and limits

Season one concerns small, original Python repositories. Forty pilot tasks precede a separate official corpus. It does not establish performance on production codebases.

Keep the alternatives visible

Observed alternatives, not admitted participants.

Read the scope and limits

Two independent organizations mention five options: Claude Code, Codex CLI, GitHub Copilot coding agent, Cursor and Gemini CLI. Claude Code, Codex CLI and Cursor appear in both accounts. Mentions are not performance results.

What these observations establish

This observed consideration set was not frozen before the rights review. It is a candidate, not an admitted roster or the denominator of a published coverage claim. The Copilot account concerns its cloud coding agent, not local CLI adoption.

Inspect the local demo

The local demonstration kit exists. Code license approved: Apache-2.0; no code-agent result is established.

Local demonstration requirements and outputs

python3 template/run.py --dry-run --output build/c5-demo

Cost 0 USD, local demonstration only; no provider call.

Use Python 3.11 or later in a virtual environment with this repository installed. Choose an output directory that does not exist.

100 synthetic inputs; alpha-local, beta-local; report.html, result.json, MISSING-EVIDENCE.json.

Verify: python3 tools/verifier_reprise.py contrat build/c5-demo

  • TORNEO-019 approves Apache-2.0 for code and CC-BY-4.0 for published measurements and rankings. Third-party archives retain their original rights. The decision does not publish the kit or open the repository.
  • Consult the applicable license files before reuse.
  • You need rights to the corpus, execution and publication.
  • Each measured tool keeps its own terms; the kit does not waive them.
  • The demonstration uses two deterministic local functions, not coding agents.
  • Synthetic inputs and numbers do not establish real performance, cost or adoption.
  • A result applies to its frozen corpus, configurations and date.
  • A download or local rehearsal does not qualify as independent adoption.
  • Freeze and externally anchor the protocol before the first measured call.
  • Preserve traces and declare missing evidence.
  • Publish the corpus scope, configurations, dates and applicable rights.
  • Independent qualification requires a review of the exact evidence.

Prepare the evidence

Preserve the record before publication.

Read the scope and limits

Freeze the protocol before execution. Keep the corpus, traces, results and missing-evidence manifest together. Publication requires applicable rights and review of the exact evidence.