Plan a coding-agent comparison
Test changes to small Python repositories.
Results are pending. No leader is established on any measure.
- Your task
- A code change
- Independent checks
Start with a decision
Small Python tasks; production relevance is unproven.
Read the scope and limits
Season one concerns small, original Python repositories. Forty pilot tasks precede a separate official corpus. It does not establish performance on production codebases.
Keep the alternatives visible
Observed alternatives, not admitted participants.
Read the scope and limits
Two independent organizations mention five options: Claude Code, Codex CLI, GitHub Copilot coding agent, Cursor and Gemini CLI. Claude Code, Codex CLI and Cursor appear in both accounts. Mentions are not performance results.
What these observations establish
This observed consideration set was not frozen before the rights review. It is a candidate, not an admitted roster or the denominator of a published coverage claim. The Copilot account concerns its cloud coding agent, not local CLI adoption.
Inspect the local demo
The local demonstration kit exists. Code license approved: Apache-2.0; no code-agent result is established.
Local demonstration requirements and outputs
Use Python 3.11 or later in a virtual environment with this repository installed. Choose an output directory that does not exist.
100 synthetic inputs; alpha-local, beta-local; report.html, result.json, MISSING-EVIDENCE.json.
Verify: python3 tools/verifier_reprise.py contrat build/c5-demo
- TORNEO-019 approves Apache-2.0 for code and CC-BY-4.0 for published measurements and rankings. Third-party archives retain their original rights. The decision does not publish the kit or open the repository.
- Consult the applicable license files before reuse.
- You need rights to the corpus, execution and publication.
- Each measured tool keeps its own terms; the kit does not waive them.
- The demonstration uses two deterministic local functions, not coding agents.
- Synthetic inputs and numbers do not establish real performance, cost or adoption.
- A result applies to its frozen corpus, configurations and date.
- A download or local rehearsal does not qualify as independent adoption.
- Freeze and externally anchor the protocol before the first measured call.
- Preserve traces and declare missing evidence.
- Publish the corpus scope, configurations, dates and applicable rights.
- Independent qualification requires a review of the exact evidence.
Prepare the evidence
Preserve the record before publication.
Read the scope and limits
Freeze the protocol before execution. Keep the corpus, traces, results and missing-evidence manifest together. Publication requires applicable rights and review of the exact evidence.