Agent orchestration benchmark methodology

A SwarmForge orchestration benchmark is a reproducible evaluation of coding tasks executed by coordinated workers under a recorded environment, model and task configuration. SwarmForge has not published orchestration benchmark results; this page defines a proposed reporting methodology, not measured outcomes.

Define the question and task set

Each study should identify its question, task selection procedure, repositories and base commits, permitted operations, expected deliverables and independent correctness checks. State whether the study measures independent tasks, dependent tasks, repository review or integrated feature delivery.

Compare like-for-like conditions and report failed attempts, timeout cases, exclusions and human interventions. Task completion reported by a worker is not equivalent to a correct, reviewed source handoff.

Record the execution environment

For each run, record the date, SwarmForge version and commit, OpenCode version, VM provider and snapshot, host/guest resources, repository commit, concurrency and queue limits, timeouts, model provider and exact model ID. Include relevant inference parameters and the manager’s instructions or versioned prompt where available.

Do not expose tokens, private repository contents or provider credentials in a public run bundle. Publish enough non-secret information to reproduce the setup and understand limits.

Publish an evidence table

Field Required evidence
Model and provider Exact IDs, endpoint protocol and model/inference settings
Task Versioned instructions, repository and base commit
Completion rate Numerator, denominator and independent grading rule
Time Queue, provisioning, coding, review and end-to-end wall time
Cost Actual provider charges when available; separately label estimates and unpriced usage
Token usage Input/output/reasoning/cache categories and measurement source
Interventions Follow-ups, restarts, retries, corrections and operator actions
Methodology Task selection, baselines, repetition count, exclusions and uncertainty
Environment and date Snapshot, resources, versions, concurrency and run timestamps
Artifacts Reports, test output, result records and source handoff commits

Separate capability from outcomes

A configured worker limit is not a benchmark of throughput. Metrics-derived cost estimates are not billing statements. Report correctness, source integration and artifact preservation separately, because a completed coding turn can still need review or recovery.

Future result pages should include their original evidence, limitations and methodology links. Do not publish placeholder scores, invented speedups or product comparison rankings before a study is complete.