new in v0.8Duo ships built in

OpenCollab

Do your agents really collaborate?

OpenCollab is an open-source multi-agent coding framework. Declare a team in one file, run it on a runtime that holds everything else fixed, and check in the event stream whether the team you declared is the team that ran.

the problem

The illusion of collaboration

Multi-agent teams on repository-level tasks usually begin by configuring roles such as an Analyst, a Coder and a Tester. Left to decide on their own, models rarely delegate. Mostly the Adopter, the agent that receives the task, bypasses its teammates and resolves the issue alone; in other cases it spends its token allowance before issuing any handoff. Adherence, the share of runs in which the team works as declared, frequently falls below half as the system collapses into a single-agent loop. The paper calls this the illusion of collaboration.

Prompts and log filters do not fix it. Mandating delegation in the prompt makes Adherence swing with the phrasing. Keeping only the runs where agents collaborated biases the comparison: agents solve easy tasks alone and ask for help when stuck, so that subset skews toward harder tasks.

Declared team versus what often runs Left: an Adopter connected by dashed declared edges to an analyst, a coder and a tester. Right: the same team where only the Adopter acts and the teammates are bypassed. declared what often runs Adopter Analyst Coder Tester works alone Adopter Analyst Coder Tester teammates bypassed
Declared edges are dashed. In many runs only the Adopter acts, and the team collapses into a single-agent loop.
shared runtimeOne gate before every model call

Every role runs through the same session loop, and pre-call admission enforces each role's token allowance, so no Adopter silently spends the run before it could delegate.

event streamRecords instead of prose logs

Every action and every token is bound to the role that produced it, so scripts can reconstruct who did what without parsing natural language.

adherenceA verdict for every run

From the stream we check each run: did every teammate act, did the Adopter delegate, did each role keep to its tools, budget and message routes. A run counts only if every check passes, which is what lets a gain be credited to the team.

how it works

Declare the organization, hold everything else, record what ran

OpenCollab separates orchestration from execution. Every role runs as a session through one ten-state loop that handles context, model calls and tools. Controllers above it govern topology, lifecycle and message routing: a model-directed Team, a scripted Workflow, or a Single agent.

declare

One file states the team

Role prompts, per-role toolsets and a directed topology of who may address whom. A Workflow is a short Python module that fixes call order, fan-out and termination.

# configs/team.example.yaml (excerpt)
entry: lead
roles:
  analyst:
    tools: [bash, file_read, grep]
  coder:
    tools: [bash, file_read,
            file_write, apply_patch,
            grep, message_agent]
  reviewer:
    tools: [bash, file_read,
            git_diff, grep]
  # lead and all prompts omitted
topology:
  lead: [analyst, coder, reviewer]
  analyst: [coder]
  coder: [reviewer]
  reviewer: [coder]
hold

Runtime gates, not post-hoc filters

Five non-organizational factors are held identical across configurations at invocation time.

  • modelone base model for every role
  • toolsa call outside the role's toolset is refused and recorded
  • budgetpre-call admission against a per-role token ceiling; an overrun halts the session with a termination record
  • contextone compaction policy for every configuration
  • topologya message off a declared edge is refused and recorded
record

An event stream, not prose logs

Every step appends a typed record stamped with its role, so scripts reconstruct the realized topology and each role's cost.

each record
steptimestamprunroletypepayload
seven event types
  • Model call
  • Tool execution
  • Agent finish
  • Session termination
  • Refused spawn
  • Worktree change
  • Context shaping

Structural adherence

Adherence asks of every run: did the team you declared actually run? Six checks answer it. For each applicable check a run gets a verdict: adherent, deviant, or unverifiable. A run is adherent only if every applicable axis passes.

participationdelegationrole boundarybudget sharinginformation flowcontext policy
adh(r) = ∏x ∈ 𝒳(Zr) 1[ vx(r) = adherent ]
CACE = ITT / α̂adh = ΔPass@1 / α̂adh

α̂adh is the share of a configuration's runs with adh(r) = 1; a run in which only one of two required teammates acts has adh(r) = 0. Under mean exclusion, ITT / α̂adh identifies the configuration-level complier effect.

239 lines

reproduce the core multi-agent pattern of edict on OpenCollab. The same protocol ships as examples/mini-edict.

edict~24k
on OpenCollab239

about 100× less code for the same organization

OpenCollab overview: a specification of task, model, role prompts, topology and policies feeds Single, Team or Workflow control; all run on a shared session runtime with isolated worktrees and a common session loop; an event stream is audited to give assigned Z, realized A, result Y and cost C.
OpenCollab overview. Single, Team, and Workflow controllers share a session runtime. Configuration and execution traces support Adherence and resource cost auditing.
results · system-level performance

OpenCollab (Duo) sets the highest Pass@1 on all three benchmarks

Five harnesses run with GPT-5.6-Luna on SWE-bench Pro, Terminal-Bench 2.1 and DeepSWE. OpenCollab (Base) is the single agent. OpenCollab (Duo) is a Workflow in which code issues every handoff, so both Coders run on every task.

HarnessPass@1 (%)Avg. tokens (M)Avg. cost ($)Cache hit (%)

Widest gain on DeepSWE

Duo 69.91% against Base 55.75%. Paired by task, Duo solves far more tasks the single agent fails than the other way round.

Fewest tokens

OpenCollab (Base) uses the fewest tokens and the lowest cost on all three benchmarks. Duo spends more, since it runs two complete solving processes and selects between their candidates, yet costs less than Claude Code on each.

results · dimensions of adherence

Changing any single dimension shifts Adherence from 47.2% to as high as 97.2%

A model-directed team is audited through its event stream while one dimension changes at a time. The reference team uses Qwen3.8-Flash, open prompt cards, unrestricted tools and 2M tokens per role, on a random sample of SWE-bench Pro tasks. Unconstrained defaults yield 47.2% Adherence; restricting tool boundaries prevents isolated execution and lifts it above 90%, so collaboration can be controlled systematically.

variantadherence · 95% Clopper–Pearson CIunverif.pass@1tokens

Two timelines on a shared time axis. (a) Read-only variant: the adopter delegates to Coders and adopts Coder A's patch; graded code passes. (b) Reference team: the adopter briefs Coders but keeps executing until it exhausts its budget, never submitting teammate work; graded code fails.
Two execution trajectories on a shared time axis. Lanes denote agents; blocks represent model calls colored by tool usage. (a) Read-only variant: restricting the adopter to read-only tools removes its ability to edit files, compelling delegation to Coders and the adoption of Coder A's patch. (b) Reference team: the adopter briefs both coders but monopolizes execution until exhausting its token budget, never submitting teammate work.
results · causal validity

Adherence decides how far a gain can be credited to the organization

Each configuration runs the same tasks once and is paired task by task with the Single agent. ITT is the Pass@1 difference from Single; CACE divides it by Adherence.

ConfigurationAdherence (%)Pass@1 (%)ITTCACE
Single–69.4––
Open card (reference)47.2 [47.2, 61.1]75.0+5.6+11.8 [+9.1, +11.8]
Mandatory card91.7 [91.7, 100.0]72.2+2.8+3.0 [+2.8, +3.0]
Budget stated, 2M75.0 [75.0, 83.3]69.40.00.0 [0.0, 0.0]
Read-only94.4 [94.4, 97.2]52.8−16.7−17.6 [−17.6, −17.1]

Brackets show sensitivity to the labeling rule for unverifiable verdicts and are not confidence intervals.

open card · split by Ar
Ar = 0not realized
+15.8
Ar = 1realized
−5.9

Under the Open card, 47.2% Adherence rescales an ITT of +5.6 to a CACE of +11.8. The +5.6 arises on the tasks where the team did not form; the adherent tasks show −5.9, so +11.8 should not be read as the complier effect. Under the Mandatory card, the paired estimate over adherent runs coincides with CACE (+3.0).

Team minus Single in Pass@1 points; axis from −32 to +32. OpenCollab records Ar for every run on tasks also run by Single, so this contrast is computed rather than assumed.

results · controlled comparison

Seven conditions for a controlled comparison

Five conditions hold non-organizational factors fixed; two verify what ran. The paper audits ten artifacts: none satisfies all seven, and OpenCollab meets all seven.

held · can the factor be set explicitly?verified · can the run be checked?
ArtifactModelToolsBudgetContextTopologyRealizedCompared
✓by the artifact's own settings and records –only through the researcher's code, or in part ×not met
quick start

Start in three commands

OpenCollab works with any OpenAI-compatible or Anthropic endpoint. Declared teams, the Duo workflow and Mini Edict are covered in the README.


      

the paper

OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime

Four panels: (a) organizations in existing frameworks, (b) single versus team under shared tasks, model and tools, (c) a configured analyst-coder-tester topology versus the observed event timeline, (d) auditing a trace against its policy.
Perspectives on multi-agent collaboration: (a) organization, (b) benefit and cost, (c) configured versus observed execution, and (d) collaboration auditing.
@misc{hsu2026opencollabmultiagentcodingframework,
  title         = {OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime},
  author        = {Chun-Wah Hsu and Kai Gong and Yu Wu and Xianhe Chen and Mengyang Liu and Jie Li and Hanyu Li and Zhixuan Liu and Naisheng Tang and Jiaying Chi and Ziheng Fan and Xuning He and Xiaokang Yang and Xue Jiang and Yihong Dong},
  year          = {2026},
  eprint        = {2609.38345},
  archivePrefix = {arXiv},
  primaryClass  = {cs.SE},
  url           = {https://arxiv.org/abs/2609.38345}
}