Background and Problems
In a single-agent setup, one model often has to handle requirement understanding, retrieval, planning, tool invocation, code generation, and result validation at the same time. As tasks become complex, prompts quickly balloon, and the context becomes filled with irrelevant information, leading to unstable outputs. As a result, many teams begin experimenting with multi-agent collaboration: assigning different roles to handle planning, execution, review, and summarization.

However, multi-agent systems are not a silver bullet. Without clear role boundaries, the system can degrade into multiple models simply forwarding text to one another; without a unified message bus, call chains can become a tangled web that is difficult to troubleshoot; without conflict resolution mechanisms, different agents may repeatedly argue based on different facts or goals, ultimately increasing latency and cost. This article focuses on three engineering questions: how to divide responsibilities, how to communicate, and how to arbitrate.
Core Concepts
Role Division: Define Contracts Before Writing Prompts
The first step in a multi-agent system is not choosing models, but splitting responsibilities. Common roles include: Planner, responsible for task decomposition and priority ordering; Researcher, responsible for retrieval and fact organization; Executor, responsible for invoking tools or executing code; Reviewer, responsible for quality checks; Arbitrator, responsible for conflict arbitration; and Memory Manager, responsible for long-term memory and summary compression.
Each role should have a clear contract: what the input is, what the output is, which tools can be invoked, what it must not do, and when it should finish. For example, the Reviewer should not directly rewrite the final solution, but should output structured review comments; the Researcher should not make business decisions, but should provide evidence with sources. This reduces role overreach and makes later unit evaluation easier.
Message Bus: Turn Conversations into Traceable Events
Many early multi-agent implementations use point-to-point calls: A calls B, B calls C, and C calls back A. This approach is hard to debug and can easily create loops. A more robust approach is to introduce a message bus and abstract interactions between agents as events. Events may include task.created, plan.proposed, tool.result.received, review.failed, conflict.detected, and so on.
A message envelope should include at least a trace ID, sender, recipient or topic, message type, payload, schema version, timestamp, and expiration time. The trace ID is used for end-to-end replay, the schema version supports compatible evolution, and the expiration time prevents stale messages from triggering side effects.
from dataclasses import dataclass
@dataclass
class AgentMessage:
msg_id: str
trace_id: str
sender: str
topic: str
payload: dict
schema_version: str = 'v1'
This code defines a unified message structure. In real projects, you can also add signatures, priority, retry counts, and idempotency keys. The message bus itself can be implemented with an in-memory queue, Redis Stream, Kafka, or a cloud messaging service. The key is not the component itself, but whether the message contract is stable.
Conflict Resolution: Move from Debate to Adjudication
Multi-agent conflicts generally fall into four categories: fact conflicts, plan conflicts, resource conflicts, and policy conflicts. Fact conflicts occur when different agents present different facts; plan conflicts arise from inconsistent execution order; resource conflicts happen when multiple tasks compete for the same tool, file, budget, or external interface; policy conflicts stem from different risk preferences—for example, one agent prefers automatic execution while another requires human confirmation.
When resolving conflicts, it is not advisable to let multiple agents engage in endless dialogue. A better approach is to define adjudication rules: safety issues should be escalated to humans first; for factual issues, prefer results with evidence, more recent timestamps, and more credible sources; planning issues should be decided by the Planner or Arbitrator according to the objective function; resource issues should be controlled through locks, queues, budgets, and idempotency keys.
def resolve_conflict(conflicts):
if any(c.kind == 'safety' for c in conflicts):
return 'human_review'
if any(c.kind == 'fact' for c in conflicts):
return pick_best_evidence(conflicts)
return planner.replan(conflicts)
This pseudocode illustrates the idea of conflict triage: safety conflicts must not be silently absorbed; factual conflicts are handled through evidence ranking; planning conflicts are returned to the planner for reordering. In production environments, you should also record the reasons for decisions to support retrospectives and evaluation.
Practical Steps and Checklist
- Step 1: Define task boundaries. First clarify which tasks are suitable for automation and which require human confirmation. Do not let multi-agent systems fully automate high-risk operations from the start.
- Step 2: Split minimal roles. Start with Planner, Executor, and Reviewer, and add Researcher, Critic, or Memory Manager only when truly necessary.
- Step 3: Design message contracts. Define fields and validation rules for each event type, and avoid passing long free-form text contexts directly.
- Step 4: Establish shared state. Place task goals, constraints, completed steps, and pending confirmations into a structured state object, rather than letting each agent maintain its own memory.
- Step 5: Add budget controls. Set maximum rounds, maximum tokens, maximum tool invocations, and maximum duration for each task.
- Step 6: Implement conflict arbitration. Make clear who has final decision authority and preserve the evidence chain.
- Step 7: Perform trajectory evaluation. Evaluate not only the final answer, but also whether the plan is reasonable, whether tool calls are necessary, and whether reviews identify issues.
Common Pitfalls and Recommendations
- Too many agents. The finer the roles, the higher the communication cost. Prioritize testability rather than making the architecture diagram look impressive.
- Blurry prompt responsibilities. If one agent can plan, execute, and review, it can easily reinforce its own errors. Responsibilities should be mutually exclusive.
- Missing trace_id. Without a trace ID, multi-agent issues are almost impossible to reconstruct. Every message, tool call, and model response should be linked to the same task.
- Unbounded context concatenation. Feeding all historical messages to every agent raises costs and dilutes attention. Use summaries, retrieval, and structured state.
- Ignoring idempotency. Message retries can cause duplicate tool calls or duplicate data writes. Tool calls should include idempotency keys, and write operations should be safely replayable.
- No stopping condition. Two agents asking each other for revisions can fall into an infinite loop. Set maximum rounds and escalation paths.
- Evaluating only the final result. Problems in multi-agent systems often appear in intermediate steps. Establish separate metrics for planning, retrieval, execution, and review.
Directions for Further Reading
- Multi-agent frameworks and protocols. You can follow developments in LangGraph, AutoGen, CrewAI, OpenAI Swarm, MCP, A2A, and related directions. For specific capabilities and interfaces, refer to the official documentation.
- Distributed system patterns. Concepts such as event sourcing, Saga, eventual consistency, idempotent consumers, and dead-letter queues are highly relevant to multi-agent communication and recovery mechanisms.
- Evaluation systems. Further explore trajectory evaluation, LLM-as-judge, human spot checks, and regression benchmark suites, rather than looking only at single outputs.
- Security and permissions. Once multi-agent systems can invoke tools, least privilege, audit logs, sandboxed execution, and human approval become more important than model capability.
References
Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.
Contact: chenxj.g@gmail.com
