[{"data":1,"prerenderedAt":96},["ShallowReactive",2],{"PortalFooter_QX9Sj6L0jYa0f3sHh8puk8D6pkI69QhELaa1ZMuWxNo":3,"article-ar-多-agent-协作落地-角色分工-消息总线与冲突消解的工程实践":10,"article-body-v-0-0-0":32,"AppImage_U2SWsmBV1mDznAcDcnH397RslvGSkZeLYKgmLBsA":33,"PortalBreadcrumb_yrk2AtPrXsKFZJhIyeTD0R4a2WAywceLd1kQYcFeOIg":40,"related-ar-article-多-agent-协作落地-角色分工-消息总线与冲突消解的工程实践":47,"PostCard_HtNHQAkw3V43kL9Q3RIE2xzdetuj6nammYg3vtL3c":75,"PostCard_5MrdQReyjlh7tlwvUW44Dx2TYhaatLpjOsynzYD4wPk":82,"PostCard_Ds2Jr0fmMZOVJLg7hvcMHXzDwdLKPhniDm4KJBvQo":89},["Island",4],{"key":5,"result":6},"PortalFooter_QX9Sj6L0jYa0f3sHh8puk8D6pkI69QhELaa1ZMuWxNo",{"head":7},{"link":8,"style":9},[],[],{"id":11,"locale":12,"slug":13,"type":14,"title":15,"summary":16,"content_html":17,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"published_at":23,"created_at":24,"updated_at":25,"alternates":26},4612,"en","多-agent-协作落地-角色分工-消息总线与冲突消解的工程实践","article","Multi-Agent Collaboration in Practice: Role Division, Message Bus, and Conflict Resolution Engineering Practices","A multi-agent system is not simply stacking multiple large models; it requires clear responsibility boundaries, communication pathways, and conflict arbitration. From an engineering perspective, this article reviews key methods for role design, message buses, and conflict resolution, and provides actionable checklists and common pitfalls.","\u003Ch2>Background and Problems\u003C\u002Fh2>\u003Cp>In a single-agent setup, one model often has to handle requirement understanding, retrieval, planning, tool invocation, code generation, and result validation at the same time. As tasks become complex, prompts quickly balloon, and the context becomes filled with irrelevant information, leading to unstable outputs. As a result, many teams begin experimenting with multi-agent collaboration: assigning different roles to handle planning, execution, review, and summarization.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Flandscape_mountain.jpg\" alt=\"观山静思\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>观山静思\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp>However, multi-agent systems are not a silver bullet. Without clear role boundaries, the system can degrade into multiple models simply forwarding text to one another; without a unified message bus, call chains can become a tangled web that is difficult to troubleshoot; without conflict resolution mechanisms, different agents may repeatedly argue based on different facts or goals, ultimately increasing latency and cost. This article focuses on three engineering questions: how to divide responsibilities, how to communicate, and how to arbitrate.\u003C\u002Fp>\u003Ch2>Core Concepts\u003C\u002Fh2>\u003Ch3>Role Division: Define Contracts Before Writing Prompts\u003C\u002Fh3>\u003Cp>The first step in a multi-agent system is not choosing models, but splitting responsibilities. Common roles include: \u003Cstrong>Planner\u003C\u002Fstrong>, responsible for task decomposition and priority ordering; \u003Cstrong>Researcher\u003C\u002Fstrong>, responsible for retrieval and fact organization; \u003Cstrong>Executor\u003C\u002Fstrong>, responsible for invoking tools or executing code; \u003Cstrong>Reviewer\u003C\u002Fstrong>, responsible for quality checks; \u003Cstrong>Arbitrator\u003C\u002Fstrong>, responsible for conflict arbitration; and \u003Cstrong>Memory Manager\u003C\u002Fstrong>, responsible for long-term memory and summary compression.\u003C\u002Fp>\u003Cp>Each role should have a clear contract: what the input is, what the output is, which tools can be invoked, what it must not do, and when it should finish. For example, the Reviewer should not directly rewrite the final solution, but should output structured review comments; the Researcher should not make business decisions, but should provide evidence with sources. This reduces role overreach and makes later unit evaluation easier.\u003C\u002Fp>\u003Ch3>Message Bus: Turn Conversations into Traceable Events\u003C\u002Fh3>\u003Cp>Many early multi-agent implementations use point-to-point calls: A calls B, B calls C, and C calls back A. This approach is hard to debug and can easily create loops. A more robust approach is to introduce a message bus and abstract interactions between agents as events. Events may include \u003Ccode>task.created\u003C\u002Fcode>, \u003Ccode>plan.proposed\u003C\u002Fcode>, \u003Ccode>tool.result.received\u003C\u002Fcode>, \u003Ccode>review.failed\u003C\u002Fcode>, \u003Ccode>conflict.detected\u003C\u002Fcode>, and so on.\u003C\u002Fp>\u003Cp>A message envelope should include at least a trace ID, sender, recipient or topic, message type, payload, schema version, timestamp, and expiration time. The trace ID is used for end-to-end replay, the schema version supports compatible evolution, and the expiration time prevents stale messages from triggering side effects.\u003C\u002Fp>\u003Cpre>\u003Ccode>from dataclasses import dataclass\n\n@dataclass\nclass AgentMessage:\n    msg_id: str\n    trace_id: str\n    sender: str\n    topic: str\n    payload: dict\n    schema_version: str = 'v1'\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>This code defines a unified message structure. In real projects, you can also add signatures, priority, retry counts, and idempotency keys. The message bus itself can be implemented with an in-memory queue, Redis Stream, Kafka, or a cloud messaging service. The key is not the component itself, but whether the message contract is stable.\u003C\u002Fp>\u003Ch3>Conflict Resolution: Move from Debate to Adjudication\u003C\u002Fh3>\u003Cp>Multi-agent conflicts generally fall into four categories: fact conflicts, plan conflicts, resource conflicts, and policy conflicts. Fact conflicts occur when different agents present different facts; plan conflicts arise from inconsistent execution order; resource conflicts happen when multiple tasks compete for the same tool, file, budget, or external interface; policy conflicts stem from different risk preferences—for example, one agent prefers automatic execution while another requires human confirmation.\u003C\u002Fp>\u003Cp>When resolving conflicts, it is not advisable to let multiple agents engage in endless dialogue. A better approach is to define adjudication rules: safety issues should be escalated to humans first; for factual issues, prefer results with evidence, more recent timestamps, and more credible sources; planning issues should be decided by the Planner or Arbitrator according to the objective function; resource issues should be controlled through locks, queues, budgets, and idempotency keys.\u003C\u002Fp>\u003Cpre>\u003Ccode>def resolve_conflict(conflicts):\n    if any(c.kind == 'safety' for c in conflicts):\n        return 'human_review'\n    if any(c.kind == 'fact' for c in conflicts):\n        return pick_best_evidence(conflicts)\n    return planner.replan(conflicts)\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>This pseudocode illustrates the idea of conflict triage: safety conflicts must not be silently absorbed; factual conflicts are handled through evidence ranking; planning conflicts are returned to the planner for reordering. In production environments, you should also record the reasons for decisions to support retrospectives and evaluation.\u003C\u002Fp>\u003Ch2>Practical Steps and Checklist\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Step 1: Define task boundaries.\u003C\u002Fstrong> First clarify which tasks are suitable for automation and which require human confirmation. Do not let multi-agent systems fully automate high-risk operations from the start.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 2: Split minimal roles.\u003C\u002Fstrong> Start with Planner, Executor, and Reviewer, and add Researcher, Critic, or Memory Manager only when truly necessary.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 3: Design message contracts.\u003C\u002Fstrong> Define fields and validation rules for each event type, and avoid passing long free-form text contexts directly.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 4: Establish shared state.\u003C\u002Fstrong> Place task goals, constraints, completed steps, and pending confirmations into a structured state object, rather than letting each agent maintain its own memory.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 5: Add budget controls.\u003C\u002Fstrong> Set maximum rounds, maximum tokens, maximum tool invocations, and maximum duration for each task.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 6: Implement conflict arbitration.\u003C\u002Fstrong> Make clear who has final decision authority and preserve the evidence chain.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 7: Perform trajectory evaluation.\u003C\u002Fstrong> Evaluate not only the final answer, but also whether the plan is reasonable, whether tool calls are necessary, and whether reviews identify issues.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Common Pitfalls and Recommendations\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Too many agents.\u003C\u002Fstrong> The finer the roles, the higher the communication cost. Prioritize testability rather than making the architecture diagram look impressive.\u003C\u002Fli>\u003Cli>\u003Cstrong>Blurry prompt responsibilities.\u003C\u002Fstrong> If one agent can plan, execute, and review, it can easily reinforce its own errors. Responsibilities should be mutually exclusive.\u003C\u002Fli>\u003Cli>\u003Cstrong>Missing trace_id.\u003C\u002Fstrong> Without a trace ID, multi-agent issues are almost impossible to reconstruct. Every message, tool call, and model response should be linked to the same task.\u003C\u002Fli>\u003Cli>\u003Cstrong>Unbounded context concatenation.\u003C\u002Fstrong> Feeding all historical messages to every agent raises costs and dilutes attention. Use summaries, retrieval, and structured state.\u003C\u002Fli>\u003Cli>\u003Cstrong>Ignoring idempotency.\u003C\u002Fstrong> Message retries can cause duplicate tool calls or duplicate data writes. Tool calls should include idempotency keys, and write operations should be safely replayable.\u003C\u002Fli>\u003Cli>\u003Cstrong>No stopping condition.\u003C\u002Fstrong> Two agents asking each other for revisions can fall into an infinite loop. Set maximum rounds and escalation paths.\u003C\u002Fli>\u003Cli>\u003Cstrong>Evaluating only the final result.\u003C\u002Fstrong> Problems in multi-agent systems often appear in intermediate steps. Establish separate metrics for planning, retrieval, execution, and review.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Directions for Further Reading\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Multi-agent frameworks and protocols.\u003C\u002Fstrong> You can follow developments in LangGraph, AutoGen, CrewAI, OpenAI Swarm, MCP, A2A, and related directions. For specific capabilities and interfaces, refer to the official documentation.\u003C\u002Fli>\u003Cli>\u003Cstrong>Distributed system patterns.\u003C\u002Fstrong> Concepts such as event sourcing, Saga, eventual consistency, idempotent consumers, and dead-letter queues are highly relevant to multi-agent communication and recovery mechanisms.\u003C\u002Fli>\u003Cli>\u003Cstrong>Evaluation systems.\u003C\u002Fstrong> Further explore trajectory evaluation, LLM-as-judge, human spot checks, and regression benchmark suites, rather than looking only at single outputs.\u003C\u002Fli>\u003Cli>\u003Cstrong>Security and permissions.\u003C\u002Fstrong> Once multi-agent systems can invoke tools, least privilege, audit logs, sandboxed execution, and human approval become more important than model capability.\u003C\u002Fli>\u003C\u002Ful>\n\u003Csection class=\"portal-sources\" style=\"margin-top:2em;font-size:14px;line-height:1.7;color:#5c5850;\">\u003Cp style=\"margin:0 0 0.5em;font-weight:600;\">References\u003C\u002Fp>\u003Cul style=\"margin:0;padding-left:1.25em;\">\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.hanyuguoxue.com\u002Fzidian\u002Fzi-22810\" rel=\"noopener noreferrer\" target=\"_blank\">多的意思,多的解释,多的拼音,多的部首,多的笔顺-汉语国学\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.duolingo.cn\u002F\" rel=\"noopener noreferrer\" target=\"_blank\">多邻国 - 全球火爆的学习平台\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.chagushici.com\u002Fzidian\u002F%E5%A4%9A\" rel=\"noopener noreferrer\" target=\"_blank\">多_多怎么读_多的意思 - 汉语字典\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzidian.gushici.net\u002F6\u002F591a.html\" rel=\"noopener noreferrer\" target=\"_blank\">多怎么读_多的拼音 - 新华字典\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fapps.microsoft.com\u002Fdetail\u002Fxpfflkcxs7f32z?hl=zh-CN&amp;gl=CN\" rel=\"noopener noreferrer\" target=\"_blank\">多邻国 - Windows官方下载 | 微软应用商店 | Microsoft Store\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>","","https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Flandscape_mountain.jpg","Guanshan Academy","ai","published","2026-09-30T01:34:55.646284+08:00","2026-09-30T01:37:00.863644+08:00","2026-09-30T16:34:32.359912+08:00",[27,30],{"locale":28,"path":29},"zh","\u002Farticles\u002F多-agent-协作落地-角色分工-消息总线与冲突消解的工程实践",{"locale":12,"path":31},"\u002Fen\u002Farticles\u002F多-agent-协作落地-角色分工-消息总线与冲突消解的工程实践","\u003Ch2>Background and Problems\u003C\u002Fh2>\u003Cp>In a single-agent setup, one model often has to handle requirement understanding, retrieval, planning, tool invocation, code generation, and result validation at the same time. As tasks become complex, prompts quickly balloon, and the context becomes filled with irrelevant information, leading to unstable outputs. As a result, many teams begin experimenting with multi-agent collaboration: assigning different roles to handle planning, execution, review, and summarization.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Flandscape_mountain.jpg\" alt=\"观山静思\" loading=\"lazy\" decoding=\"async\">\u003Cfigcaption>观山静思\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp>However, multi-agent systems are not a silver bullet. Without clear role boundaries, the system can degrade into multiple models simply forwarding text to one another; without a unified message bus, call chains can become a tangled web that is difficult to troubleshoot; without conflict resolution mechanisms, different agents may repeatedly argue based on different facts or goals, ultimately increasing latency and cost. This article focuses on three engineering questions: how to divide responsibilities, how to communicate, and how to arbitrate.\u003C\u002Fp>\u003Ch2>Core Concepts\u003C\u002Fh2>\u003Ch3>Role Division: Define Contracts Before Writing Prompts\u003C\u002Fh3>\u003Cp>The first step in a multi-agent system is not choosing models, but splitting responsibilities. Common roles include: \u003Cstrong>Planner\u003C\u002Fstrong>, responsible for task decomposition and priority ordering; \u003Cstrong>Researcher\u003C\u002Fstrong>, responsible for retrieval and fact organization; \u003Cstrong>Executor\u003C\u002Fstrong>, responsible for invoking tools or executing code; \u003Cstrong>Reviewer\u003C\u002Fstrong>, responsible for quality checks; \u003Cstrong>Arbitrator\u003C\u002Fstrong>, responsible for conflict arbitration; and \u003Cstrong>Memory Manager\u003C\u002Fstrong>, responsible for long-term memory and summary compression.\u003C\u002Fp>\u003Cp>Each role should have a clear contract: what the input is, what the output is, which tools can be invoked, what it must not do, and when it should finish. For example, the Reviewer should not directly rewrite the final solution, but should output structured review comments; the Researcher should not make business decisions, but should provide evidence with sources. This reduces role overreach and makes later unit evaluation easier.\u003C\u002Fp>\u003Ch3>Message Bus: Turn Conversations into Traceable Events\u003C\u002Fh3>\u003Cp>Many early multi-agent implementations use point-to-point calls: A calls B, B calls C, and C calls back A. This approach is hard to debug and can easily create loops. A more robust approach is to introduce a message bus and abstract interactions between agents as events. Events may include \u003Ccode>task.created\u003C\u002Fcode>, \u003Ccode>plan.proposed\u003C\u002Fcode>, \u003Ccode>tool.result.received\u003C\u002Fcode>, \u003Ccode>review.failed\u003C\u002Fcode>, \u003Ccode>conflict.detected\u003C\u002Fcode>, and so on.\u003C\u002Fp>\u003Cp>A message envelope should include at least a trace ID, sender, recipient or topic, message type, payload, schema version, timestamp, and expiration time. The trace ID is used for end-to-end replay, the schema version supports compatible evolution, and the expiration time prevents stale messages from triggering side effects.\u003C\u002Fp>\u003Cpre>\u003Ccode>from dataclasses import dataclass\n\n@dataclass\nclass AgentMessage:\n    msg_id: str\n    trace_id: str\n    sender: str\n    topic: str\n    payload: dict\n    schema_version: str = 'v1'\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>This code defines a unified message structure. In real projects, you can also add signatures, priority, retry counts, and idempotency keys. The message bus itself can be implemented with an in-memory queue, Redis Stream, Kafka, or a cloud messaging service. The key is not the component itself, but whether the message contract is stable.\u003C\u002Fp>\u003Ch3>Conflict Resolution: Move from Debate to Adjudication\u003C\u002Fh3>\u003Cp>Multi-agent conflicts generally fall into four categories: fact conflicts, plan conflicts, resource conflicts, and policy conflicts. Fact conflicts occur when different agents present different facts; plan conflicts arise from inconsistent execution order; resource conflicts happen when multiple tasks compete for the same tool, file, budget, or external interface; policy conflicts stem from different risk preferences—for example, one agent prefers automatic execution while another requires human confirmation.\u003C\u002Fp>\u003Cp>When resolving conflicts, it is not advisable to let multiple agents engage in endless dialogue. A better approach is to define adjudication rules: safety issues should be escalated to humans first; for factual issues, prefer results with evidence, more recent timestamps, and more credible sources; planning issues should be decided by the Planner or Arbitrator according to the objective function; resource issues should be controlled through locks, queues, budgets, and idempotency keys.\u003C\u002Fp>\u003Cpre>\u003Ccode>def resolve_conflict(conflicts):\n    if any(c.kind == 'safety' for c in conflicts):\n        return 'human_review'\n    if any(c.kind == 'fact' for c in conflicts):\n        return pick_best_evidence(conflicts)\n    return planner.replan(conflicts)\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>This pseudocode illustrates the idea of conflict triage: safety conflicts must not be silently absorbed; factual conflicts are handled through evidence ranking; planning conflicts are returned to the planner for reordering. In production environments, you should also record the reasons for decisions to support retrospectives and evaluation.\u003C\u002Fp>\u003Ch2>Practical Steps and Checklist\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Step 1: Define task boundaries.\u003C\u002Fstrong> First clarify which tasks are suitable for automation and which require human confirmation. Do not let multi-agent systems fully automate high-risk operations from the start.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 2: Split minimal roles.\u003C\u002Fstrong> Start with Planner, Executor, and Reviewer, and add Researcher, Critic, or Memory Manager only when truly necessary.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 3: Design message contracts.\u003C\u002Fstrong> Define fields and validation rules for each event type, and avoid passing long free-form text contexts directly.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 4: Establish shared state.\u003C\u002Fstrong> Place task goals, constraints, completed steps, and pending confirmations into a structured state object, rather than letting each agent maintain its own memory.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 5: Add budget controls.\u003C\u002Fstrong> Set maximum rounds, maximum tokens, maximum tool invocations, and maximum duration for each task.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 6: Implement conflict arbitration.\u003C\u002Fstrong> Make clear who has final decision authority and preserve the evidence chain.\u003C\u002Fli>\u003Cli>\u003Cstrong>Step 7: Perform trajectory evaluation.\u003C\u002Fstrong> Evaluate not only the final answer, but also whether the plan is reasonable, whether tool calls are necessary, and whether reviews identify issues.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Common Pitfalls and Recommendations\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Too many agents.\u003C\u002Fstrong> The finer the roles, the higher the communication cost. Prioritize testability rather than making the architecture diagram look impressive.\u003C\u002Fli>\u003Cli>\u003Cstrong>Blurry prompt responsibilities.\u003C\u002Fstrong> If one agent can plan, execute, and review, it can easily reinforce its own errors. Responsibilities should be mutually exclusive.\u003C\u002Fli>\u003Cli>\u003Cstrong>Missing trace_id.\u003C\u002Fstrong> Without a trace ID, multi-agent issues are almost impossible to reconstruct. Every message, tool call, and model response should be linked to the same task.\u003C\u002Fli>\u003Cli>\u003Cstrong>Unbounded context concatenation.\u003C\u002Fstrong> Feeding all historical messages to every agent raises costs and dilutes attention. Use summaries, retrieval, and structured state.\u003C\u002Fli>\u003Cli>\u003Cstrong>Ignoring idempotency.\u003C\u002Fstrong> Message retries can cause duplicate tool calls or duplicate data writes. Tool calls should include idempotency keys, and write operations should be safely replayable.\u003C\u002Fli>\u003Cli>\u003Cstrong>No stopping condition.\u003C\u002Fstrong> Two agents asking each other for revisions can fall into an infinite loop. Set maximum rounds and escalation paths.\u003C\u002Fli>\u003Cli>\u003Cstrong>Evaluating only the final result.\u003C\u002Fstrong> Problems in multi-agent systems often appear in intermediate steps. Establish separate metrics for planning, retrieval, execution, and review.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Directions for Further Reading\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Multi-agent frameworks and protocols.\u003C\u002Fstrong> You can follow developments in LangGraph, AutoGen, CrewAI, OpenAI Swarm, MCP, A2A, and related directions. For specific capabilities and interfaces, refer to the official documentation.\u003C\u002Fli>\u003Cli>\u003Cstrong>Distributed system patterns.\u003C\u002Fstrong> Concepts such as event sourcing, Saga, eventual consistency, idempotent consumers, and dead-letter queues are highly relevant to multi-agent communication and recovery mechanisms.\u003C\u002Fli>\u003Cli>\u003Cstrong>Evaluation systems.\u003C\u002Fstrong> Further explore trajectory evaluation, LLM-as-judge, human spot checks, and regression benchmark suites, rather than looking only at single outputs.\u003C\u002Fli>\u003Cli>\u003Cstrong>Security and permissions.\u003C\u002Fstrong> Once multi-agent systems can invoke tools, least privilege, audit logs, sandboxed execution, and human approval become more important than model capability.\u003C\u002Fli>\u003C\u002Ful>\n\u003Csection class=\"portal-sources\" style=\"margin-top:2em;font-size:14px;line-height:1.7;color:#5c5850;\">\u003Cp style=\"margin:0 0 0.5em;font-weight:600;\">References\u003C\u002Fp>\u003Cul style=\"margin:0;padding-left:1.25em;\">\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.hanyuguoxue.com\u002Fzidian\u002Fzi-22810\" rel=\"noopener noreferrer\">多的意思,多的解释,多的拼音,多的部首,多的笔顺-汉语国学\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.duolingo.cn\u002F\" rel=\"noopener noreferrer\">多邻国 - 全球火爆的学习平台\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.chagushici.com\u002Fzidian\u002F%E5%A4%9A\" rel=\"noopener noreferrer\">多_多怎么读_多的意思 - 汉语字典\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzidian.gushici.net\u002F6\u002F591a.html\" rel=\"noopener noreferrer\">多怎么读_多的拼音 - 新华字典\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fapps.microsoft.com\u002Fdetail\u002Fxpfflkcxs7f32z?hl=zh-CN&amp;gl=CN\" rel=\"noopener noreferrer\">多邻国 - Windows官方下载 | 微软应用商店 | Microsoft Store\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",["Island",34],{"key":35,"result":36},"AppImage_U2SWsmBV1mDznAcDcnH397RslvGSkZeLYKgmLBsA",{"head":37},{"link":38,"style":39},[],[],["Island",41],{"key":42,"result":43},"PortalBreadcrumb_yrk2AtPrXsKFZJhIyeTD0R4a2WAywceLd1kQYcFeOIg",{"head":44},{"link":45,"style":46},[],[],[48,57,66],{"id":49,"locale":12,"slug":50,"type":14,"title":51,"summary":52,"content_html":53,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"published_at":54,"created_at":55,"updated_at":56},4654,"提示词工程实践-系统提示-few-shot-与结构化输出怎么写更稳","Prompt Engineering in Practice: How to Write More Reliable System Prompts, Few-shot Examples, and Structured Outputs","From an engineering perspective, this article breaks down prompt design, focusing on the applicable scenarios for system prompts, Few-shot examples, and structured output. It provides a practical checklist, common pitfalls, and testing recommendations to help developers improve the stability and parseability of large model outputs.","\u003Ch2>Background and Problem: A Prompt Is Not a Chat Message, but Interface Design\u003C\u002Fh2>\u003Cp>When many teams integrate large models for the first time, they often encounter a gap: the model appears highly capable, but behaves inconsistently once applied to real business scenarios. Answers may be too broad, formats may be unparsable, or the model may drift under edge conditions. On the surface, this seems like the model is not smart enough. In practice, the more common cause is that the prompt does not clearly specify the task objective, input constraints, output format, and failure handling.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Flandscape_mountain.jpg\" alt=\"观山静思\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>观山静思\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\u003Cp>From an engineering perspective, a prompt is not merely a single sentence; it is a lightweight protocol for the model to execute. Like API documentation, it should be explicit about: who the role is, what the input is, what the output should look like, what must not be done, and how to handle uncertainty. System prompts, Few-shot examples, and structured output are the three most common levers in this protocol.\u003C\u002Fp>\u003Ch2>Core Concepts: What Problems Do These Three Techniques Solve?\u003C\u002Fh2>\u003Ch3>1. System Prompts: Stabilize Role, Style, and Boundaries\u003C\u002Fh3>\u003Cp>A system prompt is typically used to define the model's overall behavioral framework. Compared with user inputs in each turn, it is better suited for long-lived rules, such as role definition, tone and style, safety boundaries, output preferences, and tool-calling constraints. A common mistake is to write the system prompt as a vague slogan, such as “Be professional, accurate, and concise.” Such descriptions are not entirely useless, but they lack executable criteria. A more reliable approach is to break requirements down into actions the model can follow.\u003C\u002Fp>\u003Cp>For example, instead of writing “Please answer professionally,” write: “You are a technical documentation editor. Before answering, determine whether the user's question belongs to frontend engineering. If information is insufficient, ask no more than three clarification questions. Do not fabricate dependency versions that were not provided.” The advantage of this kind of prompt is that it reduces the model's freedom to improvise and makes later regression testing easier.\u003C\u002Fp>\u003Ch3>2. Few-shot: Use Examples to Calibrate Format and Judgment Criteria\u003C\u002Fh3>\u003Cp>Few-shot is not simply piling up examples. It uses examples to tell the model what kind of output should be produced for what kind of input. It is especially suitable for tasks such as classification, extraction, rewriting, review, and intent recognition. Examples provide two types of value: format calibration and boundary calibration. Format calibration tells the model the fields, order, length, and tone. Boundary calibration tells the model how to classify ambiguous input, or whether to return a null value.\u003C\u002Fp>\u003Cp>In real projects, Few-shot examples do not need to be numerous; they need to be representative. Prioritize three types of samples: standard samples, boundary samples, and samples that are easily misjudged. Standard samples establish the basic format. Boundary samples define ambiguous areas. Misjudgment samples correct common errors. If the rules among examples conflict, the model can easily learn unstable patterns.\u003C\u002Fp>\u003Ch3>3. Structured Output: Make Results Parsable, Validatable, and Storable\u003C\u002Fh3>\u003Cp>When prompts enter production systems, outputs often cannot be only natural language; they need to be JSON, tables, enumerated fields, or fixed templates. The core of structured output is not simply telling the model to “output JSON.” Instead, the schema, field meanings, allowed value ranges, and default strategies must be clearly specified. Many failures occur not because the model cannot produce JSON, but because the prompt does not state whether to fill a missing field with null, an empty string, or omit the field.\u003C\u002Fp>\u003Cp>According to public documentation, mainstream large models generally support reducing format errors through prompt constraints or official structured-output capabilities. However, different models vary in their support for JSON Schema, function calling, and response formats; always refer to the official documentation. From an engineering standpoint, always keep a layer of server-side validation and do not treat model output directly as trusted data.\u003C\u002Fp>\u003Ch2>Practical Steps: A Deployable Prompt Design Checklist\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Clarify the task type:\u003C\u002Fstrong> First determine whether the task is open-ended generation, information extraction, classification\u002Fdecision-making, or code generation. Different types suit different prompting strategies.\u003C\u002Fli>\u003Cli>\u003Cstrong>Specify input sources:\u003C\u002Fstrong> Tell the model which content comes from user input, which comes from system context, and which is untrusted. When necessary, use tags to separate data from instructions.\u003C\u002Fli>\u003Cli>\u003Cstrong>Define the output contract:\u003C\u002Fstrong> Specify field names, types, lengths, enumerated values, whether empty values are allowed, and what to return on failure.\u003C\u002Fli>\u003Cli>\u003Cstrong>Add the minimum necessary examples:\u003C\u002Fstrong> Start with 2 to 4 high-quality examples covering normal, boundary, and abnormal scenarios.\u003C\u002Fli>\u003Cli>\u003Cstrong>Set refusal and fallback strategies:\u003C\u002Fstrong> For example, “If the information cannot be extracted from the given materials, return empty: true. Do not guess.”\u003C\u002Fli>\u003Cli>\u003Cstrong>Run regression tests:\u003C\u002Fstrong> Prepare a fixed test set. Every time the prompt changes, compare output differences to avoid local optimizations causing overall regressions.\u003C\u002Fli>\u003C\u002Ful>\u003Cp>Below is a structured prompt example for ticket classification. It places the system role, input area, Few-shot examples, and JSON output requirements into the same template, making programmatic assembly and future maintenance easier.\u003C\u002Fp>\u003Cpre>\u003Ccode>system:\nYou are a ticket classification assistant. Classify tickets only based on the ticket content provided by the user. Do not infer information that is not given.\nThe output must be valid JSON with the following fields: category, urgency, reason.\ncategory must be one of account, payment, bug, other.\nurgency must be one of low, medium, high.\nIf the information is insufficient, set category to other and clearly explain the reason.\n\nExample 1:\nInput: I cannot log in. It says the password is incorrect, but I am not receiving the reset email.\nOutput: {&quot;category&quot;:&quot;account&quot;,&quot;urgency&quot;:&quot;high&quot;,&quot;reason&quot;:&quot;Login is blocked and the reset email is not working&quot;}\n\nExample 2:\nInput: A button on the page does not respond after being clicked, and the console shows a 500 error.\nOutput: {&quot;category&quot;:&quot;bug&quot;,&quot;urgency&quot;:&quot;medium&quot;,&quot;reason&quot;:&quot;A feature is unavailable and a server-side error is present&quot;}\n\nUser input:\n{{ticket_content}}\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>The key point of this example is not the JSON itself, but that it turns “do not infer,” “enumerated values,” and “what to do when information is insufficient” into executable rules. For programs, this kind of output is easier to validate and route.\u003C\u002Fp>\u003Ch2>Common Pitfalls and Recommendations\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Writing only the goal, not the constraints:\u003C\u002Fstrong> For example, saying “Summarize this for me” without specifying length, target audience, or whether risk items should be preserved. Include acceptance criteria in the prompt.\u003C\u002Fli>\u003Cli>\u003Cstrong>Too many examples that contradict one another:\u003C\u002Fstrong> Too many examples consume context and amplify noise. Keep samples that can actually change the model's judgment.\u003C\u002Fli>\u003Cli>\u003Cstrong>Mixing untrusted input directly into instructions:\u003C\u002Fstrong> User input may contain manipulation or injection content. Wrap external text in explicit tags and declare that instructions inside it are invalid.\u003C\u002Fli>\u003Cli>\u003Cstrong>Over-relying on the model's self-discipline:\u003C\u002Fstrong> Production environments should add JSON Schema validation, field sanitization, and retry mechanisms. Model output should be treated only as candidate results.\u003C\u002Fli>\u003Cli>\u003Cstrong>Ignoring model differences:\u003C\u002Fstrong> The same prompt may perform noticeably differently across models. Re-evaluate before switching models; do not assume direct transferability.\u003C\u002Fli>\u003Cli>\u003Cstrong>Hard-coding prompts once and for all:\u003C\u002Fstrong> Prompts should be version-controlled like code, with change reasons, test sets, and performance metrics recorded.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Directions for Further Reading\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Context engineering:\u003C\u002Fstrong> Explore how to organize system prompts, retrieved materials, conversation history, and tool results within limited tokens.\u003C\u002Fli>\u003Cli>\u003Cstrong>Structured output and function calling:\u003C\u002Fstrong> Follow official capabilities for JSON output, tool calling, and response formats across models.\u003C\u002Fli>\u003Cli>\u003Cstrong>Prompt evaluation:\u003C\u002Fstrong> Build golden sample sets, automatic scorers, and human spot-check workflows.\u003C\u002Fli>\u003Cli>\u003Cstrong>Security and injection protection:\u003C\u002Fstrong> Understand risks such as prompt injection, privilege escalation, and sensitive information leakage, along with mitigation strategies.\u003C\u002Fli>\u003Cli>\u003Cstrong>RAG and knowledge boundaries:\u003C\u002Fstrong> When tasks depend on private knowledge, combine retrieval, citations, and confidence control.\u003C\u002Fli>\u003C\u002Ful>\n\u003Csection class=\"portal-sources\" style=\"margin-top:2em;font-size:14px;line-height:1.7;color:#5c5850;\">\u003Cp style=\"margin:0 0 0.5em;font-weight:600;\">References\u003C\u002Fp>\u003Cul style=\"margin:0;padding-left:1.25em;\">\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.runoob.com\u002Fai-agent\u002Fprompt-engineering.html\" rel=\"noopener noreferrer\" target=\"_blank\">提示词工程（Prompt Engineering） | 菜鸟教程\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>","2026-09-30T15:52:16.289011+08:00","2026-09-30T15:55:41.641326+08:00","2026-09-30T16:34:32.304768+08:00",{"id":58,"locale":12,"slug":59,"type":14,"title":60,"summary":61,"content_html":62,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"published_at":63,"created_at":64,"updated_at":65},4628,"私有化部署大模型-硬件选型-容器编排与监控告警实践","Private Deployment of Large Models: Hardware Selection, Container Orchestration, and Monitoring and Alerting Practices","This article outlines the key path for privately deploying large models from an engineering implementation perspective: how to choose hardware based on GPU memory, concurrency, and cost; how to use containers and orchestration systems to deliver stable inference services; and how to establish a monitoring and alerting system for GPU and business metrics.","\u003Ch2>Background and Problems\u003C\u002Fh2>\u003Cp>More and more teams want to place large models in their own data centers or dedicated clouds for reasons such as data compliance, intranet integration, controllable costs, and customized inference pipelines. However, private deployment is not as simple as downloading weights and starting a service. In real projects, the most common failure points are often not the model itself, but underestimation of hardware resources, chaotic container environment dependencies, unstable GPU scheduling, and lack of observability after launch.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Flandscape_mountain.jpg\" alt=\"观山静思\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>观山静思\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp>From an engineering perspective, a privately deployed large-model service involves at least three layers: the underlying compute layer, including GPUs, CPUs, memory, storage, and networking; the middle runtime layer, including drivers, CUDA, container runtimes, and inference frameworks; and the upper service layer, including API gateways, authentication, rate limiting, monitoring, and alerting. A weakness in any layer can lead to timeouts, GPU memory overflows, or node drift.\u003C\u002Fp>\u003Ch2>Core Concepts\u003C\u002Fh2>\u003Ch3>1. GPU Memory Is Not the Only Metric\u003C\u002Fh3>\u003Cp>Many people's first reaction is to look at GPU memory. GPU memory determines whether the model weights can fit, but the inference experience is also affected by memory bandwidth, compute unit performance, PCIe or NVLink interconnects, and CPU preprocessing capability. According to public materials, in long-context scenarios, the KV cache can significantly increase GPU memory usage.\u003C\u002Fp>\u003Cp>A rough estimate: the GPU memory required for weights is approximately the number of parameters multiplied by the number of bytes per parameter. For example, a 7B model in FP16 requires about 14GB for weights. Adding KV cache, runtime buffers, and fragmentation, 24GB of GPU memory provides a more comfortable margin. If INT8 or INT4 quantization is used, weight memory decreases, but quality loss and framework compatibility should be evaluated. Please refer to official documentation.\u003C\u002Fp>\u003Ch3>2. Inference Frameworks and Service Deployment\u003C\u002Fh3>\u003Cp>In production environments, it is not recommended to load models directly with scripts. Instead, choose mature inference serving frameworks, such as vLLM, Text Generation Inference, Ollama, or vendor inference services. They usually provide batching, continuous batching, streaming output, OpenAI-compatible interfaces, and basic metrics. Interface standardization is more important than peak performance, because gateways, monitoring, and clients all rely on stable protocols.\u003C\u002Fp>\u003Ch3>3. The Value of Container Orchestration\u003C\u002Fh3>\u003Cp>Containerization isolates dependencies, and orchestration systems can handle replicas, health checks, resource limits, and node scheduling. For GPU workloads, NVIDIA drivers, Container Toolkit, GPU Operator, or device plugins must also work together. Kubernetes is suitable for multi-team scenarios; for single-machine or small-scale environments, Docker Compose can reduce complexity.\u003C\u002Fp>\u003Ch3>4. Monitoring and Alerting Should Cover Both GPUs and Business Metrics\u003C\u002Fh3>\u003Cp>Looking only at GPU utilization is not enough. High utilization does not necessarily mean the service is healthy; it may simply indicate request backlog. Low utilization does not necessarily indicate a failure; it may be due to batching policies or client-side rate limiting. A more reasonable approach is to collect hardware, runtime, and business metrics at the same time: GPU memory, temperature, power consumption, request latency, time to first token, generation speed, queue length, error rate, and replica restart count.\u003C\u002Fp>\u003Ch2>Practical Steps and Checklist\u003C\u002Fh2>\u003Ch3>Step 1: Define the Business Profile\u003C\u002Fh3>\u003Cul>\u003Cli>Model scale: 7B, 13B, 30B, or larger.\u003C\u002Fli>\u003Cli>Context length: short conversations, long documents, or code completion.\u003C\u002Fli>\u003Cli>Concurrency targets: peak QPS, maximum tokens per request, and whether queuing is allowed.\u003C\u002Fli>\u003Cli>Security boundaries: whether it must be fully offline, and whether auditing and permission control are required.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3>Step 2: Hardware Selection\u003C\u002Fh3>\u003Cul>\u003Cli>GPU: prioritize GPU memory capacity, memory bandwidth, and driver ecosystem. For long-context workloads, reserve at least 30% GPU memory headroom.\u003C\u002Fli>\u003Cli>CPU: there is no need to blindly pursue top-tier models, but ensure that tokenization, data loading, logging, and sidecars do not become bottlenecks.\u003C\u002Fli>\u003Cli>Memory: it is generally recommended to have at least twice the size of the model weights, while accounting for system and container overhead.\u003C\u002Fli>\u003Cli>Storage: model files are large, so high-speed SSDs are recommended; separate the image registry from the model cache.\u003C\u002Fli>\u003Cli>Networking: for multi-machine deployments, pay attention to high-speed Ethernet or RDMA; for single-machine deployments, pay attention to PCIe topology and cooling.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3>Step 3: Prepare the Runtime Environment\u003C\u002Fh3>\u003Cp>It is recommended to pin the versions of drivers, CUDA, container runtime, and inference framework. Do not upgrade drivers on production nodes on the fly. Check items include:\u003C\u002Fp>\u003Cul>\u003Cli>nvidia-smi can stably detect GPUs.\u003C\u002Fli>\u003Cli>Device nodes are accessible inside containers.\u003C\u002Fli>\u003Cli>Model files are mounted read-only or cached from object storage.\u003C\u002Fli>\u003Cli>Service ports, health check paths, and timeout values are configured.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3>Step 4: Containerize the Inference Service\u003C\u002Fh3>\u003Cp>Below is a simplified docker run example for starting an inference service compatible with the OpenAI interface. In actual deployments, replace the image, model path, and resource parameters, and refer to official documentation.\u003C\u002Fp>\u003Cpre>\u003Ccode>docker run --gpus all -v \u002Fdata\u002Fmodels:\u002Fmodels -p 8000:8000 vllm\u002Fvllm-openai:latest --model \u002Fmodels\u002Fyour-model --max-model-len 8192\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>This example mounts the model directory into the container and exposes an HTTP interface. In production, also add restart policies, resource limits, log rotation, and health checks. If using Kubernetes, you can manage replicas and scheduling through Deployment, Service, HPA, and node affinity.\u003C\u002Fp>\u003Ch3>Step 5: Integrate Monitoring and Alerting\u003C\u002Fh3>\u003Cp>A common approach is to scrape metrics with Prometheus, visualize them in Grafana, and send alerts with Alertmanager. GPU metrics can be exposed through DCGM Exporter or vendor exporters. Business metrics can be provided by the inference framework or aggregated at the gateway layer. It is recommended to configure at least the following alerts: GPU memory continuously approaching the limit, abnormal GPU temperature, rising request error rate, time to first token exceeding a threshold, and frequent replica restarts.\u003C\u002Fp>\u003Ch3>Pre-launch Checklist\u003C\u002Fh3>\u003Cul>\u003Cli>Load testing: cover short requests, long requests, concurrency spikes, and streaming output.\u003C\u002Fli>\u003Cli>Failure drills: simulate GPU device loss, container OOM, and node restarts.\u003C\u002Fli>\u003Cli>Rate limiting: set maximum concurrency and request body size at the gateway layer.\u003C\u002Fli>\u003Cli>Logging: retain request IDs, durations, token counts, and error codes, but do not log sensitive prompts.\u003C\u002Fli>\u003Cli>Capacity: reserve growth headroom for GPU memory, disk, and memory.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Common Pitfalls and Recommendations\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Estimating GPU memory based only on weights:\u003C\u002Fstrong> Ignoring KV cache and batching can cause direct OOM in long-context or high-concurrency scenarios. It is recommended to run load tests with real business samples.\u003C\u002Fli>\u003Cli>\u003Cstrong>Blindly pursuing multiple GPUs:\u003C\u002Fstrong> If the model can run on a single GPU, multi-GPU parallelism may not improve throughput and can increase communication and scheduling complexity. First clarify whether you need tensor parallelism, pipeline parallelism, or replica scaling.\u003C\u002Fli>\u003Cli>\u003Cstrong>Oversized container images and version drift:\u003C\u002Fstrong> Separate model files from runtime images, use fixed version tags, and avoid using latest directly in production.\u003C\u002Fli>\u003Cli>\u003Cstrong>GPU scheduling affinity issues in Kubernetes:\u003C\u002Fstrong> Ensure that device plugins, node labels, taints, and resource requests are consistent; otherwise Pods may remain Pending for a long time.\u003C\u002Fli>\u003Cli>\u003Cstrong>Alerting only on GPU utilization:\u003C\u002Fstrong> Combine time to first token, queue length, and error rate. Otherwise, metrics may look normal while user experience is poor.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Directions for Further Reading\u003C\u002Fh2>\u003Cul>\u003Cli>Batching, PagedAttention, continuous batching, and quantization strategies in inference frameworks.\u003C\u002Fli>\u003Cli>Kubernetes GPU Operator, Device Plugin, and node affinity design.\u003C\u002Fli>\u003Cli>Metric modeling for AI services with DCGM, Prometheus, and OpenTelemetry.\u003C\u002Fli>\u003Cli>Model gateway design: authentication, quotas, caching, content auditing, and multi-model routing.\u003C\u002Fli>\u003C\u002Ful>\u003Cp>Overall, privately deploying large models is a systems engineering effort. Hardware determines the upper bound, orchestration determines stability, and monitoring determines maintainability. Teams should first run through the full pipeline in a small-scale, reproducible environment, then gradually increase concurrency and model scale. Avoid piling up expensive hardware from the start while losing points on engineering and operations details.\u003C\u002Fp>\n\u003Csection class=\"portal-sources\" style=\"margin-top:2em;font-size:14px;line-height:1.7;color:#5c5850;\">\u003Cp style=\"margin:0 0 0.5em;font-weight:600;\">References\u003C\u002Fp>\u003Cul style=\"margin:0;padding-left:1.25em;\">\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.zhihu.com\u002F\" rel=\"noopener noreferrer\" target=\"_blank\">知乎 - 有问题，就会有答案\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.zhihu.com\u002Ftardis\u002Fzm\u002Fart\u002F280070583\" rel=\"noopener noreferrer\" target=\"_blank\">2026年 9月 CPU天梯图（更新250\u002F270K Plus&amp;amp;9850X3D）\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.zhihu.com\u002Forg\u002Fyan-yan-gu-shi\" rel=\"noopener noreferrer\" target=\"_blank\">盐言故事 - 知乎\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F1961517406183757165\" rel=\"noopener noreferrer\" target=\"_blank\">gpu是显卡吗？gpu和cpu的区别对比 - 知乎\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fblog.csdn.net\u002Fm0_69925214\u002Farticle\u002Fdetails\u002F157173376\" rel=\"noopener noreferrer\" target=\"_blank\">一次性厘清 CPU、显卡、GPU到底是什么？之间的关系？\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>","2026-09-30T07:39:02.852675+08:00","2026-09-30T07:41:41.029186+08:00","2026-09-30T16:34:32.400708+08:00",{"id":67,"locale":12,"slug":68,"type":14,"title":69,"summary":70,"content_html":71,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"published_at":72,"created_at":73,"updated_at":74},4620,"语音与大模型-asr-tts-与实时对话链路设计实践","Voice and Large Models: Practical Design of ASR, TTS, and Real-Time Conversation Pipelines","From voice input to model response to speech playback, a real-time conversation system is not simply ASR, LLM, and TTS stitched together. This article outlines pipeline modules, latency budgets, streaming chunking, interruption control, and an engineering checklist to help developers build more natural voice AI products.","\u003Ch2>Background and Challenges\u003C\u002Fh2>\n\u003Cp>Voice interaction is evolving from a communication tool into an AI entry point. On one hand, public information from KOOK and YY Voice shows that voice rooms, cross-device connectivity, and low-latency co-hosting are already mature requirements. On the other hand, online TTS tools such as NowVoice and Airvoz are rapidly popularizing text-to-speech for dubbing, audio content, and accessibility scenarios. For developers, the real challenge is not building a standalone ASR or TTS demo, but combining voice input, large-model inference, and voice output into a stable, low-latency, interruptible real-time conversation pipeline.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Flandscape_mountain.jpg\" alt=\"观山静思\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>观山静思\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp>Common pain points include: the system responds long after the user finishes speaking; the AI is mistakenly interrupted by background voices; ASR errors cause the model to answer irrelevantly; the TTS voice sounds natural but the first audio packet is too slow; and network jitter causes audio to break up. These problems must be addressed at the pipeline design level.\u003C\u002Fp>\n\u003Ch2>Core Concepts: A Real-Time Voice Conversation Pipeline\u003C\u002Fh2>\n\u003Ch3>1. Audio Capture and Preprocessing\u003C\u002Fh3>\n\u003Cp>After the client captures PCM or Opus audio, it usually goes through echo cancellation, noise suppression, automatic gain control, and silence detection. If the AI is playing audio and the user suddenly speaks, echo cancellation directly affects whether ASR mistakenly recognizes the AI's own voice as user input.\u003C\u002Fp>\n\u003Ch3>2. VAD and Endpoint Detection\u003C\u002Fh3>\n\u003Cp>VAD determines whether the user is speaking, while endpoint detection decides when to send the speech segment to ASR. A real-time conversation should not wait until the user has paused for a long time before submitting, nor should it cut off too early during a normal pause.\u003C\u002Fp>\n\u003Ch3>3. Streaming ASR\u003C\u002Fh3>\n\u003Cp>ASR converts speech into text. Real-time pipelines usually upload audio via WebSocket or gRPC streaming and continuously receive intermediate and final results. In engineering, pay attention to sample rate, encoding format, hotwords, timestamps, and segmentation strategy.\u003C\u002Fp>\n\u003Ch3>4. LLM Conversation Orchestration\u003C\u002Fh3>\n\u003Cp>The LLM generates responses based on context. Voice scenarios are better suited to streaming output: first receive the initial token, then generate sentences progressively. If tool calls, knowledge retrieval, or safety review are involved, timeout and fallback strategies should be clearly defined in the orchestration layer.\u003C\u002Fp>\n\u003Ch3>5. Streaming TTS Synthesis\u003C\u002Fh3>\n\u003Cp>TTS converts text into speech. Public materials for products such as NowVoice and Airvoz often highlight capabilities like multiple voices, natural tone, and speed adjustment. However, in real-time conversations, first-packet latency, streaming return, text cleaning, and long-sentence splitting deserve even more attention.\u003C\u002Fp>\n\u003Ch3>6. Playback, Interruption, and State Machine\u003C\u002Fh3>\n\u003Cp>The conversation system needs clear states: idle, listening, thinking, speaking, and interrupted. When the user makes a valid interruption, playback should stop immediately, unplayed audio should be canceled, and the system should return to the listening state.\u003C\u002Fp>\n\u003Ch2>Practical Steps and Checklist\u003C\u002Fh2>\n\u003Ch3>Step 1: Define the Latency Budget\u003C\u002Fh3>\n\u003Cp>For example, set the end-to-end target to 800 milliseconds to 1.5 seconds, and break it down across capture, network, ASR, LLM first token, TTS first packet, playback buffering, and other stages. Specific values vary by provider and deployment method; refer to official documentation.\u003C\u002Fp>\n\u003Ch3>Step 2: Make Everything Streaming\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>\u003Cstrong>ASR\u003C\u002Fstrong>: Support recognition while speaking and return incremental text.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>LLM\u003C\u002Fstrong>: Enable streaming generation to avoid waiting for a complete response.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>TTS\u003C\u002Fstrong>: Support synthesis by sentence or segment, generating and playing at the same time.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>Step 3: Design a TTS Chunking Strategy\u003C\u002Fh3>\n\u003Cp>Do not wait until the entire response is complete before synthesizing. Split by punctuation, semantics, and minimum length to avoid fragments that are too small and sound mechanical.\u003C\u002Fp>\n\u003Cpre>\u003Ccode>def split_for_tts(text):\n    parts = []\n    buf = ''\n    for ch in text:\n        buf += ch\n        if ch in '.!?;,':\n            if len(buf) >= 8:\n                parts.append(buf)\n                buf = ''\n    if buf:\n        parts.append(buf)\n    return parts\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>This example splits LLM output into fragments suitable for TTS. A production system can also consider complete semantics, maximum character count, and pause control.\u003C\u002Fp>\n\u003Ch3>Step 4: Implement Reliable Interruption Control\u003C\u002Fh3>\n\u003Cp>After detecting valid user speech, stop playback and cancel queued TTS tasks. To reduce false interruptions, add checks such as energy threshold, VAD confidence, and minimum speech duration.\u003C\u002Fp>\n\u003Ch3>Step 5: Establish Monitoring Metrics\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>Time from when the user stops speaking to the final ASR text.\u003C\u002Fli>\n\u003Cli>LLM first-token latency.\u003C\u002Fli>\n\u003Cli>TTS first-packet audio latency.\u003C\u002Fli>\n\u003Cli>End-to-end response latency.\u003C\u002Fli>\n\u003Cli>False interruption rate, successful interruption rate, and reconnection rate.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Common Pitfalls and Recommendations\u003C\u002Fh2>\n\u003Ch3>Pitfall 1: Focusing Only on Voice Quality While Ignoring Latency\u003C\u002Fh3>\n\u003Cp>High-quality voices are suitable for offline dubbing, but real-time conversations care more about whether the system responds promptly. Start with a low-latency voice, then improve audio quality based on user feedback.\u003C\u002Fp>\n\u003Ch3>Pitfall 2: Sending the Full Model Response to TTS at Once\u003C\u002Fh3>\n\u003Cp>This significantly increases waiting time. The right approach is: the model generates one sentence, TTS synthesizes one sentence, and the player buffers one sentence.\u003C\u002Fp>\n\u003Ch3>Pitfall 3: Overlooking Voice-Oriented Prompts\u003C\u002Fh3>\n\u003Cp>Voice responses should be short and conversational, avoiding long lists, code blocks, and complex tables. You can constrain output in the system prompt.\u003C\u002Fp>\n\u003Cpre>\u003Ccode>SYSTEM_PROMPT = (\n    'You are a voice assistant. Keep answers conversational and brief, '\n    'avoid lists and code. If unsure, say you need to confirm.'\n)\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>This example helps the LLM generate responses that are more suitable for spoken playback.\u003C\u002Fp>\n\u003Ch3>Pitfall 4: Missing Text Normalization\u003C\u002Fh3>\n\u003Cp>Numbers, dates, English abbreviations, and emojis must be converted or filtered; otherwise TTS may read them incorrectly. For example, 10:30 should be read as a time, and AI can be pronounced as English letters or in Chinese depending on the product.\u003C\u002Fp>\n\u003Ch3>Pitfall 5: Lacking End-to-End Replay\u003C\u002Fh3>\n\u003Cp>It is recommended to record audio, ASR text, LLM responses, TTS audio, and state events. During review, this makes it possible to determine whether the issue lies in recognition, understanding, generation, or synthesis.\u003C\u002Fp>\n\u003Ch2>Directions for Further Reading\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Streaming ASR and VAD\u003C\u002Fstrong>: Learn about public solutions such as WebRTC, Silero VAD, and RNNoise.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Voice large-model architecture\u003C\u002Fstrong>: Compare cascaded ASR + LLM + TTS with end-to-end speech models; refer to official documentation.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>TTS control\u003C\u002Fstrong>: Pay attention to SSML, prosody, pauses, emotional voices, and multi-speaker synthesis.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Real-time transport\u003C\u002Fstrong>: WebSocket, WebRTC, Opus, jitter buffer, and packet loss recovery.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Experience evaluation\u003C\u002Fstrong>: WER, MOS, first-packet latency, task completion rate, and interruption naturalness.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Overall, combining voice with large models is not simply stitching together three APIs. It requires joint optimization across audio engineering, model serving, product interaction, and monitoring. First get the pipeline working, then continuously iterate around latency, interruption, voice quality, and context memory to build a truly usable real-time voice AI product.\u003C\u002Fp>\n\u003Csection class=\"portal-sources\" style=\"margin-top:2em;font-size:14px;line-height:1.7;color:#5c5850;\">\u003Cp style=\"margin:0 0 0.5em;font-weight:600;\">References\u003C\u002Fp>\u003Cul style=\"margin:0;padding-left:1.25em;\">\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.kookapp.cn\u002F\" rel=\"noopener noreferrer\" target=\"_blank\">KOOK,一个好用的语音沟通工具 - 杭州逍遥一下科技有限公司\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.yy.com\u002Fyy8\u002F\" rel=\"noopener noreferrer\" target=\"_blank\">YY语音官网\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.yy.com\u002Fweb\u002Fpcyy_download\u002F\" rel=\"noopener noreferrer\" target=\"_blank\">YY语音官网\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fnowvoice.ai\u002Fzh\u002F\" rel=\"noopener noreferrer\" target=\"_blank\">免费在线文字转语音与 AI 配音 - NowVoice\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fairvoz.com\u002Fzh\u002Fai-voice-generator\" rel=\"noopener noreferrer\" target=\"_blank\">拥有 500 多种逼真语音的 AI 语音生成器 | 免费文本转语音 ...\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>","2026-09-30T05:36:03.744086+08:00","2026-09-30T05:40:36.329033+08:00","2026-09-30T16:34:32.224323+08:00",["Island",76],{"key":77,"result":78},"PostCard_HtNHQAkw3V43kL9Q3RIE2xzdetuj6nammYg3vtL3c",{"head":79},{"link":80,"style":81},[],[],["Island",83],{"key":84,"result":85},"PostCard_5MrdQReyjlh7tlwvUW44Dx2TYhaatLpjOsynzYD4wPk",{"head":86},{"link":87,"style":88},[],[],["Island",90],{"key":91,"result":92},"PostCard_Ds2Jr0fmMZOVJLg7hvcMHXzDwdLKPhniDm4KJBvQo",{"head":93},{"link":94,"style":95},[],[],1790758653605]