[{"data":1,"prerenderedAt":100},["ShallowReactive",2],{"PortalFooter_QX9Sj6L0jYa0f3sHh8puk8D6pkI69QhELaa1ZMuWxNo":3,"article-hi-ai-agent落地-任务合同五步法":10,"article-body-v-0-0-0":33,"AppImage_od7VkN3TQEMayCJHzz678SjYh8qOLP0UFf0OSmnpOQ0":34,"related-hi-article-ai-agent落地-任务合同五步法":41,"PostCard_qpmVZd2HxBPzTAwZVmapcRp1hQaClaZDeoq5pD3ScE":72,"PostCard_1qIfKluRxAEgu4zE5q9ieIEwJ2xc9VgikWnFqSZUv8":79,"PostCard_s3bL7u0aW72TxlDkceQOL6kJFw1kFChtRgFgelLFm2w":86,"PortalBreadcrumb_BIVLKrH4LzeCPlI6dec7SVr3op8YpHohLmcc3Q7jxs":93},["Island",4],{"key":5,"result":6},"PortalFooter_QX9Sj6L0jYa0f3sHh8puk8D6pkI69QhELaa1ZMuWxNo",{"head":7},{"link":8,"style":9},[],[],{"id":11,"locale":12,"slug":13,"type":14,"title":15,"summary":16,"content_html":17,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"source_job_id":23,"published_at":24,"created_at":25,"updated_at":26,"alternates":27},190,"en","ai-agent落地-任务合同五步法","article","Deploying AI Agents: The Five-Step Task Contract Method","After reading, you can turn vague requirements into a task contract that specifies inputs, outputs, tool permissions, failure budgets, and rollback lines, so you can start the launch review today.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you can turn vague requirements into a task contract that specifies inputs, outputs, tool permissions, failure budgets, and rollback lines, so you can start the launch review today.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxue Studio · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Approximately 3,923 words · About 14 minutes to read\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, a team building a customer-service agent showed me a demo: input ticket text, output handling suggestions, and it ran smoothly on the spot. But once it was placed into a real ticket pool, problems appeared immediately: Will it modify user profiles? Will it promise refunds to customers? Can errors be reversed? Questions like these are hard to answer with prompts alone. A truly deployable agent first needs a task contract.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can think of this contract as the agent’s job description: what it is responsible for, what it can see, what tools it can call, how to roll back when it fails, and who must approve before it touches critical systems. By the end of this article, you will have a five-step launch method, an example set of contract fields, and an architecture-choice comparison table, so you can start the review today.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Let boundaries speak first, then let the model perform.\u003C\u002Fem>\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">1. Define the Agent’s Boundaries First\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Split Requirements into Four Categories; Don’t Rush to Connect Tools\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Many requirements start out vague: build an intelligent customer-service assistant, build a report-summarization agent, build a ticket-triage bot. They all sound runnable, but the boundaries have not been split. The first step in deployment is not choosing a model; it is writing down the goal, inputs, outputs, and prohibited actions clearly.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Goals must be verifiable. For example, assign tickets to the correct queue and provide the next-step suggestion. Inputs must be limited. For example, read only user descriptions, historical ticket summaries, and product rules; do not read phone numbers, addresses, or payment information. Outputs must have a fixed format. For example, output \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">JSON\u003C\u002Fcode>, containing category, priority, confidence, and suggested action.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Prohibited Actions Matter More Than Automatic Actions\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Before an agent goes live, the most dangerous thing is not what it cannot do, but what it does on its own. Refunds, sending coupons, modifying orders, writing emails, creating external accounts, and deleting files—these must first be listed as prohibited actions.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can divide actions into three tiers: automatic, approval-required, and must-disable. Low-risk actions can be automatic, such as generating summaries, classifying labels, and suggesting talk tracks. Medium- and high-risk actions require approval, such as sending emails, submitting tickets, and creating tasks. High-risk actions are disabled outright, such as refunds, price changes, exporting sensitive data, and deleting records.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Replacing a lengthy PRD with a task contract is not meant to reduce communication, but to let product, engineering, and business confirm on the same page whether this agent can actually go live.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— A contract is not a document; it is an executable boundary.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRPXG5kVn7Gty6umlEGr75ibyL0fUJgic2SPNpg4aKSl7BBia8MDTsQGRqcMSEh9PiaMVnLp9be3uNYhWN71RIZN3lqiate8ericeOu0Y\u002F0?from=appmsg\" alt=\"Image: Spring Field Soft Light\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Image: Spring Field Soft Light\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">2. Core Fields of the Task Contract\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">A Minimum Viable Task Contract\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Contract fields do not need to be long, but they must be executable. Below is a YAML example that can be placed directly in the project repository, suitable for scenarios such as ticket triage and report summarization.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">contract:  id: ticket-triage-v1  objective: Classify incoming tickets into the correct queues and provide next-step handling suggestions  context:    ticket_pool_snapshot: current_day    knowledge_base_version: rules_v1    business_rules_version: cs_v3  input:    allowed_fields:      - ticket_text      - user_type      - product_name    forbidden_fields:      - phone      - address      - payment_id  output:    format: json    required:      - category      - priority      - confidence      - next_action  tools:    read_only:      - knowledge_base_search      - ticket_history_query    write_allowed: []    approval_required:      - create_followup_ticket  dependencies:    mcp:      allowed_servers:        - ticket_service      read_only: true    function_calling:      whitelist:        - knowledge_base_search        - ticket_history_query  forbidden:    - refund    - price_change    - export_user_data    - delete_ticket  data_boundary:    retention: 24h    redaction:      - email      - phone  failure_budget:    auto_retry: 1    max_error_rate: threshold_by_business  rollback_line:    mode: human_queue    fallback: rule_based_router\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Every Field Must Be Acceptance-Testable\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The objective field cannot simply say “improve efficiency.” It must be written as conditions that can be counted from logs and spot-checked from results. For example, the category field must match an enum value; tasks with confidence below the threshold must go to the human queue.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Tool permissions cannot simply say “allow calling the knowledge base.” They must specify which API, which parameters, read-only or write, and whether a \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">trace id\u003C\u002Fcode> is recorded. Dependencies such as \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">MCP\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">Function Calling\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">RAG\u003C\u002Fcode> knowledge bases must become verifiable conditions: can it be called, what was called, what was returned, and what happens on failure.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Context fields are equally important. The current ticket-pool snapshot, knowledge-base version, and business-rule version should all be written into the contract. Otherwise, the same type of question may be answered correctly yesterday and incorrectly today, making it hard during troubleshooting to determine whether the issue is caused by the model, the data, or rule drift.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Data boundaries determine the security line. Which fields may enter the model context, which must be redacted, and which cannot be read at all must be hardcoded in the contract. Failure budgets determine the operations line. How many retries are allowed and what fallback triggers when thresholds are exceeded cannot be decided only after errors occur.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Rollback lines determine business confidence. When errors occur, whether to switch to the human queue, downgrade to a rule engine, or pause writes entirely must have a designated owner in advance.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Try read-only first, then enable writes.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkRPCjzBokF0OezvXu42jvAgJFeR30Y9iczutnzLpx1HMBibeXfadjmGWtSPYAwoAG8JHMwlVQHibK9NXo4x5hBYAWmdNhnNgnM0SJ0\u002F0?from=appmsg\" alt=\"Image: Garden Path\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Image: Garden Path\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">3. The Five-Step Deployment Practice\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Step 1: Review the Contract\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Bring product, engineering, and business together for a short meeting and review only the contract fields. Do not discuss model capabilities; first confirm output standards, permission boundaries, and prohibited actions.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">During the review, ask three questions: Who decides when an output is wrong? Who blocks tool calls that exceed boundaries? Can the business side accept the worst-case outcome? If none of the three parties has a clear answer, do not launch yet.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Step 2: Read-Only Pilot\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">A read-only pilot is the safest way to start. The agent can only read data and generate suggestions; it cannot write back to any system. For example, it can only tag tickets, recommend talk tracks for customer-service agents, and generate report outlines.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">During the pilot, focus on three types of evidence: logs that show the full path, metrics that show the error rate, and spot checks that show real business judgment. According to public materials and industry observation, many teams initially hesitate to deploy agents not because the models are not smart enough, but because they lack traceable paths.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Step 3: Gradual Rollout with Approval\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Low-risk actions can be executed automatically. Medium- and high-risk actions go through an approval queue. For example, generating a refund explanation can be automatic, but submitting a refund cannot; creating a follow-up ticket can be automatic, but modifying user contact information must require approval.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The rollout strategy can be opened gradually by business domain, user type, time window, or ticket label. Each expansion must be able to roll back independently.\u003C\u002Fp>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>","","https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg","进学斋","ai","published",3491457,"2026-09-29T17:47:32+08:00","2026-09-29T18:30:42.828261+08:00","2026-09-29T21:35:02.184734+08:00",[28,31],{"locale":29,"path":30},"zh","\u002Farticles\u002Fai-agent落地-任务合同五步法",{"locale":12,"path":32},"\u002Fen\u002Farticles\u002Fai-agent落地-任务合同五步法","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you can turn vague requirements into a task contract that specifies inputs, outputs, tool permissions, failure budgets, and rollback lines, so you can start the launch review today.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\">\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxue Studio · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Approximately 3,923 words · About 14 minutes to read\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\">\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, a team building a customer-service agent showed me a demo: input ticket text, output handling suggestions, and it ran smoothly on the spot. But once it was placed into a real ticket pool, problems appeared immediately: Will it modify user profiles? Will it promise refunds to customers? Can errors be reversed? Questions like these are hard to answer with prompts alone. A truly deployable agent first needs a task contract.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can think of this contract as the agent’s job description: what it is responsible for, what it can see, what tools it can call, how to roll back when it fails, and who must approve before it touches critical systems. By the end of this article, you will have a five-step launch method, an example set of contract fields, and an architecture-choice comparison table, so you can start the review today.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Let boundaries speak first, then let the model perform.\u003C\u002Fem>\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">1. Define the Agent’s Boundaries First\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Split Requirements into Four Categories; Don’t Rush to Connect Tools\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Many requirements start out vague: build an intelligent customer-service assistant, build a report-summarization agent, build a ticket-triage bot. They all sound runnable, but the boundaries have not been split. The first step in deployment is not choosing a model; it is writing down the goal, inputs, outputs, and prohibited actions clearly.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Goals must be verifiable. For example, assign tickets to the correct queue and provide the next-step suggestion. Inputs must be limited. For example, read only user descriptions, historical ticket summaries, and product rules; do not read phone numbers, addresses, or payment information. Outputs must have a fixed format. For example, output \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">JSON\u003C\u002Fcode>, containing category, priority, confidence, and suggested action.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Prohibited Actions Matter More Than Automatic Actions\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Before an agent goes live, the most dangerous thing is not what it cannot do, but what it does on its own. Refunds, sending coupons, modifying orders, writing emails, creating external accounts, and deleting files—these must first be listed as prohibited actions.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can divide actions into three tiers: automatic, approval-required, and must-disable. Low-risk actions can be automatic, such as generating summaries, classifying labels, and suggesting talk tracks. Medium- and high-risk actions require approval, such as sending emails, submitting tickets, and creating tasks. High-risk actions are disabled outright, such as refunds, price changes, exporting sensitive data, and deleting records.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Replacing a lengthy PRD with a task contract is not meant to reduce communication, but to let product, engineering, and business confirm on the same page whether this agent can actually go live.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— A contract is not a document; it is an executable boundary.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRPXG5kVn7Gty6umlEGr75ibyL0fUJgic2SPNpg4aKSl7BBia8MDTsQGRqcMSEh9PiaMVnLp9be3uNYhWN71RIZN3lqiate8ericeOu0Y\u002F0?from=appmsg\" alt=\"Image: Spring Field Soft Light\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\">\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Image: Spring Field Soft Light\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">2. Core Fields of the Task Contract\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">A Minimum Viable Task Contract\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Contract fields do not need to be long, but they must be executable. Below is a YAML example that can be placed directly in the project repository, suitable for scenarios such as ticket triage and report summarization.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">contract:  id: ticket-triage-v1  objective: Classify incoming tickets into the correct queues and provide next-step handling suggestions  context:    ticket_pool_snapshot: current_day    knowledge_base_version: rules_v1    business_rules_version: cs_v3  input:    allowed_fields:      - ticket_text      - user_type      - product_name    forbidden_fields:      - phone      - address      - payment_id  output:    format: json    required:      - category      - priority      - confidence      - next_action  tools:    read_only:      - knowledge_base_search      - ticket_history_query    write_allowed: []    approval_required:      - create_followup_ticket  dependencies:    mcp:      allowed_servers:        - ticket_service      read_only: true    function_calling:      whitelist:        - knowledge_base_search        - ticket_history_query  forbidden:    - refund    - price_change    - export_user_data    - delete_ticket  data_boundary:    retention: 24h    redaction:      - email      - phone  failure_budget:    auto_retry: 1    max_error_rate: threshold_by_business  rollback_line:    mode: human_queue    fallback: rule_based_router\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Every Field Must Be Acceptance-Testable\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The objective field cannot simply say “improve efficiency.” It must be written as conditions that can be counted from logs and spot-checked from results. For example, the category field must match an enum value; tasks with confidence below the threshold must go to the human queue.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Tool permissions cannot simply say “allow calling the knowledge base.” They must specify which API, which parameters, read-only or write, and whether a \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">trace id\u003C\u002Fcode> is recorded. Dependencies such as \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">MCP\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">Function Calling\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">RAG\u003C\u002Fcode> knowledge bases must become verifiable conditions: can it be called, what was called, what was returned, and what happens on failure.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Context fields are equally important. The current ticket-pool snapshot, knowledge-base version, and business-rule version should all be written into the contract. Otherwise, the same type of question may be answered correctly yesterday and incorrectly today, making it hard during troubleshooting to determine whether the issue is caused by the model, the data, or rule drift.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Data boundaries determine the security line. Which fields may enter the model context, which must be redacted, and which cannot be read at all must be hardcoded in the contract. Failure budgets determine the operations line. How many retries are allowed and what fallback triggers when thresholds are exceeded cannot be decided only after errors occur.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Rollback lines determine business confidence. When errors occur, whether to switch to the human queue, downgrade to a rule engine, or pause writes entirely must have a designated owner in advance.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Try read-only first, then enable writes.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkRPCjzBokF0OezvXu42jvAgJFeR30Y9iczutnzLpx1HMBibeXfadjmGWtSPYAwoAG8JHMwlVQHibK9NXo4x5hBYAWmdNhnNgnM0SJ0\u002F0?from=appmsg\" alt=\"Image: Garden Path\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\">\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Image: Garden Path\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">3. The Five-Step Deployment Practice\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Step 1: Review the Contract\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Bring product, engineering, and business together for a short meeting and review only the contract fields. Do not discuss model capabilities; first confirm output standards, permission boundaries, and prohibited actions.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">During the review, ask three questions: Who decides when an output is wrong? Who blocks tool calls that exceed boundaries? Can the business side accept the worst-case outcome? If none of the three parties has a clear answer, do not launch yet.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Step 2: Read-Only Pilot\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">A read-only pilot is the safest way to start. The agent can only read data and generate suggestions; it cannot write back to any system. For example, it can only tag tickets, recommend talk tracks for customer-service agents, and generate report outlines.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">During the pilot, focus on three types of evidence: logs that show the full path, metrics that show the error rate, and spot checks that show real business judgment. According to public materials and industry observation, many teams initially hesitate to deploy agents not because the models are not smart enough, but because they lack traceable paths.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Step 3: Gradual Rollout with Approval\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Low-risk actions can be executed automatically. Medium- and high-risk actions go through an approval queue. For example, generating a refund explanation can be automatic, but submitting a refund cannot; creating a follow-up ticket can be automatic, but modifying user contact information must require approval.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The rollout strategy can be opened gradually by business domain, user type, time window, or ticket label. Each expansion must be able to roll back independently.\u003C\u002Fp>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>\u003C\u002Fsection>",["Island",35],{"key":36,"result":37},"AppImage_od7VkN3TQEMayCJHzz678SjYh8qOLP0UFf0OSmnpOQ0",{"head":38},{"link":39,"style":40},[],[],[42,52,62],{"id":43,"locale":12,"slug":44,"type":14,"title":45,"summary":46,"content_html":47,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"source_job_id":48,"published_at":49,"created_at":50,"updated_at":51},4592,"agent试点验收-七类失败复盘表","Agent Pilot Acceptance: A Postmortem Table for Seven Failure Types","After reading, you will get a seven-category failure attribution table, three acceptance checklists, and autonomy thresholds. Follow the steps to run a read-only pilot first, then decide which actions to automate.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will get a seven-category failure attribution table, three acceptance checklists, and autonomy thresholds. Follow the steps to run a read-only pilot first, then decide which actions should be executed automatically.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text: about 3,633 characters · estimated reading time: 13 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, many teams have placed AI agents into real workflows: customer support tickets, operations inspections, data organization, internal approvals. The demo runs smoothly, but problems appear once multiple people, multiple systems, and long-running execution are involved. On the surface, performance seems unstable; the real blocker is that failure signals have not been broken down.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article provides an acceptance method: a seven-category failure postmortem table, three checklists, a three-day read-only pilot, and staged delegation. You can follow the steps to run a read-only pilot first, then decide which actions to automate.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">I. Make Failure Explicit First: Seven Signals Determine Whether to Continue\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">1. Task Understanding Drift: It Completes a Different Task\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If the user asks the agent to organize customer complaints and generate reply drafts, it may only classify them. If the user asks it to fill in fields, it may rewrite the entire copy. The root cause of drift usually lies in goal decomposition, not the model itself.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can monitor first-pass success rate, human intervention rate, and the proportion of low-quality outputs. If the same intent frequently goes off track, first improve the task template and acceptance fields before considering prompts.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">2. Context Loss: Key Evidence Does Not Reach the Next Step\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">In multi-turn tasks, earlier constraints, customer tier, budget definitions, and historical handling conclusions may be dropped by the time tools are called. The result may look reasonable, but it cannot withstand review.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Observe context reference count, key-field accuracy, and task duration. If repeated queries and repeated confirmations increase, the pipeline is not passing evidence downstream.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">3. Tool Overreach: Actions That Should Not Be Taken Are Taken\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">A read API becomes a write API, a draft becomes a sent message, and a query becomes a deletion. This kind of issue is more dangerous than a wrong answer because it directly changes external state.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Record the number of sensitive-action hits, tool-call input parameters, caller identity, and target system. When overreach is found, first narrow the tool whitelist rather than adding a sentence saying “do not delete.”\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">4. Unverifiable Results: The Output Looks Good, but Cannot Be Traced\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The agent gives a conclusion but leaves no source of evidence, calculation definition, tool return value, or confidence indication. Business colleagues do not dare to sign off, and engineering colleagues cannot reproduce it.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Acceptance should examine key-field accuracy, completeness of the evidence chain, and whether users accept the result. Without verifiable fields, actions should not be handed to automated processes.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">5. Cost Runaway: One Task Turns into a Loop\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Repeated retrieval, repeated summarization, multiple model calls, and multiple tool requests cause the cost per task to keep rising. If the team only watches tokens, it can miss latency, retries, and manual review.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Track repeated call count, task duration, error recovery time, and per-task cost trend. Cost anomalies often mean the workflow has a loop; break the loop before scaling.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">6. Missing Approval: Automated Actions Bypass Humans\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">High-risk actions have no approval point, or approval is merely a formality. When something goes wrong, the chain of responsibility breaks, and the postmortem cannot recover why it was approved.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Check the missing-approval rate, number of external write actions, and retention period for approval records. Approval is not adding a button; it is binding each high-risk action to a person and a reason.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">7. Missing Rollback: Changes Cannot Be Reverted\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Editing documents, sending messages, updating tickets, and adjusting configurations—after failure, there is no rollback ID, no compensating action, and no fallback template. One success may conceal the next incident.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Record the rollback owner, error recovery time, idempotency, and audit ID. If write operations cannot be rolled back, allow them only in low-impact pilot scopes.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— First collect runtime logs, then determine whether the problem lies in the model, prompt, tools, or workflow.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not immediately switch frameworks, expand permissions, or add prompts. First make failures observable; only then does the team have a basis for discussion.\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkROMdibe5hjUkfGygfhy7nUib2bicXAV0oicfMrpdCPHs1NaYib56OA6H166X6IVxU9xib3DlnNuy5icy5bxiaONJ9iaQ9yGpAoR7bgTjfJ0\u002F0?from=appmsg\" alt=\"Illustration · Lakeside Dusk\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Lakeside Dusk\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">II. Define Acceptance Criteria: Replace “Looks Usable” with Three Checklists\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">1. Business Acceptance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The business side asks only whether the result can be used. You can fix five items: task goal completion rate, key-field accuracy, whether users accept the result, proportion of low-quality outputs, and manual review ratio. Each item must specify its definition.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For example, key customer-complaint fields include customer identity, request type, responsible team, and handling deadline. If one is missing, the task is not complete.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">2. Engineering Acceptance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The engineering side must be able to reproduce, track, and stop losses. You can fix: timeout rate, retry count, idempotency, log-field completeness, dependency-service error codes, and call-chain latency distribution.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Logs must include at least \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">run_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">agent_version\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_name\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_hash\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">output_hash\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">duration\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">error_code\u003C\u002Fcode>.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">event: agent.plan\u003Cbr>task_id: T-1024\u003Cbr>run_id: R-77\u003Cbr>user_intent: Organize customer complaints and generate reply drafts\u003Cbr>planned_steps: retrieve, summarize, draft\u003Cbr>tool_call: search_tickets status=open\u003Cbr>risk_level: read_only\u003Cbr>approval_required: false\u003Cbr>idempotency_key: T-1024-R-77-search-001\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This log is not for appearance; it is for attribution after the three-day pilot. If there are only results and no process, the postmortem becomes guesswork.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">3. Security and Compliance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The security focus is not model capability, but action boundaries. You need to check: exposure of sensitive data, external write actions, approvals for deletion\u002Fpayment\u002Fpublishing operations, retention period for audit records, and rollback owner.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Customer-facing outbound messages, fund changes, production configuration changes, and record deletion should not be automatically executed by default. Even if the business is pressing, run them first in a sandbox or shadow system.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Actionable Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Establish a three-key log of \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">run_id\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">agent_version\u003C\u002Fcode> to ensure every execution is traceable.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Record an input summary, output summary, duration, error code, and idempotency key for each tool call.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Place external writes, deletions, payments, publishing, and customer-facing outbound actions on a default block list.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Write clear definitions, criteria, and reviewers for business acceptance fields; do not accept vague evaluations.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Set up daily … for failure samples.\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",3491491,"2026-09-29T20:53:49+08:00","2026-09-29T21:12:09.724901+08:00","2026-09-29T21:35:02.881325+08:00",{"id":53,"locale":12,"slug":54,"type":14,"title":55,"summary":56,"content_html":57,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"source_job_id":58,"published_at":59,"created_at":60,"updated_at":61},4590,"agent观测回滚-落地实战清单","Agent Observability and Rollback: A Practical Implementation Checklist","After reading, you will get a checklist of runtime observability fields for agents and an error-tier rollback table: start with low-traffic logging, then gradually grant more authority.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will get a checklist of runtime observability fields for agents and an error-tier rollback table: start with low-traffic logging, then gradually grant more authority.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text: about 3,580 characters · Reading time: about 12 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Several recent agent projects have entered the runtime phase, and the problems are becoming more complex: the model's answers look good, but the workflow is unstable. Issues range from missing prompts to excessive permissions, high costs, and unnoticed errors.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article focuses only on post-launch runtime data governance. You will get four observability dashboards, three levels of rollback actions, and a seven-day implementation plan. Start by recording low-volume traffic, then gradually grant more authority.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">1. Implementation Bottlenecks: From 'The Model Can Answer' to 'The System Is Manageable'\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">After connecting agents to tools, many teams assume that 'being able to call a tool' means 'being ready for production.' In reality, failures rarely remain in the answer text; they more often appear in broken tool-call chains, permission overreach, uncontrolled costs, and unmanaged human takeover.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For example, a query tool returns an empty value, and the agent continues to generate a conclusion; a write-back API times out, causing a task to be executed repeatedly; a sensitive operation is performed without approval and directly enters a production account. These problems cannot be fundamentally solved by ad hoc prompt edits.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">What must be done during runtime is to turn every execution into a traceable object. First, define stable execution: tasks must be traceable, errors must be stoppable, and costs must be rollback-capable.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Traceable: Every Task Must Have an Identity\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Without \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode> and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">trace_id\u003C\u002Fcode>, agent calls are like anonymous requests. You do not know which role, workflow, or user batch they belong to.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Each call should record at least: task identity, scene, input digest, prompt version, model version, start time, and status.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Stoppable: Errors Must Be Tiered\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Not all errors should be retried. Parameter validation failures can prompt the user to make corrections; repeated tool failures should trigger circuit breaking; operations without approved permissions should stop immediately.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Stop conditions should be written into the system, not left to operations staff watching dashboards.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Rollbackable: Actions Need Exit Paths\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If an agent only generates text, rollback is simple. Once it writes to a database, sends messages, or modifies tickets, you must know where the rollback point is.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Being rollback-capable does not mean undoing everything every time; it means ensuring that errors do not spread.\u003C\u002Fp>\u003Cp align=\"center\" style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Only when it is visible can it be managed.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkRNicSJdZqVvCiasq28Q9qpTaIzccpTIdMFMV6CG1VdagjPZLT84B1Yicz5ppSlAZAShzH8OIJJ570JSEorQEyRGSLus3EPRJJYnLM\u002F0?from=appmsg\" alt=\"配图·山谷柔绿\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Valley in Soft Green\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">2. Observability Model: Four Data Dashboards\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not pile all logs into a single table. During agent runtime, you should split monitoring into at least four dashboards: requests, tools, cost and quality, and audit. Each dashboard answers a different question.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Request Dashboard: Whether a Task Was Triggered Correctly\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It records who the request came from, what task it belongs to, and why it was triggered. It is useful for investigating why this agent started running.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can start by collecting a set of fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">scene\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_digest\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">prompt_version\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">status\u003C\u002Fcode>. The number of fields matters less than being able to link them together.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">trace_id=agent-ticket-001\ntask_id=ticket-8827\nscene=ticket_first_reply\ninput_digest=order_delay_customer_question\nprompt_version=ticket_reply_v3\nstart_time=2026-09-29T10:20:00\nstatus=success\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Tool Dashboard: Where the Chain Breaks\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The tool dashboard is more important than the answer dashboard. It records call order, return codes, latency, and failure points.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It is best to break it down by step: step 1 queries the order, step 2 reads the knowledge base, step 3 generates a draft, and step 4 writes back to the ticket. If step 3 has no write-back, the problem is in the tool layer, not the model layer.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">step_no\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_name\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">args_digest\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">return_code\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">latency_ms\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">retry_count\u003C\u002Fcode>.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Cost and Quality Dashboard: Whether It Is Worth Continuing to Grant Authority\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Agents are not better simply because they are more automated; you need to calculate the costs. Token usage, tool API fees, manual takeover rate, and repeat-question rate all determine whether traffic can be expanded.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For quality metrics, do not look only at accuracy. You can also examine first-time completion rate, number of manual corrections, number of repeated triggers, and whether the result ultimately enters the business system.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tokens_in\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tokens_out\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_cost_estimate\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">manual_takeover_reason\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">repeat_question_count\u003C\u002Fcode>.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Audit Dashboard: Who Approved Which Action\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">When actions involve ticket closure, customer notifications, refund requests, or inventory adjustments, auditing must be a separate dashboard.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It records sensitive operations, approvers, execution results, and whether rollback occurred. Audit is not for the model; it is for the process.\u003C\u002Fp>\u003Cp align=\"center\" style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Rollback is not a step backward; it preserves the right to continue running.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRMBOBM1DicBljibhC3e83iaaT74NwgoO06223ibzZsRbHtgsXfCgp4pEpRIZX9A2WOh9PY4pW622DbAYVMfFWvjXzznsnYqkMvPwGA\u002F0?from=appmsg\" alt=\"配图·静湖云烟\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Still Lake and Misty Clouds\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">3. Rollback Mechanism: Three Action Levels\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Many teams rely on gut feeling for rollback, which is dangerous. Real production deployment requires three levels of action: read-only observation, shadow execution, and minimal closed loop. Each level has permission boundaries and trigger conditions.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Read-Only Observation: Record First, Do Not Touch the Business\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Read-only observation is suitable for the first week of a new scenario. The agent can plan, query, and generate drafts, but it cannot write to production databases or send external messages.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The goal is not to replace humans, but to obtain evidence of the execution chain. You can...\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",3491490,"2026-09-29T20:49:00+08:00","2026-09-29T21:00:41.908764+08:00","2026-09-29T21:35:01.807382+08:00",{"id":63,"locale":12,"slug":64,"type":14,"title":65,"summary":66,"content_html":67,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"source_job_id":68,"published_at":69,"created_at":70,"updated_at":71},186,"agent落地-运行闭环七步法","AI Agent Deployment: The Seven-Step Operational Closed Loop","After reading, you will be able to build a post-launch closed loop for AI agent monitoring, alerting, human takeover, rollback, and evaluation, and use a two-week pilot to decide whether to continue, degrade, or stop.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will be able to build a post-launch closed loop for AI agent monitoring, alerting, human takeover, rollback, and evaluation, and use a two-week pilot to decide whether to continue, degrade, or stop.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text about 3,943 characters · Estimated reading time 14 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, while reviewing several teams that deliver AI agent implementations, I found that their on-site demos were mostly impressive: they could summarize, query databases, and fill in fields. But once connected to real requests, problems appeared: tool timeouts, mistaken permission denials, output drift, and no clear owner for human takeover. This is not because the prompts are not flashy enough; it is because the operational closed loop has not been established.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article focuses only on the post-launch operational closed loop: how to define normal, degraded, and circuit-breaking states; how to replay failures; what minimum viable monitoring looks like; how to prepare human takeover, regression evaluation, and rollback packages; and finally how to use a two-week pilot to decide whether to continue, degrade, or stop.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">According to industry observations, most AI agent incidents are not caused by a single model error, but by operational boundaries that were not clearly defined in advance. Only when the closed loop is built can teams rely less on ad-hoc firefighting.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">1. First Define an Operational State Table: Normal, Degraded, and Circuit-Broken\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Many teams reduce AI agent states to success or failure. In production, failures often expand gradually. A parsing error may first cause missing fields, then trigger tool retries, and finally drive up costs. The meaning of a state table is to let the system know when to continue, when to reduce privileges, and when to stop.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">The State Table Should Answer Four Things\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It must define entry conditions, system actions, recovery conditions, and responsible owners. Without these four items, on-duty personnel can only rely on intuition. The table below can be directly adapted into your team’s version.\u003C\u002Fp>\u003Ctable style=\"width:100%;border-collapse:collapse;font-size:13px;margin:16px 0;\">\u003Ctr>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">State\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Entry Condition\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">System Action\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Recovery Condition\u003C\u002Fth>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Normal\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Core tools are available, output validation passes, costs are within baseline, and sensitive operations are authorized.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Low-risk tasks are executed automatically; medium-risk tasks generate suggestions; high-risk tasks are routed for approval.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Evaluation and monitoring continue to pass, and the agent enters the next review cycle.\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Degraded\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Parsing failures, tool errors, or user corrections rise continuously; quality is unstable for a certain type of task; or costs grow abnormally.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">High-risk automatic execution is disabled, while read-only queries, draft suggestions, and human confirmation are retained.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The failure rate falls within the continuous observation window, regression tests pass, and the responsible person approves restoration.\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Circuit-Broken\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Permission boundaries are crossed, critical dependencies are unavailable, output causes business impact, or cost or data risk becomes uncontrollable.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The entry point is disabled, evidence is preserved, the process is switched back to the previous workflow, and the responsible person and affected users are notified.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The root cause has been fixed, the incident retrospective is complete, and a small-scale pilot passes.\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftable>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">There is a practical principle here: degradation is not turning the AI agent into a decoration, but narrowing its capabilities to a range that is explainable, auditable, and recoverable.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">First Clarify Permissions for Three Task Types\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If permission boundaries are not clearly written before launch, problems will appear during operation: the model suggests refunds, customer service directly modifies orders, or external actions are submitted without user confirmation. At minimum, divide tasks into three types:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Can be executed automatically\u003C\u002Fstrong>: knowledge base retrieval, material summarization, initial ticket classification, field-extraction drafts, and read-only status queries. Outputs must be verifiable and must not directly change external state.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Suggestion only\u003C\u002Fstrong>: contract clause modifications, fee explanations, risk scripts, and cross-department process recommendations. Progress is allowed only after human confirmation.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Approval required\u003C\u002Fstrong>: refunds, account permission changes, data deletion, external sending, financial operations, and write actions that affect user rights. The approver, reason, and evidence must be retained.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Key Outputs Must Include Evidence Fields\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Without evidence fields, the group chat will only leave a sentence saying “something went wrong again.” Evidence should help people quickly answer: where the input came from, what the tool returned, whether validation passed, and which layer failed.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_source: ticket_id, doc_id, user_context_version\ntool_calls: retrieval_status, extraction_status, validation_status\noutput_check: schema_passed, safety_passed, business_rule_passed\nfailure_reason: plan_timeout, tool_error, permission_denied, user_correction\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">These fields do not necessarily all need to be shown to users, but they must be present in the logs.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— The state table is the control plane; evidence logs are the microscope.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkROdDNdbMQuptW8ogdgFPnJdzGnbEwAEXKeb6NmUrRcxgqkCV0g9eTkHU57XJbrodBxjX5DDzOvFeOvS5qxic8vAXcL9hTibLu7qk\u002F0?from=appmsg\" alt=\"Illustration · Warm Autumn Hues\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Warm Autumn Hues\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">2. Minimum Monitoring and Evidence Checklist\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not start monitoring by building a large dashboard. First capture five failure types; they are more actionable than overall accuracy.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Start with Five Failure Types\u003C\u002Fh3>\u003Cul>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Parsing failure\u003C\u002Fstrong>: fields are not extracted, JSON does not match the schema, or table rows and columns are misaligned. The problem is often in input format and schema constraints.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Planning timeout\u003C\u002Fstrong>: task decomposition is too long, multi-round retrieval jumps repeatedly, or tool selection is hesitant. The problem is often in task boundaries and loop control.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Tool error\u003C\u002Fstrong>: a dependent API returns an exception, permissions expire, or parameter mapping is wrong. The problem is often in the tool chain and error propagation.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Permission denial\u003C\u002Fstrong>: the AI agent calls an action it should not call, or accesses unauthorized data. The problem is often in policy configuration and the approval chain.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">User correction\u003C\u002Fstrong>: the user explicitly says the answer is irrelevant, the evidence is insufficient, or the suggestion cannot be executed. The problem is often in task value and matching real scenarios.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Build Task-Level Logs\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The goal of logs is replayability. Given a request_id, you should be able to reconstruct which tools were called, which version was used, where it got stuck, and whether human takeover occurred. Fields do not need to be exhaustive, but they need to be stable.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">request_id\ntask_type\nagent_version\nprompt_version\ntool_call_chain\nlatency\ncost\nvalidation_status\nfail_layer\nhuman_takeover_count\nuser_feedback_id\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Set Actionable Thresholds\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Thresholds do not need complex models; first cover changes that can cause incidents. You can use fixed windows or sliding windows, but every threshold must correspond to an action.\u003C\u002Fp>\u003Col>\u003Cli>If the same task fails consecutively up to a preset count, automatically switch to suggestion mode and pause write operations.\u003C\u002Fli>\u003Cli>If denials of sensitive tools surge, freeze related actions and check permission mapping and input sources.\u003C\u002Fli>\u003Cli>If costs rise abnormally compared with the historical baseline, enable caching, rate limiting, and short-answer mode.\u003C\u002Fli>\u003Cli>If human takeover and user corrections rise continuously, trigger regression evaluation and rollback assessment.\u003C\u002Fli>\u003C\u002Fol>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Minimum Monitoring and Evidence Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Each request has a unique request_id;\u003Cbr>☐ Each tool call has status, latency, and error code;\u003Cbr>☐ Each failure can be attributed to one of five layers: input, planning, tool, output, and permission;\u003Cbr>☐ Each degradation switch has a responsible owner;\u003Cbr>☐ Each human takeover has an approval record and evidence package;\u003Cbr>☐ Each version change has an old entry point and rollback path;\u003Cbr>☐ Each alert has a clear action, rather than only sending a notification.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— The meaning of a closed loop is to ensure that the next anomaly no longer relies on ad-hoc firefighting.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRP5WeeqJZQxz7mAH0BoQwyJMKcKP4sOtfgAIljZh9YTMqB0bVF1SKgREkNLmfricnn6MdyYpIeW1g6YRWhZOoflWgeuOuq5lXko\u002F0?from=appmsg\" alt=\"Illustration · Bamboo Shadows by the River\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Bamboo Shadows by the River\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",3491459,"2026-09-29T17:56:37+08:00","2026-09-29T18:08:27.848109+08:00","2026-09-29T21:35:02.049844+08:00",["Island",73],{"key":74,"result":75},"PostCard_qpmVZd2HxBPzTAwZVmapcRp1hQaClaZDeoq5pD3ScE",{"head":76},{"link":77,"style":78},[],[],["Island",80],{"key":81,"result":82},"PostCard_1qIfKluRxAEgu4zE5q9ieIEwJ2xc9VgikWnFqSZUv8",{"head":83},{"link":84,"style":85},[],[],["Island",87],{"key":88,"result":89},"PostCard_s3bL7u0aW72TxlDkceQOL6kJFw1kFChtRgFgelLFm2w",{"head":90},{"link":91,"style":92},[],[],["Island",94],{"key":95,"result":96},"PortalBreadcrumb_BIVLKrH4LzeCPlI6dec7SVr3op8YpHohLmcc3Q7jxs",{"head":97},{"link":98,"style":99},[],[],1790694125983]