[{"data":1,"prerenderedAt":100},["ShallowReactive",2],{"PortalFooter_QX9Sj6L0jYa0f3sHh8puk8D6pkI69QhELaa1ZMuWxNo":3,"article-id-agent-落地实战-把失败写成可恢复状态":10,"article-body-v-0-0-0":33,"AppImage_HG4CFo1pNtDrqi6BQmOttFDghx49PDmMht7TnFM7uE":34,"PortalBreadcrumb_HtABvfeO9pJVbSh0jamzxbCsdXXISUm1ppNzPvFh1I":41,"related-id-article-agent-落地实战-把失败写成可恢复状态":48,"PostCard_J6VZJE8frnTuFgkSyZmdy1L28Nz3U6IhAq3R0H9WSY":79,"PostCard_RvrKxYkuXkpbD77MaDmpwO8eFJX1aWd2cQyVUSdoA":86,"PostCard_Pugn6ocQ4tEtMeBqWpF1UbcMQFVJx689gQ7Db4wQo":93},["Island",4],{"key":5,"result":6},"PortalFooter_QX9Sj6L0jYa0f3sHh8puk8D6pkI69QhELaa1ZMuWxNo",{"head":7},{"link":8,"style":9},[],[],{"id":11,"locale":12,"slug":13,"type":14,"title":15,"summary":16,"content_html":17,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"source_job_id":23,"published_at":24,"created_at":25,"updated_at":26,"alternates":27},160,"en","agent-落地实战-把失败写成可恢复状态","article","Agent in Production: Turning Failures into Recoverable States","After reading, you get a business event-sharding table, a failure-recovery checklist, and pre-launch validation steps to complete a controlled Agent pilot in two days.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you get a business event-sharding table, a failure-recovery checklist, and pre-launch validation steps to complete a controlled Agent pilot in two days.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxue Studio · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">About 3,789 characters · about 13 minutes to read\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article gives you an event-sharding table, a failure-recovery checklist, and pre-launch inspection steps. You can use them to run a low-risk Agent pilot: it can look up materials, generate drafts, and record state; when it fails, you know who takes over, how to recover, and what evidence supports the review.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, many teams have turned Agents into “conference-room stars”: they can read tickets, search knowledge bases, and write drafts; in a demo, one sentence can pull the materials together. But once they are connected to real business processes, the problems appear: sequential calls need state confirmation, failures need retries, write operations need traces, and audits need traceability.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">I. Background: Why Demos Stop in the Conference Room\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Q&amp;A, retrieval, and drafting are short tasks with light failures, so they are good for demos. Business processes need continuous execution: first query the ticket, then read the knowledge base, then generate a reply draft, then write the result into the system.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Once the chain gets long, failure is no longer “the result is inaccurate,” but “whether it has already run,” “whether it ran twice,” and “whether it can be undone.” Based on public materials and industry observation, many pilots settle quickly on model selection but are slow to fill in event sharding, state recording, and rollback actions.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can first choose a low-risk closed loop: internal ticket summarization, log classification, or customer-service knowledge retrieval. The key is not that the task is simple, but that failure will not contaminate production data.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">— Being able to run is only the starting point; being recoverable is what earns production eligibility.\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRPEsnGZzO4vcZZgrgUYbxicOtvtrT7HInqmLrqxBlV1M3bQZEicHjzAyR2IgaszGn3ibX5cz9S5tbzASwuax6ppLRG4JvhiaATWP9Y\u002F0?from=appmsg\" alt=\"Illustration · Soft Light in Spring Fields\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Soft Light in Spring Fields\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">II. Mechanism: State Sharding and Runtime Evidence\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">1. Cut Long Workflows into Stoppable, Auditable Steps\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">A runnable Agent should not be treated as a one-shot black box. A more stable approach is to split the workflow into short nodes such as event, decision, invocation, validation, and delivery. Each node needs inputs, outputs, timeouts, retries, and human takeover points.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For example, “customer-service ticket summarization” can be split into: read ticket, search knowledge base, generate draft, check sensitive fields, save draft, and notify a human. Each step has its own state, so when it fails you know where it is stuck.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">2. Leave a Minimal Evidence Package for Each Node\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The evidence package is not a chat log, nor free-form text. It must answer: who, at which step, which tool was used, what action was taken, and whether it continued automatically.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The minimal fields can be few: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">event_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">step_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_digest\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_name\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">params_masked\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">result_status\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">duration_ms\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">next_action\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">retryable\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">human_required\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">rollback_basis\u003C\u002Fcode>.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">{\n  &quot;task_id&quot;: &quot;ticket-042&quot;,\n  &quot;event_id&quot;: &quot;cust-export-timeout&quot;,\n  &quot;step_id&quot;: &quot;summarize-draft&quot;,\n  &quot;input_digest&quot;: &quot;Customer report: export report timed out&quot;,\n  &quot;tool_name&quot;: &quot;kb.search&quot;,\n  &quot;params_masked&quot;: &quot;tenant=t-a1b2|query=export&quot;,\n  &quot;result_status&quot;: &quot;success&quot;,\n  &quot;duration_ms&quot;: 1800,\n  &quot;next_action&quot;: &quot;save_draft&quot;,\n  &quot;retryable&quot;: false,\n  &quot;human_required&quot;: false,\n  &quot;rollback_basis&quot;: &quot;no_write&quot;\n}\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This record does not aim to be elegant; it aims to withstand scrutiny. When auditors see the status as \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">success\u003C\u002Fcode>, they also need to see whether the next step is \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">save_draft\u003C\u002Fcode>; when they see a timeout as \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">timeout\u003C\u002Fcode>, they also need to know whether it entered human confirmation.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">3. Do Not Rely on Free-Form Text for State\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Descriptions such as “draft generated” feel natural, but they can become unreliable during audits. More reliable state fields are: whether it has executed, whether it is retryable, whether human confirmation is required, and whether there is a rollback basis.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If a write operation has been sent but the external system times out, you cannot simply mark it as failed, nor automatically resend it. It needs at least three markers: “may have occurred,” “requires human confirmation,” and “compensation action available.”\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">— Engineering is not about making machines talk better; it is about making every step withstand scrutiny.\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkRN0bZC7VuyXtxwZJiaykyZUcUovfqyfLshkRXiaA7yEYMteC3VReWzdeVxcNSVOXSVAIqkJNoqKIBLXxrEtWEFN7K8NTqdd63dbk\u002F0?from=appmsg\" alt=\"Illustration · Garden Path\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Garden Path\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">III. Steps: Run a Small Closed Loop in Two Days\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If the schedule is tight, you can fold the Day 3 checks into the end of Day 2. The core is to make the evidence complete first, then make the workflow run smoothly.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Day1: Establish Admission and Permissions\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Choose a low-risk event, such as internal ticket summarization. First write a task admission table: what events can enter, which fields must be masked, and which results must not be sent automatically.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Then write a tool permission table: allow only read-only queries, summary generation, and draft saving. Do not give the Agent actions such as production database writes, message sending, or refund compensation at first.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Event-Sharding Table\u003C\u002Fh3>\u003Ctable style=\"width:100%;border-collapse:collapse;font-size:13px;margin:16px 0;\">\u003Ctr>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Node\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Input\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Output\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Timeout\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Failure action\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Human checkpoint\u003C\u002Fth>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Read ticket\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Ticket ID, tenant identifier\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Problem summary, field list\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">API contract\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Skip and notify a human\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Confirm when fields are missing\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Search knowledge base\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Summary keywords\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>","","https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg","进学斋","ai","published",3491397,"2026-09-29T12:15:42+08:00","2026-09-29T12:54:47.846895+08:00","2026-09-29T21:35:01.865938+08:00",[28,31],{"locale":29,"path":30},"zh","\u002Farticles\u002Fagent-落地实战-把失败写成可恢复状态",{"locale":12,"path":32},"\u002Fen\u002Farticles\u002Fagent-落地实战-把失败写成可恢复状态","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you get a business event-sharding table, a failure-recovery checklist, and pre-launch validation steps to complete a controlled Agent pilot in two days.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\">\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxue Studio · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">About 3,789 characters · about 13 minutes to read\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\">\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article gives you an event-sharding table, a failure-recovery checklist, and pre-launch inspection steps. You can use them to run a low-risk Agent pilot: it can look up materials, generate drafts, and record state; when it fails, you know who takes over, how to recover, and what evidence supports the review.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, many teams have turned Agents into “conference-room stars”: they can read tickets, search knowledge bases, and write drafts; in a demo, one sentence can pull the materials together. But once they are connected to real business processes, the problems appear: sequential calls need state confirmation, failures need retries, write operations need traces, and audits need traceability.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">I. Background: Why Demos Stop in the Conference Room\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Q&amp;A, retrieval, and drafting are short tasks with light failures, so they are good for demos. Business processes need continuous execution: first query the ticket, then read the knowledge base, then generate a reply draft, then write the result into the system.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Once the chain gets long, failure is no longer “the result is inaccurate,” but “whether it has already run,” “whether it ran twice,” and “whether it can be undone.” Based on public materials and industry observation, many pilots settle quickly on model selection but are slow to fill in event sharding, state recording, and rollback actions.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can first choose a low-risk closed loop: internal ticket summarization, log classification, or customer-service knowledge retrieval. The key is not that the task is simple, but that failure will not contaminate production data.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">— Being able to run is only the starting point; being recoverable is what earns production eligibility.\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRPEsnGZzO4vcZZgrgUYbxicOtvtrT7HInqmLrqxBlV1M3bQZEicHjzAyR2IgaszGn3ibX5cz9S5tbzASwuax6ppLRG4JvhiaATWP9Y\u002F0?from=appmsg\" alt=\"Illustration · Soft Light in Spring Fields\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\">\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Soft Light in Spring Fields\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">II. Mechanism: State Sharding and Runtime Evidence\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">1. Cut Long Workflows into Stoppable, Auditable Steps\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">A runnable Agent should not be treated as a one-shot black box. A more stable approach is to split the workflow into short nodes such as event, decision, invocation, validation, and delivery. Each node needs inputs, outputs, timeouts, retries, and human takeover points.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For example, “customer-service ticket summarization” can be split into: read ticket, search knowledge base, generate draft, check sensitive fields, save draft, and notify a human. Each step has its own state, so when it fails you know where it is stuck.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">2. Leave a Minimal Evidence Package for Each Node\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The evidence package is not a chat log, nor free-form text. It must answer: who, at which step, which tool was used, what action was taken, and whether it continued automatically.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The minimal fields can be few: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">event_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">step_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_digest\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_name\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">params_masked\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">result_status\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">duration_ms\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">next_action\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">retryable\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">human_required\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">rollback_basis\u003C\u002Fcode>.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">{\n  \"task_id\": \"ticket-042\",\n  \"event_id\": \"cust-export-timeout\",\n  \"step_id\": \"summarize-draft\",\n  \"input_digest\": \"Customer report: export report timed out\",\n  \"tool_name\": \"kb.search\",\n  \"params_masked\": \"tenant=t-a1b2|query=export\",\n  \"result_status\": \"success\",\n  \"duration_ms\": 1800,\n  \"next_action\": \"save_draft\",\n  \"retryable\": false,\n  \"human_required\": false,\n  \"rollback_basis\": \"no_write\"\n}\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This record does not aim to be elegant; it aims to withstand scrutiny. When auditors see the status as \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">success\u003C\u002Fcode>, they also need to see whether the next step is \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">save_draft\u003C\u002Fcode>; when they see a timeout as \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">timeout\u003C\u002Fcode>, they also need to know whether it entered human confirmation.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">3. Do Not Rely on Free-Form Text for State\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Descriptions such as “draft generated” feel natural, but they can become unreliable during audits. More reliable state fields are: whether it has executed, whether it is retryable, whether human confirmation is required, and whether there is a rollback basis.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If a write operation has been sent but the external system times out, you cannot simply mark it as failed, nor automatically resend it. It needs at least three markers: “may have occurred,” “requires human confirmation,” and “compensation action available.”\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">— Engineering is not about making machines talk better; it is about making every step withstand scrutiny.\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkRN0bZC7VuyXtxwZJiaykyZUcUovfqyfLshkRXiaA7yEYMteC3VReWzdeVxcNSVOXSVAIqkJNoqKIBLXxrEtWEFN7K8NTqdd63dbk\u002F0?from=appmsg\" alt=\"Illustration · Garden Path\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\">\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Garden Path\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">III. Steps: Run a Small Closed Loop in Two Days\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If the schedule is tight, you can fold the Day 3 checks into the end of Day 2. The core is to make the evidence complete first, then make the workflow run smoothly.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Day1: Establish Admission and Permissions\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Choose a low-risk event, such as internal ticket summarization. First write a task admission table: what events can enter, which fields must be masked, and which results must not be sent automatically.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Then write a tool permission table: allow only read-only queries, summary generation, and draft saving. Do not give the Agent actions such as production database writes, message sending, or refund compensation at first.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Event-Sharding Table\u003C\u002Fh3>\u003Ctable style=\"width:100%;border-collapse:collapse;font-size:13px;margin:16px 0;\">\u003Ctbody>\u003Ctr>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Node\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Input\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Output\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Timeout\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Failure action\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Human checkpoint\u003C\u002Fth>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Read ticket\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Ticket ID, tenant identifier\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Problem summary, field list\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">API contract\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Skip and notify a human\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Confirm when fields are missing\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Search knowledge base\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Summary keywords\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003C\u002Fsection>",["Island",35],{"key":36,"result":37},"AppImage_HG4CFo1pNtDrqi6BQmOttFDghx49PDmMht7TnFM7uE",{"head":38},{"link":39,"style":40},[],[],["Island",42],{"key":43,"result":44},"PortalBreadcrumb_HtABvfeO9pJVbSh0jamzxbCsdXXISUm1ppNzPvFh1I",{"head":45},{"link":46,"style":47},[],[],[49,59,69],{"id":50,"locale":12,"slug":51,"type":14,"title":52,"summary":53,"content_html":54,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"source_job_id":55,"published_at":56,"created_at":57,"updated_at":58},4592,"agent试点验收-七类失败复盘表","Agent Pilot Acceptance: A Postmortem Table for Seven Failure Types","After reading, you will get a seven-category failure attribution table, three acceptance checklists, and autonomy thresholds. Follow the steps to run a read-only pilot first, then decide which actions to automate.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will get a seven-category failure attribution table, three acceptance checklists, and autonomy thresholds. Follow the steps to run a read-only pilot first, then decide which actions should be executed automatically.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text: about 3,633 characters · estimated reading time: 13 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, many teams have placed AI agents into real workflows: customer support tickets, operations inspections, data organization, internal approvals. The demo runs smoothly, but problems appear once multiple people, multiple systems, and long-running execution are involved. On the surface, performance seems unstable; the real blocker is that failure signals have not been broken down.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article provides an acceptance method: a seven-category failure postmortem table, three checklists, a three-day read-only pilot, and staged delegation. You can follow the steps to run a read-only pilot first, then decide which actions to automate.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">I. Make Failure Explicit First: Seven Signals Determine Whether to Continue\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">1. Task Understanding Drift: It Completes a Different Task\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If the user asks the agent to organize customer complaints and generate reply drafts, it may only classify them. If the user asks it to fill in fields, it may rewrite the entire copy. The root cause of drift usually lies in goal decomposition, not the model itself.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can monitor first-pass success rate, human intervention rate, and the proportion of low-quality outputs. If the same intent frequently goes off track, first improve the task template and acceptance fields before considering prompts.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">2. Context Loss: Key Evidence Does Not Reach the Next Step\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">In multi-turn tasks, earlier constraints, customer tier, budget definitions, and historical handling conclusions may be dropped by the time tools are called. The result may look reasonable, but it cannot withstand review.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Observe context reference count, key-field accuracy, and task duration. If repeated queries and repeated confirmations increase, the pipeline is not passing evidence downstream.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">3. Tool Overreach: Actions That Should Not Be Taken Are Taken\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">A read API becomes a write API, a draft becomes a sent message, and a query becomes a deletion. This kind of issue is more dangerous than a wrong answer because it directly changes external state.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Record the number of sensitive-action hits, tool-call input parameters, caller identity, and target system. When overreach is found, first narrow the tool whitelist rather than adding a sentence saying “do not delete.”\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">4. Unverifiable Results: The Output Looks Good, but Cannot Be Traced\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The agent gives a conclusion but leaves no source of evidence, calculation definition, tool return value, or confidence indication. Business colleagues do not dare to sign off, and engineering colleagues cannot reproduce it.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Acceptance should examine key-field accuracy, completeness of the evidence chain, and whether users accept the result. Without verifiable fields, actions should not be handed to automated processes.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">5. Cost Runaway: One Task Turns into a Loop\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Repeated retrieval, repeated summarization, multiple model calls, and multiple tool requests cause the cost per task to keep rising. If the team only watches tokens, it can miss latency, retries, and manual review.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Track repeated call count, task duration, error recovery time, and per-task cost trend. Cost anomalies often mean the workflow has a loop; break the loop before scaling.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">6. Missing Approval: Automated Actions Bypass Humans\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">High-risk actions have no approval point, or approval is merely a formality. When something goes wrong, the chain of responsibility breaks, and the postmortem cannot recover why it was approved.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Check the missing-approval rate, number of external write actions, and retention period for approval records. Approval is not adding a button; it is binding each high-risk action to a person and a reason.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">7. Missing Rollback: Changes Cannot Be Reverted\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Editing documents, sending messages, updating tickets, and adjusting configurations—after failure, there is no rollback ID, no compensating action, and no fallback template. One success may conceal the next incident.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Record the rollback owner, error recovery time, idempotency, and audit ID. If write operations cannot be rolled back, allow them only in low-impact pilot scopes.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— First collect runtime logs, then determine whether the problem lies in the model, prompt, tools, or workflow.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not immediately switch frameworks, expand permissions, or add prompts. First make failures observable; only then does the team have a basis for discussion.\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkROMdibe5hjUkfGygfhy7nUib2bicXAV0oicfMrpdCPHs1NaYib56OA6H166X6IVxU9xib3DlnNuy5icy5bxiaONJ9iaQ9yGpAoR7bgTjfJ0\u002F0?from=appmsg\" alt=\"Illustration · Lakeside Dusk\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Lakeside Dusk\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">II. Define Acceptance Criteria: Replace “Looks Usable” with Three Checklists\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">1. Business Acceptance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The business side asks only whether the result can be used. You can fix five items: task goal completion rate, key-field accuracy, whether users accept the result, proportion of low-quality outputs, and manual review ratio. Each item must specify its definition.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For example, key customer-complaint fields include customer identity, request type, responsible team, and handling deadline. If one is missing, the task is not complete.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">2. Engineering Acceptance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The engineering side must be able to reproduce, track, and stop losses. You can fix: timeout rate, retry count, idempotency, log-field completeness, dependency-service error codes, and call-chain latency distribution.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Logs must include at least \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">run_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">agent_version\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_name\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_hash\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">output_hash\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">duration\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">error_code\u003C\u002Fcode>.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">event: agent.plan\u003Cbr>task_id: T-1024\u003Cbr>run_id: R-77\u003Cbr>user_intent: Organize customer complaints and generate reply drafts\u003Cbr>planned_steps: retrieve, summarize, draft\u003Cbr>tool_call: search_tickets status=open\u003Cbr>risk_level: read_only\u003Cbr>approval_required: false\u003Cbr>idempotency_key: T-1024-R-77-search-001\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This log is not for appearance; it is for attribution after the three-day pilot. If there are only results and no process, the postmortem becomes guesswork.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">3. Security and Compliance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The security focus is not model capability, but action boundaries. You need to check: exposure of sensitive data, external write actions, approvals for deletion\u002Fpayment\u002Fpublishing operations, retention period for audit records, and rollback owner.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Customer-facing outbound messages, fund changes, production configuration changes, and record deletion should not be automatically executed by default. Even if the business is pressing, run them first in a sandbox or shadow system.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Actionable Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Establish a three-key log of \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">run_id\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">agent_version\u003C\u002Fcode> to ensure every execution is traceable.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Record an input summary, output summary, duration, error code, and idempotency key for each tool call.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Place external writes, deletions, payments, publishing, and customer-facing outbound actions on a default block list.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Write clear definitions, criteria, and reviewers for business acceptance fields; do not accept vague evaluations.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Set up daily … for failure samples.\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",3491491,"2026-09-29T20:53:49+08:00","2026-09-29T21:12:09.724901+08:00","2026-09-29T21:35:02.881325+08:00",{"id":60,"locale":12,"slug":61,"type":14,"title":62,"summary":63,"content_html":64,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"source_job_id":65,"published_at":66,"created_at":67,"updated_at":68},4590,"agent观测回滚-落地实战清单","Agent Observability and Rollback: A Practical Implementation Checklist","After reading, you will get a checklist of runtime observability fields for agents and an error-tier rollback table: start with low-traffic logging, then gradually grant more authority.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will get a checklist of runtime observability fields for agents and an error-tier rollback table: start with low-traffic logging, then gradually grant more authority.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text: about 3,580 characters · Reading time: about 12 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Several recent agent projects have entered the runtime phase, and the problems are becoming more complex: the model's answers look good, but the workflow is unstable. Issues range from missing prompts to excessive permissions, high costs, and unnoticed errors.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article focuses only on post-launch runtime data governance. You will get four observability dashboards, three levels of rollback actions, and a seven-day implementation plan. Start by recording low-volume traffic, then gradually grant more authority.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">1. Implementation Bottlenecks: From 'The Model Can Answer' to 'The System Is Manageable'\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">After connecting agents to tools, many teams assume that 'being able to call a tool' means 'being ready for production.' In reality, failures rarely remain in the answer text; they more often appear in broken tool-call chains, permission overreach, uncontrolled costs, and unmanaged human takeover.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For example, a query tool returns an empty value, and the agent continues to generate a conclusion; a write-back API times out, causing a task to be executed repeatedly; a sensitive operation is performed without approval and directly enters a production account. These problems cannot be fundamentally solved by ad hoc prompt edits.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">What must be done during runtime is to turn every execution into a traceable object. First, define stable execution: tasks must be traceable, errors must be stoppable, and costs must be rollback-capable.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Traceable: Every Task Must Have an Identity\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Without \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode> and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">trace_id\u003C\u002Fcode>, agent calls are like anonymous requests. You do not know which role, workflow, or user batch they belong to.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Each call should record at least: task identity, scene, input digest, prompt version, model version, start time, and status.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Stoppable: Errors Must Be Tiered\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Not all errors should be retried. Parameter validation failures can prompt the user to make corrections; repeated tool failures should trigger circuit breaking; operations without approved permissions should stop immediately.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Stop conditions should be written into the system, not left to operations staff watching dashboards.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Rollbackable: Actions Need Exit Paths\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If an agent only generates text, rollback is simple. Once it writes to a database, sends messages, or modifies tickets, you must know where the rollback point is.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Being rollback-capable does not mean undoing everything every time; it means ensuring that errors do not spread.\u003C\u002Fp>\u003Cp align=\"center\" style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Only when it is visible can it be managed.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkRNicSJdZqVvCiasq28Q9qpTaIzccpTIdMFMV6CG1VdagjPZLT84B1Yicz5ppSlAZAShzH8OIJJ570JSEorQEyRGSLus3EPRJJYnLM\u002F0?from=appmsg\" alt=\"配图·山谷柔绿\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Valley in Soft Green\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">2. Observability Model: Four Data Dashboards\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not pile all logs into a single table. During agent runtime, you should split monitoring into at least four dashboards: requests, tools, cost and quality, and audit. Each dashboard answers a different question.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Request Dashboard: Whether a Task Was Triggered Correctly\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It records who the request came from, what task it belongs to, and why it was triggered. It is useful for investigating why this agent started running.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can start by collecting a set of fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">scene\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_digest\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">prompt_version\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">status\u003C\u002Fcode>. The number of fields matters less than being able to link them together.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">trace_id=agent-ticket-001\ntask_id=ticket-8827\nscene=ticket_first_reply\ninput_digest=order_delay_customer_question\nprompt_version=ticket_reply_v3\nstart_time=2026-09-29T10:20:00\nstatus=success\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Tool Dashboard: Where the Chain Breaks\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The tool dashboard is more important than the answer dashboard. It records call order, return codes, latency, and failure points.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It is best to break it down by step: step 1 queries the order, step 2 reads the knowledge base, step 3 generates a draft, and step 4 writes back to the ticket. If step 3 has no write-back, the problem is in the tool layer, not the model layer.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">step_no\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_name\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">args_digest\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">return_code\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">latency_ms\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">retry_count\u003C\u002Fcode>.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Cost and Quality Dashboard: Whether It Is Worth Continuing to Grant Authority\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Agents are not better simply because they are more automated; you need to calculate the costs. Token usage, tool API fees, manual takeover rate, and repeat-question rate all determine whether traffic can be expanded.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For quality metrics, do not look only at accuracy. You can also examine first-time completion rate, number of manual corrections, number of repeated triggers, and whether the result ultimately enters the business system.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tokens_in\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tokens_out\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_cost_estimate\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">manual_takeover_reason\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">repeat_question_count\u003C\u002Fcode>.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Audit Dashboard: Who Approved Which Action\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">When actions involve ticket closure, customer notifications, refund requests, or inventory adjustments, auditing must be a separate dashboard.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It records sensitive operations, approvers, execution results, and whether rollback occurred. Audit is not for the model; it is for the process.\u003C\u002Fp>\u003Cp align=\"center\" style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Rollback is not a step backward; it preserves the right to continue running.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRMBOBM1DicBljibhC3e83iaaT74NwgoO06223ibzZsRbHtgsXfCgp4pEpRIZX9A2WOh9PY4pW622DbAYVMfFWvjXzznsnYqkMvPwGA\u002F0?from=appmsg\" alt=\"配图·静湖云烟\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Still Lake and Misty Clouds\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">3. Rollback Mechanism: Three Action Levels\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Many teams rely on gut feeling for rollback, which is dangerous. Real production deployment requires three levels of action: read-only observation, shadow execution, and minimal closed loop. Each level has permission boundaries and trigger conditions.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Read-Only Observation: Record First, Do Not Touch the Business\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Read-only observation is suitable for the first week of a new scenario. The agent can plan, query, and generate drafts, but it cannot write to production databases or send external messages.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The goal is not to replace humans, but to obtain evidence of the execution chain. You can...\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",3491490,"2026-09-29T20:49:00+08:00","2026-09-29T21:00:41.908764+08:00","2026-09-29T21:35:01.807382+08:00",{"id":70,"locale":12,"slug":71,"type":14,"title":72,"summary":73,"content_html":74,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"source_job_id":75,"published_at":76,"created_at":77,"updated_at":78},186,"agent落地-运行闭环七步法","AI Agent Deployment: The Seven-Step Operational Closed Loop","After reading, you will be able to build a post-launch closed loop for AI agent monitoring, alerting, human takeover, rollback, and evaluation, and use a two-week pilot to decide whether to continue, degrade, or stop.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will be able to build a post-launch closed loop for AI agent monitoring, alerting, human takeover, rollback, and evaluation, and use a two-week pilot to decide whether to continue, degrade, or stop.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text about 3,943 characters · Estimated reading time 14 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, while reviewing several teams that deliver AI agent implementations, I found that their on-site demos were mostly impressive: they could summarize, query databases, and fill in fields. But once connected to real requests, problems appeared: tool timeouts, mistaken permission denials, output drift, and no clear owner for human takeover. This is not because the prompts are not flashy enough; it is because the operational closed loop has not been established.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article focuses only on the post-launch operational closed loop: how to define normal, degraded, and circuit-breaking states; how to replay failures; what minimum viable monitoring looks like; how to prepare human takeover, regression evaluation, and rollback packages; and finally how to use a two-week pilot to decide whether to continue, degrade, or stop.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">According to industry observations, most AI agent incidents are not caused by a single model error, but by operational boundaries that were not clearly defined in advance. Only when the closed loop is built can teams rely less on ad-hoc firefighting.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">1. First Define an Operational State Table: Normal, Degraded, and Circuit-Broken\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Many teams reduce AI agent states to success or failure. In production, failures often expand gradually. A parsing error may first cause missing fields, then trigger tool retries, and finally drive up costs. The meaning of a state table is to let the system know when to continue, when to reduce privileges, and when to stop.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">The State Table Should Answer Four Things\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It must define entry conditions, system actions, recovery conditions, and responsible owners. Without these four items, on-duty personnel can only rely on intuition. The table below can be directly adapted into your team’s version.\u003C\u002Fp>\u003Ctable style=\"width:100%;border-collapse:collapse;font-size:13px;margin:16px 0;\">\u003Ctr>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">State\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Entry Condition\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">System Action\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Recovery Condition\u003C\u002Fth>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Normal\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Core tools are available, output validation passes, costs are within baseline, and sensitive operations are authorized.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Low-risk tasks are executed automatically; medium-risk tasks generate suggestions; high-risk tasks are routed for approval.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Evaluation and monitoring continue to pass, and the agent enters the next review cycle.\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Degraded\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Parsing failures, tool errors, or user corrections rise continuously; quality is unstable for a certain type of task; or costs grow abnormally.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">High-risk automatic execution is disabled, while read-only queries, draft suggestions, and human confirmation are retained.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The failure rate falls within the continuous observation window, regression tests pass, and the responsible person approves restoration.\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Circuit-Broken\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Permission boundaries are crossed, critical dependencies are unavailable, output causes business impact, or cost or data risk becomes uncontrollable.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The entry point is disabled, evidence is preserved, the process is switched back to the previous workflow, and the responsible person and affected users are notified.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The root cause has been fixed, the incident retrospective is complete, and a small-scale pilot passes.\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftable>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">There is a practical principle here: degradation is not turning the AI agent into a decoration, but narrowing its capabilities to a range that is explainable, auditable, and recoverable.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">First Clarify Permissions for Three Task Types\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If permission boundaries are not clearly written before launch, problems will appear during operation: the model suggests refunds, customer service directly modifies orders, or external actions are submitted without user confirmation. At minimum, divide tasks into three types:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Can be executed automatically\u003C\u002Fstrong>: knowledge base retrieval, material summarization, initial ticket classification, field-extraction drafts, and read-only status queries. Outputs must be verifiable and must not directly change external state.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Suggestion only\u003C\u002Fstrong>: contract clause modifications, fee explanations, risk scripts, and cross-department process recommendations. Progress is allowed only after human confirmation.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Approval required\u003C\u002Fstrong>: refunds, account permission changes, data deletion, external sending, financial operations, and write actions that affect user rights. The approver, reason, and evidence must be retained.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Key Outputs Must Include Evidence Fields\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Without evidence fields, the group chat will only leave a sentence saying “something went wrong again.” Evidence should help people quickly answer: where the input came from, what the tool returned, whether validation passed, and which layer failed.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_source: ticket_id, doc_id, user_context_version\ntool_calls: retrieval_status, extraction_status, validation_status\noutput_check: schema_passed, safety_passed, business_rule_passed\nfailure_reason: plan_timeout, tool_error, permission_denied, user_correction\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">These fields do not necessarily all need to be shown to users, but they must be present in the logs.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— The state table is the control plane; evidence logs are the microscope.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkROdDNdbMQuptW8ogdgFPnJdzGnbEwAEXKeb6NmUrRcxgqkCV0g9eTkHU57XJbrodBxjX5DDzOvFeOvS5qxic8vAXcL9hTibLu7qk\u002F0?from=appmsg\" alt=\"Illustration · Warm Autumn Hues\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Warm Autumn Hues\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">2. Minimum Monitoring and Evidence Checklist\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not start monitoring by building a large dashboard. First capture five failure types; they are more actionable than overall accuracy.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Start with Five Failure Types\u003C\u002Fh3>\u003Cul>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Parsing failure\u003C\u002Fstrong>: fields are not extracted, JSON does not match the schema, or table rows and columns are misaligned. The problem is often in input format and schema constraints.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Planning timeout\u003C\u002Fstrong>: task decomposition is too long, multi-round retrieval jumps repeatedly, or tool selection is hesitant. The problem is often in task boundaries and loop control.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Tool error\u003C\u002Fstrong>: a dependent API returns an exception, permissions expire, or parameter mapping is wrong. The problem is often in the tool chain and error propagation.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Permission denial\u003C\u002Fstrong>: the AI agent calls an action it should not call, or accesses unauthorized data. The problem is often in policy configuration and the approval chain.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">User correction\u003C\u002Fstrong>: the user explicitly says the answer is irrelevant, the evidence is insufficient, or the suggestion cannot be executed. The problem is often in task value and matching real scenarios.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Build Task-Level Logs\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The goal of logs is replayability. Given a request_id, you should be able to reconstruct which tools were called, which version was used, where it got stuck, and whether human takeover occurred. Fields do not need to be exhaustive, but they need to be stable.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">request_id\ntask_type\nagent_version\nprompt_version\ntool_call_chain\nlatency\ncost\nvalidation_status\nfail_layer\nhuman_takeover_count\nuser_feedback_id\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Set Actionable Thresholds\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Thresholds do not need complex models; first cover changes that can cause incidents. You can use fixed windows or sliding windows, but every threshold must correspond to an action.\u003C\u002Fp>\u003Col>\u003Cli>If the same task fails consecutively up to a preset count, automatically switch to suggestion mode and pause write operations.\u003C\u002Fli>\u003Cli>If denials of sensitive tools surge, freeze related actions and check permission mapping and input sources.\u003C\u002Fli>\u003Cli>If costs rise abnormally compared with the historical baseline, enable caching, rate limiting, and short-answer mode.\u003C\u002Fli>\u003Cli>If human takeover and user corrections rise continuously, trigger regression evaluation and rollback assessment.\u003C\u002Fli>\u003C\u002Fol>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Minimum Monitoring and Evidence Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Each request has a unique request_id;\u003Cbr>☐ Each tool call has status, latency, and error code;\u003Cbr>☐ Each failure can be attributed to one of five layers: input, planning, tool, output, and permission;\u003Cbr>☐ Each degradation switch has a responsible owner;\u003Cbr>☐ Each human takeover has an approval record and evidence package;\u003Cbr>☐ Each version change has an old entry point and rollback path;\u003Cbr>☐ Each alert has a clear action, rather than only sending a notification.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— The meaning of a closed loop is to ensure that the next anomaly no longer relies on ad-hoc firefighting.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRP5WeeqJZQxz7mAH0BoQwyJMKcKP4sOtfgAIljZh9YTMqB0bVF1SKgREkNLmfricnn6MdyYpIeW1g6YRWhZOoflWgeuOuq5lXko\u002F0?from=appmsg\" alt=\"Illustration · Bamboo Shadows by the River\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Bamboo Shadows by the River\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",3491459,"2026-09-29T17:56:37+08:00","2026-09-29T18:08:27.848109+08:00","2026-09-29T21:35:02.049844+08:00",["Island",80],{"key":81,"result":82},"PostCard_J6VZJE8frnTuFgkSyZmdy1L28Nz3U6IhAq3R0H9WSY",{"head":83},{"link":84,"style":85},[],[],["Island",87],{"key":88,"result":89},"PostCard_RvrKxYkuXkpbD77MaDmpwO8eFJX1aWd2cQyVUSdoA",{"head":90},{"link":91,"style":92},[],[],["Island",94],{"key":95,"result":96},"PostCard_Pugn6ocQ4tEtMeBqWpF1UbcMQFVJx689gQ7Db4wQo",{"head":97},{"link":98,"style":99},[],[],1790693843589]