[{"data":1,"prerenderedAt":101},["ShallowReactive",2],{"PortalFooter_QX9Sj6L0jYa0f3sHh8puk8D6pkI69QhELaa1ZMuWxNo":3,"article-es-ai-安全与对齐工程实践-幻觉控制-越狱防护与内容审核":10,"article-body-v-0-0-0":32,"AppImage_9p9lpUVpZy2pD6REnZg0xaYGC1nDQaRwJHjeEcUGY":33,"PortalBreadcrumb_exPgsYbuf6eIe4kQ2vQPfgqaphcgw25zrXUWLm8j10":40,"related-es-article-ai-安全与对齐工程实践-幻觉控制-越狱防护与内容审核":47,"PostCard_a2w4ZbPC2Qsn0vOlO1ap62B7zycplcayE7leDB99dM":80,"PostCard_xcXmSSw9o0dI9EuzgsNuCIvIiJ0gbS8F5HgA2i5yw":87,"PostCard_bPu9Vhxv9zAEdfGtN97KY1CZNjWwniUaZrJPUPeQg":94},["Island",4],{"key":5,"result":6},"PortalFooter_QX9Sj6L0jYa0f3sHh8puk8D6pkI69QhELaa1ZMuWxNo",{"head":7},{"link":8,"style":9},[],[],{"id":11,"locale":12,"slug":13,"type":14,"title":15,"summary":16,"content_html":17,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"published_at":23,"created_at":24,"updated_at":25,"alternates":26},170,"en","ai-安全与对齐工程实践-幻觉控制-越狱防护与内容审核","article","AI Safety and Alignment Engineering in Practice: Hallucination Control, Jailbreak Prevention, and Content Moderation","From an engineering implementation perspective, this article reviews the core issues of large model safety and alignment. It outlines layered defenses, red-team testing, and launch checklists for hallucination control, jailbreak prevention, and content moderation, helping developers build more reliable generative AI applications.","\u003Ch2>Background and Problems\u003C\u002Fh2>\n\u003Cp>Over the past two years, large models have moved from demo tools into production systems such as customer service, R&D assistants, knowledge Q&A, and marketing content generation. Once connected to real business workflows, problems quickly become concrete: a model may invent nonexistent API parameters, produce prohibited content after multiple rounds of prompting, or treat hidden text on a web page as an instruction to execute. For engineering teams, AI safety and alignment are not abstract ethical slogans, but a set of engineering constraints that must be incorporated into requirements, testing, release, and operations.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Ftea_garden.jpg\" alt=\"茶园春色\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>茶园春色\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp>Common risks can be summarized in three categories: first, hallucinations, where the model generates content that sounds plausible but lacks factual basis; second, jailbreaks and prompt injection, where attackers bypass safety policies through phrasing, encoding, role-playing, or external data; third, content compliance risks, including illegal content, violence, discrimination, privacy leaks, and high-risk professional advice. A truly mature system does not assume the model is always correct; instead, it uses multiple layers of mechanisms to make errors discoverable, blockable, and traceable.\u003C\u002Fp>\n\u003Ch2>Core Concepts\u003C\u002Fh2>\n\u003Ch3>Safety, Alignment, and Controllability\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>\u003Cstrong>Safety\u003C\u002Fstrong>: The system avoids producing harmful, illegal, or misleading outputs as much as possible across diverse inputs.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Alignment\u003C\u002Fstrong>: Model behavior remains consistent with user intent, product goals, organizational policies, and social norms.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Controllability\u003C\u002Fstrong>: Developers can define the boundaries of model capabilities, audit key behaviors, and stop or degrade service when anomalies occur.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>The relationship among the three can be understood as follows: alignment addresses what the model should do, safety addresses what the model must not do, and controllability addresses how humans can supervise and correct it. If there are only prompts without permission boundaries, the model can easily be manipulated; if there are only filters without factual grounding, the system may easily block legitimate requests.\u003C\u002Fp>\n\u003Ch3>Hallucination Control\u003C\u002Fh3>\n\u003Cp>Hallucination does not mean the model is lying; it is more like the model continuing to speak with high confidence even when information is insufficient. In engineering practice, control usually comes from three directions: first, reduce unsupported generation, for example by introducing retrieval augmentation, knowledge base citations, and tool queries; second, lower confidence in uncertain answers, for example by requiring the model to distinguish facts, speculation, and unknowns; third, establish verification mechanisms, such as citation validation, answer consistency checks, and human spot checks.\u003C\u002Fp>\n\u003Cp>For fact-based tasks, it is advisable to divide answers into verifiable and unverifiable categories. Verifiable questions should go through retrieval or databases whenever possible, while unverifiable questions should clearly indicate uncertainty. This is more effective than simply telling the model not to hallucinate.\u003C\u002Fp>\n\u003Ch3>Jailbreak Prevention\u003C\u002Fh3>\n\u003Cp>Jailbreaks usually exploit the model’s generalization ability in natural language, for example by pretending to write fiction, debug code, discuss academic topics, or act as a system administrator, or by using Base64, Unicode, multilingual mixing, and other techniques to evade detection. Prompt injection goes a step further by hiding malicious instructions in user input, web page content, document attachments, or tool responses.\u003C\u002Fp>\n\u003Cp>The focus of prevention is not to find a universal keyword list, but to build layered defenses: the input layer identifies high-risk intents, the system layer separates user instructions from system policies, the tool layer restricts permissions for files, network access, databases, and code execution, the output layer performs safety review, and the logging layer preserves auditable evidence.\u003C\u002Fp>\n\u003Ch3>Content Moderation\u003C\u002Fh3>\n\u003Cp>Content moderation is the last line of defense before model outputs reach users. It should not be just a black-box classifier; it should include clear policies: which content must be blocked, which content needs rewriting or downgrading, which content can be allowed with warnings, and which scenarios must be escalated to humans. Moderation policies should match business risk levels. For example, scenarios involving healthcare, legal advice, finance, or minors should be significantly stricter.\u003C\u002Fp>\n\u003Ch2>Practical Steps and Checklists\u003C\u002Fh2>\n\u003Ch3>1. Define Risk Levels First, Then Choose Model Capabilities\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>Clarify whether the application involves funds, healthcare, legal advice, minors, privacy, or automated execution.\u003C\u002Fli>\n\u003Cli>Based on the risk level, decide whether internet access is allowed, whether tool calls are allowed, and whether executable code generation is allowed.\u003C\u002Fli>\n\u003Cli>For high-risk scenarios, set up human review, rate limiting, allowlists, and strong auditing.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>2. Build a Layered Defense Pipeline\u003C\u002Fh3>\n\u003Cp>A common production pipeline includes input cleaning, intent recognition, permission checks, retrieval augmentation, model generation, fact checking, output moderation, and logging. Below is a simplified pseudocode example showing where each checkpoint can be placed.\u003C\u002Fp>\n\u003Cpre>\u003Ccode>def guardrail(prompt, context):\n    if contains_secret(prompt):\n        return deny('Do not submit secrets or sensitive personal information')\n\n    intent = classify_intent(prompt)\n    if intent in ('medical', 'legal', 'finance'):\n        return answer_with_disclaimer(prompt, require_review=True)\n\n    if intent == 'fact':\n        context = retrieve_sources(prompt)\n\n    answer = llm_generate(prompt, context, temperature=0.2)\n\n    if intent == 'fact' and not has_citation(answer):\n        answer = add_uncertainty_note(answer)\n\n    return output_filter(answer)\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>This code is not a complete safety solution, but it illustrates the engineering approach: first handle sensitive information and high-risk intents, then decide whether retrieval is needed based on task type, and finally moderate the output and add uncertainty notes where appropriate.\u003C\u002Fp>\n\u003Ch3>3. Key Actions for Hallucination Control\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>Connect fact-based questions to retrieval, knowledge bases, or structured queries, and require source citations whenever possible.\u003C\u002Fli>\n\u003Cli>Restrict the model from freely improvising when evidence is lacking; configure refusal templates or escalation to humans.\u003C\u002Fli>\n\u003Cli>Use chunked retrieval for long-document Q&A to prevent the model from stitching together incorrect conclusions across sections.\u003C\u002Fli>\n\u003Cli>Build offline evaluation sets, focusing on factual accuracy, citation hit rate, and the reasonableness of refusals.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>4. Jailbreak and Prompt Injection Testing\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>Create test cases involving role-playing, reverse instructions, multi-turn manipulation, encoding bypasses, and multilingual mixing.\u003C\u002Fli>\n\u003Cli>Check whether system prompts can be leaked and whether tool calls can be indirectly controlled by users.\u003C\u002Fli>\n\u003Cli>Isolate external content from web scraping, file parsing, and database responses to prevent indirect injection.\u003C\u002Fli>\n\u003Cli>Run regression red-team testing after every model upgrade, prompt adjustment, or tool change.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>5. Content Moderation and Operational Closed Loop\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>Output moderation should cover illegal content, violence, discrimination, privacy, self-harm, minors, and high-risk professional advice.\u003C\u002Fli>\n\u003Cli>Record reasons for blocked cases, review false positives, and continuously update policies.\u003C\u002Fli>\n\u003Cli>Provide specialized response scripts and human escalation paths for scenarios such as customer service, education, and healthcare.\u003C\u002Fli>\n\u003Cli>Store logs after data masking, and clearly define retention periods and access permissions in accordance with applicable laws and official compliance requirements.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Common Pitfalls and Recommendations\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Relying only on system prompts\u003C\u002Fstrong>: System prompts are important, but they cannot be the only line of defense. Attackers can change model behavior through multi-turn conversations or external content, so permission controls and output moderation are also required.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Oversimplified keyword filtering\u003C\u002Fstrong>: Keyword lists can easily be bypassed with synonyms, pinyin, encoding, or metaphors. Use semantic classification, contextual judgment, and risk tiering.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Attributing hallucinations only to prompts\u003C\u002Fstrong>: Without reliable knowledge sources and evaluation mechanisms, simply asking the model to answer cautiously often has limited effect.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Excessive refusals harming usability\u003C\u002Fstrong>: Safety policies should not block everything. For edge cases, provide explanations, alternatives, or human support to avoid repeatedly rejecting users.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Ignoring indirect injection\u003C\u002Fstrong>: When the model reads web pages, PDFs, emails, or tool responses, external text may carry instructions. Clearly separate data content from executable instructions.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>No continuous evaluation\u003C\u002Fstrong>: Model versions, prompts, knowledge bases, and tools all change. Without regression testing, safety capabilities can quietly degrade over iterations.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Directions for Further Reading\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>OWASP security risk lists for large model applications and generative AI security practices.\u003C\u002Fli>\n\u003Cli>Public governance frameworks such as the NIST AI Risk Management Framework.\u003C\u002Fli>\n\u003Cli>According to publicly available materials, safety alignment and red-team testing methods proposed by institutions such as Google SAIF and Anthropic.\u003C\u002Fli>\n\u003Cli>Alignment technical routes such as RLHF, DPO, and Constitutional AI, along with their applicable boundaries.\u003C\u002Fli>\n\u003Cli>RAG evaluation metrics such as faithfulness, answer relevance, citation accuracy, and refusal rate.\u003C\u002Fli>\n\u003Cli>Automated red-teaming, adversarial prompt generation, and model behavior monitoring toolchains.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Overall, AI safety and alignment are not one-time projects, but ongoing operations. Engineering teams need to bring hallucination control, jailbreak prevention, and content moderation into the same quality system, managing model risk in ways that are testable, observable, and rollback-capable.\u003C\u002Fp>\n\u003Csection class=\"portal-sources\" style=\"margin-top:2em;font-size:14px;line-height:1.7;color:#5c5850;\">\u003Cp style=\"margin:0 0 0.5em;font-weight:600;\">References\u003C\u002Fp>\u003Cul style=\"margin:0;padding-left:1.25em;\">\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.chagushici.com\u002Fzidian\u002Fzuci\u002F%E5%A4%A7\" rel=\"noopener noreferrer\" target=\"_blank\">大组词_大字组词_大的词语 - 汉语词典\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F636078956\" rel=\"noopener noreferrer\" target=\"_blank\">AI安全与对齐: When, Why, What, and How - 知乎\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fnotes.hehugo.com\u002Fresearch\u002Fai\u002Fengineering\u002Fai-safety-alignment-guide\" rel=\"noopener noreferrer\" target=\"_blank\">AI 安全与对齐完全指南 | 何雨果的知识库\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F666115472\" rel=\"noopener noreferrer\" target=\"_blank\">AI安全前沿 #1 | AI安全四大抓手：对齐、鲁棒性、监测 ...\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fblog.csdn.net\u002Fmieshizhishou\u002Farticle\u002Fdetails\u002F140318746\" rel=\"noopener noreferrer\" target=\"_blank\">【有啥问啥】LLM大模型应用中的安全对齐的简单理解 ...\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>","","https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Ftea_garden.jpg","Guanshan Academy","ai","published","2026-09-29T14:25:47.577226+08:00","2026-09-29T14:30:35.364115+08:00","2026-09-29T21:35:03.202401+08:00",[27,30],{"locale":28,"path":29},"zh","\u002Farticles\u002Fai-安全与对齐工程实践-幻觉控制-越狱防护与内容审核",{"locale":12,"path":31},"\u002Fen\u002Farticles\u002Fai-安全与对齐工程实践-幻觉控制-越狱防护与内容审核","\u003Ch2>Background and Problems\u003C\u002Fh2>\n\u003Cp>Over the past two years, large models have moved from demo tools into production systems such as customer service, R&amp;D assistants, knowledge Q&amp;A, and marketing content generation. Once connected to real business workflows, problems quickly become concrete: a model may invent nonexistent API parameters, produce prohibited content after multiple rounds of prompting, or treat hidden text on a web page as an instruction to execute. For engineering teams, AI safety and alignment are not abstract ethical slogans, but a set of engineering constraints that must be incorporated into requirements, testing, release, and operations.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Ftea_garden.jpg\" alt=\"茶园春色\" loading=\"lazy\" decoding=\"async\">\u003Cfigcaption>茶园春色\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp>Common risks can be summarized in three categories: first, hallucinations, where the model generates content that sounds plausible but lacks factual basis; second, jailbreaks and prompt injection, where attackers bypass safety policies through phrasing, encoding, role-playing, or external data; third, content compliance risks, including illegal content, violence, discrimination, privacy leaks, and high-risk professional advice. A truly mature system does not assume the model is always correct; instead, it uses multiple layers of mechanisms to make errors discoverable, blockable, and traceable.\u003C\u002Fp>\n\u003Ch2>Core Concepts\u003C\u002Fh2>\n\u003Ch3>Safety, Alignment, and Controllability\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>\u003Cstrong>Safety\u003C\u002Fstrong>: The system avoids producing harmful, illegal, or misleading outputs as much as possible across diverse inputs.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Alignment\u003C\u002Fstrong>: Model behavior remains consistent with user intent, product goals, organizational policies, and social norms.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Controllability\u003C\u002Fstrong>: Developers can define the boundaries of model capabilities, audit key behaviors, and stop or degrade service when anomalies occur.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>The relationship among the three can be understood as follows: alignment addresses what the model should do, safety addresses what the model must not do, and controllability addresses how humans can supervise and correct it. If there are only prompts without permission boundaries, the model can easily be manipulated; if there are only filters without factual grounding, the system may easily block legitimate requests.\u003C\u002Fp>\n\u003Ch3>Hallucination Control\u003C\u002Fh3>\n\u003Cp>Hallucination does not mean the model is lying; it is more like the model continuing to speak with high confidence even when information is insufficient. In engineering practice, control usually comes from three directions: first, reduce unsupported generation, for example by introducing retrieval augmentation, knowledge base citations, and tool queries; second, lower confidence in uncertain answers, for example by requiring the model to distinguish facts, speculation, and unknowns; third, establish verification mechanisms, such as citation validation, answer consistency checks, and human spot checks.\u003C\u002Fp>\n\u003Cp>For fact-based tasks, it is advisable to divide answers into verifiable and unverifiable categories. Verifiable questions should go through retrieval or databases whenever possible, while unverifiable questions should clearly indicate uncertainty. This is more effective than simply telling the model not to hallucinate.\u003C\u002Fp>\n\u003Ch3>Jailbreak Prevention\u003C\u002Fh3>\n\u003Cp>Jailbreaks usually exploit the model’s generalization ability in natural language, for example by pretending to write fiction, debug code, discuss academic topics, or act as a system administrator, or by using Base64, Unicode, multilingual mixing, and other techniques to evade detection. Prompt injection goes a step further by hiding malicious instructions in user input, web page content, document attachments, or tool responses.\u003C\u002Fp>\n\u003Cp>The focus of prevention is not to find a universal keyword list, but to build layered defenses: the input layer identifies high-risk intents, the system layer separates user instructions from system policies, the tool layer restricts permissions for files, network access, databases, and code execution, the output layer performs safety review, and the logging layer preserves auditable evidence.\u003C\u002Fp>\n\u003Ch3>Content Moderation\u003C\u002Fh3>\n\u003Cp>Content moderation is the last line of defense before model outputs reach users. It should not be just a black-box classifier; it should include clear policies: which content must be blocked, which content needs rewriting or downgrading, which content can be allowed with warnings, and which scenarios must be escalated to humans. Moderation policies should match business risk levels. For example, scenarios involving healthcare, legal advice, finance, or minors should be significantly stricter.\u003C\u002Fp>\n\u003Ch2>Practical Steps and Checklists\u003C\u002Fh2>\n\u003Ch3>1. Define Risk Levels First, Then Choose Model Capabilities\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>Clarify whether the application involves funds, healthcare, legal advice, minors, privacy, or automated execution.\u003C\u002Fli>\n\u003Cli>Based on the risk level, decide whether internet access is allowed, whether tool calls are allowed, and whether executable code generation is allowed.\u003C\u002Fli>\n\u003Cli>For high-risk scenarios, set up human review, rate limiting, allowlists, and strong auditing.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>2. Build a Layered Defense Pipeline\u003C\u002Fh3>\n\u003Cp>A common production pipeline includes input cleaning, intent recognition, permission checks, retrieval augmentation, model generation, fact checking, output moderation, and logging. Below is a simplified pseudocode example showing where each checkpoint can be placed.\u003C\u002Fp>\n\u003Cpre>\u003Ccode>def guardrail(prompt, context):\n    if contains_secret(prompt):\n        return deny('Do not submit secrets or sensitive personal information')\n\n    intent = classify_intent(prompt)\n    if intent in ('medical', 'legal', 'finance'):\n        return answer_with_disclaimer(prompt, require_review=True)\n\n    if intent == 'fact':\n        context = retrieve_sources(prompt)\n\n    answer = llm_generate(prompt, context, temperature=0.2)\n\n    if intent == 'fact' and not has_citation(answer):\n        answer = add_uncertainty_note(answer)\n\n    return output_filter(answer)\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>This code is not a complete safety solution, but it illustrates the engineering approach: first handle sensitive information and high-risk intents, then decide whether retrieval is needed based on task type, and finally moderate the output and add uncertainty notes where appropriate.\u003C\u002Fp>\n\u003Ch3>3. Key Actions for Hallucination Control\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>Connect fact-based questions to retrieval, knowledge bases, or structured queries, and require source citations whenever possible.\u003C\u002Fli>\n\u003Cli>Restrict the model from freely improvising when evidence is lacking; configure refusal templates or escalation to humans.\u003C\u002Fli>\n\u003Cli>Use chunked retrieval for long-document Q&amp;A to prevent the model from stitching together incorrect conclusions across sections.\u003C\u002Fli>\n\u003Cli>Build offline evaluation sets, focusing on factual accuracy, citation hit rate, and the reasonableness of refusals.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>4. Jailbreak and Prompt Injection Testing\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>Create test cases involving role-playing, reverse instructions, multi-turn manipulation, encoding bypasses, and multilingual mixing.\u003C\u002Fli>\n\u003Cli>Check whether system prompts can be leaked and whether tool calls can be indirectly controlled by users.\u003C\u002Fli>\n\u003Cli>Isolate external content from web scraping, file parsing, and database responses to prevent indirect injection.\u003C\u002Fli>\n\u003Cli>Run regression red-team testing after every model upgrade, prompt adjustment, or tool change.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>5. Content Moderation and Operational Closed Loop\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>Output moderation should cover illegal content, violence, discrimination, privacy, self-harm, minors, and high-risk professional advice.\u003C\u002Fli>\n\u003Cli>Record reasons for blocked cases, review false positives, and continuously update policies.\u003C\u002Fli>\n\u003Cli>Provide specialized response scripts and human escalation paths for scenarios such as customer service, education, and healthcare.\u003C\u002Fli>\n\u003Cli>Store logs after data masking, and clearly define retention periods and access permissions in accordance with applicable laws and official compliance requirements.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Common Pitfalls and Recommendations\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Relying only on system prompts\u003C\u002Fstrong>: System prompts are important, but they cannot be the only line of defense. Attackers can change model behavior through multi-turn conversations or external content, so permission controls and output moderation are also required.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Oversimplified keyword filtering\u003C\u002Fstrong>: Keyword lists can easily be bypassed with synonyms, pinyin, encoding, or metaphors. Use semantic classification, contextual judgment, and risk tiering.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Attributing hallucinations only to prompts\u003C\u002Fstrong>: Without reliable knowledge sources and evaluation mechanisms, simply asking the model to answer cautiously often has limited effect.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Excessive refusals harming usability\u003C\u002Fstrong>: Safety policies should not block everything. For edge cases, provide explanations, alternatives, or human support to avoid repeatedly rejecting users.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Ignoring indirect injection\u003C\u002Fstrong>: When the model reads web pages, PDFs, emails, or tool responses, external text may carry instructions. Clearly separate data content from executable instructions.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>No continuous evaluation\u003C\u002Fstrong>: Model versions, prompts, knowledge bases, and tools all change. Without regression testing, safety capabilities can quietly degrade over iterations.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Directions for Further Reading\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>OWASP security risk lists for large model applications and generative AI security practices.\u003C\u002Fli>\n\u003Cli>Public governance frameworks such as the NIST AI Risk Management Framework.\u003C\u002Fli>\n\u003Cli>According to publicly available materials, safety alignment and red-team testing methods proposed by institutions such as Google SAIF and Anthropic.\u003C\u002Fli>\n\u003Cli>Alignment technical routes such as RLHF, DPO, and Constitutional AI, along with their applicable boundaries.\u003C\u002Fli>\n\u003Cli>RAG evaluation metrics such as faithfulness, answer relevance, citation accuracy, and refusal rate.\u003C\u002Fli>\n\u003Cli>Automated red-teaming, adversarial prompt generation, and model behavior monitoring toolchains.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Overall, AI safety and alignment are not one-time projects, but ongoing operations. Engineering teams need to bring hallucination control, jailbreak prevention, and content moderation into the same quality system, managing model risk in ways that are testable, observable, and rollback-capable.\u003C\u002Fp>\n\u003Csection class=\"portal-sources\" style=\"margin-top:2em;font-size:14px;line-height:1.7;color:#5c5850;\">\u003Cp style=\"margin:0 0 0.5em;font-weight:600;\">References\u003C\u002Fp>\u003Cul style=\"margin:0;padding-left:1.25em;\">\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.chagushici.com\u002Fzidian\u002Fzuci\u002F%E5%A4%A7\" rel=\"noopener noreferrer\">大组词_大字组词_大的词语 - 汉语词典\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F636078956\" rel=\"noopener noreferrer\">AI安全与对齐: When, Why, What, and How - 知乎\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fnotes.hehugo.com\u002Fresearch\u002Fai\u002Fengineering\u002Fai-safety-alignment-guide\" rel=\"noopener noreferrer\">AI 安全与对齐完全指南 | 何雨果的知识库\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F666115472\" rel=\"noopener noreferrer\">AI安全前沿 #1 | AI安全四大抓手：对齐、鲁棒性、监测 ...\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fblog.csdn.net\u002Fmieshizhishou\u002Farticle\u002Fdetails\u002F140318746\" rel=\"noopener noreferrer\">【有啥问啥】LLM大模型应用中的安全对齐的简单理解 ...\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",["Island",34],{"key":35,"result":36},"AppImage_9p9lpUVpZy2pD6REnZg0xaYGC1nDQaRwJHjeEcUGY",{"head":37},{"link":38,"style":39},[],[],["Island",41],{"key":42,"result":43},"PortalBreadcrumb_exPgsYbuf6eIe4kQ2vQPfgqaphcgw25zrXUWLm8j10",{"head":44},{"link":45,"style":46},[],[],[48,60,70],{"id":49,"locale":12,"slug":50,"type":14,"title":51,"summary":52,"content_html":53,"video_url":18,"cover_url":54,"author":55,"category":21,"status":22,"source_job_id":56,"published_at":57,"created_at":58,"updated_at":59},4592,"agent试点验收-七类失败复盘表","Agent Pilot Acceptance: A Postmortem Table for Seven Failure Types","After reading, you will get a seven-category failure attribution table, three acceptance checklists, and autonomy thresholds. Follow the steps to run a read-only pilot first, then decide which actions to automate.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will get a seven-category failure attribution table, three acceptance checklists, and autonomy thresholds. Follow the steps to run a read-only pilot first, then decide which actions should be executed automatically.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text: about 3,633 characters · estimated reading time: 13 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, many teams have placed AI agents into real workflows: customer support tickets, operations inspections, data organization, internal approvals. The demo runs smoothly, but problems appear once multiple people, multiple systems, and long-running execution are involved. On the surface, performance seems unstable; the real blocker is that failure signals have not been broken down.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article provides an acceptance method: a seven-category failure postmortem table, three checklists, a three-day read-only pilot, and staged delegation. You can follow the steps to run a read-only pilot first, then decide which actions to automate.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">I. Make Failure Explicit First: Seven Signals Determine Whether to Continue\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">1. Task Understanding Drift: It Completes a Different Task\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If the user asks the agent to organize customer complaints and generate reply drafts, it may only classify them. If the user asks it to fill in fields, it may rewrite the entire copy. The root cause of drift usually lies in goal decomposition, not the model itself.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can monitor first-pass success rate, human intervention rate, and the proportion of low-quality outputs. If the same intent frequently goes off track, first improve the task template and acceptance fields before considering prompts.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">2. Context Loss: Key Evidence Does Not Reach the Next Step\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">In multi-turn tasks, earlier constraints, customer tier, budget definitions, and historical handling conclusions may be dropped by the time tools are called. The result may look reasonable, but it cannot withstand review.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Observe context reference count, key-field accuracy, and task duration. If repeated queries and repeated confirmations increase, the pipeline is not passing evidence downstream.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">3. Tool Overreach: Actions That Should Not Be Taken Are Taken\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">A read API becomes a write API, a draft becomes a sent message, and a query becomes a deletion. This kind of issue is more dangerous than a wrong answer because it directly changes external state.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Record the number of sensitive-action hits, tool-call input parameters, caller identity, and target system. When overreach is found, first narrow the tool whitelist rather than adding a sentence saying “do not delete.”\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">4. Unverifiable Results: The Output Looks Good, but Cannot Be Traced\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The agent gives a conclusion but leaves no source of evidence, calculation definition, tool return value, or confidence indication. Business colleagues do not dare to sign off, and engineering colleagues cannot reproduce it.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Acceptance should examine key-field accuracy, completeness of the evidence chain, and whether users accept the result. Without verifiable fields, actions should not be handed to automated processes.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">5. Cost Runaway: One Task Turns into a Loop\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Repeated retrieval, repeated summarization, multiple model calls, and multiple tool requests cause the cost per task to keep rising. If the team only watches tokens, it can miss latency, retries, and manual review.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Track repeated call count, task duration, error recovery time, and per-task cost trend. Cost anomalies often mean the workflow has a loop; break the loop before scaling.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">6. Missing Approval: Automated Actions Bypass Humans\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">High-risk actions have no approval point, or approval is merely a formality. When something goes wrong, the chain of responsibility breaks, and the postmortem cannot recover why it was approved.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Check the missing-approval rate, number of external write actions, and retention period for approval records. Approval is not adding a button; it is binding each high-risk action to a person and a reason.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">7. Missing Rollback: Changes Cannot Be Reverted\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Editing documents, sending messages, updating tickets, and adjusting configurations—after failure, there is no rollback ID, no compensating action, and no fallback template. One success may conceal the next incident.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Record the rollback owner, error recovery time, idempotency, and audit ID. If write operations cannot be rolled back, allow them only in low-impact pilot scopes.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— First collect runtime logs, then determine whether the problem lies in the model, prompt, tools, or workflow.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not immediately switch frameworks, expand permissions, or add prompts. First make failures observable; only then does the team have a basis for discussion.\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkROMdibe5hjUkfGygfhy7nUib2bicXAV0oicfMrpdCPHs1NaYib56OA6H166X6IVxU9xib3DlnNuy5icy5bxiaONJ9iaQ9yGpAoR7bgTjfJ0\u002F0?from=appmsg\" alt=\"Illustration · Lakeside Dusk\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Lakeside Dusk\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">II. Define Acceptance Criteria: Replace “Looks Usable” with Three Checklists\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">1. Business Acceptance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The business side asks only whether the result can be used. You can fix five items: task goal completion rate, key-field accuracy, whether users accept the result, proportion of low-quality outputs, and manual review ratio. Each item must specify its definition.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For example, key customer-complaint fields include customer identity, request type, responsible team, and handling deadline. If one is missing, the task is not complete.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">2. Engineering Acceptance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The engineering side must be able to reproduce, track, and stop losses. You can fix: timeout rate, retry count, idempotency, log-field completeness, dependency-service error codes, and call-chain latency distribution.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Logs must include at least \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">run_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">agent_version\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_name\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_hash\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">output_hash\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">duration\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">error_code\u003C\u002Fcode>.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">event: agent.plan\u003Cbr>task_id: T-1024\u003Cbr>run_id: R-77\u003Cbr>user_intent: Organize customer complaints and generate reply drafts\u003Cbr>planned_steps: retrieve, summarize, draft\u003Cbr>tool_call: search_tickets status=open\u003Cbr>risk_level: read_only\u003Cbr>approval_required: false\u003Cbr>idempotency_key: T-1024-R-77-search-001\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This log is not for appearance; it is for attribution after the three-day pilot. If there are only results and no process, the postmortem becomes guesswork.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">3. Security and Compliance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The security focus is not model capability, but action boundaries. You need to check: exposure of sensitive data, external write actions, approvals for deletion\u002Fpayment\u002Fpublishing operations, retention period for audit records, and rollback owner.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Customer-facing outbound messages, fund changes, production configuration changes, and record deletion should not be automatically executed by default. Even if the business is pressing, run them first in a sandbox or shadow system.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Actionable Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Establish a three-key log of \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">run_id\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">agent_version\u003C\u002Fcode> to ensure every execution is traceable.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Record an input summary, output summary, duration, error code, and idempotency key for each tool call.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Place external writes, deletions, payments, publishing, and customer-facing outbound actions on a default block list.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Write clear definitions, criteria, and reviewers for business acceptance fields; do not accept vague evaluations.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Set up daily … for failure samples.\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>","https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg","进学斋",3491491,"2026-09-29T20:53:49+08:00","2026-09-29T21:12:09.724901+08:00","2026-09-29T21:35:02.881325+08:00",{"id":61,"locale":12,"slug":62,"type":14,"title":63,"summary":64,"content_html":65,"video_url":18,"cover_url":54,"author":55,"category":21,"status":22,"source_job_id":66,"published_at":67,"created_at":68,"updated_at":69},4590,"agent观测回滚-落地实战清单","Agent Observability and Rollback: A Practical Implementation Checklist","After reading, you will get a checklist of runtime observability fields for agents and an error-tier rollback table: start with low-traffic logging, then gradually grant more authority.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will get a checklist of runtime observability fields for agents and an error-tier rollback table: start with low-traffic logging, then gradually grant more authority.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text: about 3,580 characters · Reading time: about 12 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Several recent agent projects have entered the runtime phase, and the problems are becoming more complex: the model's answers look good, but the workflow is unstable. Issues range from missing prompts to excessive permissions, high costs, and unnoticed errors.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article focuses only on post-launch runtime data governance. You will get four observability dashboards, three levels of rollback actions, and a seven-day implementation plan. Start by recording low-volume traffic, then gradually grant more authority.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">1. Implementation Bottlenecks: From 'The Model Can Answer' to 'The System Is Manageable'\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">After connecting agents to tools, many teams assume that 'being able to call a tool' means 'being ready for production.' In reality, failures rarely remain in the answer text; they more often appear in broken tool-call chains, permission overreach, uncontrolled costs, and unmanaged human takeover.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For example, a query tool returns an empty value, and the agent continues to generate a conclusion; a write-back API times out, causing a task to be executed repeatedly; a sensitive operation is performed without approval and directly enters a production account. These problems cannot be fundamentally solved by ad hoc prompt edits.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">What must be done during runtime is to turn every execution into a traceable object. First, define stable execution: tasks must be traceable, errors must be stoppable, and costs must be rollback-capable.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Traceable: Every Task Must Have an Identity\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Without \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode> and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">trace_id\u003C\u002Fcode>, agent calls are like anonymous requests. You do not know which role, workflow, or user batch they belong to.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Each call should record at least: task identity, scene, input digest, prompt version, model version, start time, and status.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Stoppable: Errors Must Be Tiered\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Not all errors should be retried. Parameter validation failures can prompt the user to make corrections; repeated tool failures should trigger circuit breaking; operations without approved permissions should stop immediately.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Stop conditions should be written into the system, not left to operations staff watching dashboards.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Rollbackable: Actions Need Exit Paths\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If an agent only generates text, rollback is simple. Once it writes to a database, sends messages, or modifies tickets, you must know where the rollback point is.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Being rollback-capable does not mean undoing everything every time; it means ensuring that errors do not spread.\u003C\u002Fp>\u003Cp align=\"center\" style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Only when it is visible can it be managed.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkRNicSJdZqVvCiasq28Q9qpTaIzccpTIdMFMV6CG1VdagjPZLT84B1Yicz5ppSlAZAShzH8OIJJ570JSEorQEyRGSLus3EPRJJYnLM\u002F0?from=appmsg\" alt=\"配图·山谷柔绿\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Valley in Soft Green\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">2. Observability Model: Four Data Dashboards\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not pile all logs into a single table. During agent runtime, you should split monitoring into at least four dashboards: requests, tools, cost and quality, and audit. Each dashboard answers a different question.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Request Dashboard: Whether a Task Was Triggered Correctly\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It records who the request came from, what task it belongs to, and why it was triggered. It is useful for investigating why this agent started running.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can start by collecting a set of fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">scene\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_digest\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">prompt_version\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">status\u003C\u002Fcode>. The number of fields matters less than being able to link them together.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">trace_id=agent-ticket-001\ntask_id=ticket-8827\nscene=ticket_first_reply\ninput_digest=order_delay_customer_question\nprompt_version=ticket_reply_v3\nstart_time=2026-09-29T10:20:00\nstatus=success\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Tool Dashboard: Where the Chain Breaks\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The tool dashboard is more important than the answer dashboard. It records call order, return codes, latency, and failure points.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It is best to break it down by step: step 1 queries the order, step 2 reads the knowledge base, step 3 generates a draft, and step 4 writes back to the ticket. If step 3 has no write-back, the problem is in the tool layer, not the model layer.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">step_no\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_name\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">args_digest\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">return_code\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">latency_ms\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">retry_count\u003C\u002Fcode>.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Cost and Quality Dashboard: Whether It Is Worth Continuing to Grant Authority\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Agents are not better simply because they are more automated; you need to calculate the costs. Token usage, tool API fees, manual takeover rate, and repeat-question rate all determine whether traffic can be expanded.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For quality metrics, do not look only at accuracy. You can also examine first-time completion rate, number of manual corrections, number of repeated triggers, and whether the result ultimately enters the business system.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tokens_in\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tokens_out\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_cost_estimate\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">manual_takeover_reason\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">repeat_question_count\u003C\u002Fcode>.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Audit Dashboard: Who Approved Which Action\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">When actions involve ticket closure, customer notifications, refund requests, or inventory adjustments, auditing must be a separate dashboard.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It records sensitive operations, approvers, execution results, and whether rollback occurred. Audit is not for the model; it is for the process.\u003C\u002Fp>\u003Cp align=\"center\" style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Rollback is not a step backward; it preserves the right to continue running.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRMBOBM1DicBljibhC3e83iaaT74NwgoO06223ibzZsRbHtgsXfCgp4pEpRIZX9A2WOh9PY4pW622DbAYVMfFWvjXzznsnYqkMvPwGA\u002F0?from=appmsg\" alt=\"配图·静湖云烟\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Still Lake and Misty Clouds\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">3. Rollback Mechanism: Three Action Levels\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Many teams rely on gut feeling for rollback, which is dangerous. Real production deployment requires three levels of action: read-only observation, shadow execution, and minimal closed loop. Each level has permission boundaries and trigger conditions.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Read-Only Observation: Record First, Do Not Touch the Business\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Read-only observation is suitable for the first week of a new scenario. The agent can plan, query, and generate drafts, but it cannot write to production databases or send external messages.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The goal is not to replace humans, but to obtain evidence of the execution chain. You can...\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",3491490,"2026-09-29T20:49:00+08:00","2026-09-29T21:00:41.908764+08:00","2026-09-29T21:35:01.807382+08:00",{"id":71,"locale":12,"slug":72,"type":14,"title":73,"summary":74,"content_html":75,"video_url":18,"cover_url":54,"author":55,"category":21,"status":22,"source_job_id":76,"published_at":77,"created_at":78,"updated_at":79},186,"agent落地-运行闭环七步法","AI Agent Deployment: The Seven-Step Operational Closed Loop","After reading, you will be able to build a post-launch closed loop for AI agent monitoring, alerting, human takeover, rollback, and evaluation, and use a two-week pilot to decide whether to continue, degrade, or stop.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will be able to build a post-launch closed loop for AI agent monitoring, alerting, human takeover, rollback, and evaluation, and use a two-week pilot to decide whether to continue, degrade, or stop.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text about 3,943 characters · Estimated reading time 14 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, while reviewing several teams that deliver AI agent implementations, I found that their on-site demos were mostly impressive: they could summarize, query databases, and fill in fields. But once connected to real requests, problems appeared: tool timeouts, mistaken permission denials, output drift, and no clear owner for human takeover. This is not because the prompts are not flashy enough; it is because the operational closed loop has not been established.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article focuses only on the post-launch operational closed loop: how to define normal, degraded, and circuit-breaking states; how to replay failures; what minimum viable monitoring looks like; how to prepare human takeover, regression evaluation, and rollback packages; and finally how to use a two-week pilot to decide whether to continue, degrade, or stop.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">According to industry observations, most AI agent incidents are not caused by a single model error, but by operational boundaries that were not clearly defined in advance. Only when the closed loop is built can teams rely less on ad-hoc firefighting.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">1. First Define an Operational State Table: Normal, Degraded, and Circuit-Broken\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Many teams reduce AI agent states to success or failure. In production, failures often expand gradually. A parsing error may first cause missing fields, then trigger tool retries, and finally drive up costs. The meaning of a state table is to let the system know when to continue, when to reduce privileges, and when to stop.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">The State Table Should Answer Four Things\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It must define entry conditions, system actions, recovery conditions, and responsible owners. Without these four items, on-duty personnel can only rely on intuition. The table below can be directly adapted into your team’s version.\u003C\u002Fp>\u003Ctable style=\"width:100%;border-collapse:collapse;font-size:13px;margin:16px 0;\">\u003Ctr>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">State\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Entry Condition\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">System Action\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Recovery Condition\u003C\u002Fth>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Normal\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Core tools are available, output validation passes, costs are within baseline, and sensitive operations are authorized.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Low-risk tasks are executed automatically; medium-risk tasks generate suggestions; high-risk tasks are routed for approval.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Evaluation and monitoring continue to pass, and the agent enters the next review cycle.\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Degraded\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Parsing failures, tool errors, or user corrections rise continuously; quality is unstable for a certain type of task; or costs grow abnormally.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">High-risk automatic execution is disabled, while read-only queries, draft suggestions, and human confirmation are retained.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The failure rate falls within the continuous observation window, regression tests pass, and the responsible person approves restoration.\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Circuit-Broken\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Permission boundaries are crossed, critical dependencies are unavailable, output causes business impact, or cost or data risk becomes uncontrollable.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The entry point is disabled, evidence is preserved, the process is switched back to the previous workflow, and the responsible person and affected users are notified.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The root cause has been fixed, the incident retrospective is complete, and a small-scale pilot passes.\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftable>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">There is a practical principle here: degradation is not turning the AI agent into a decoration, but narrowing its capabilities to a range that is explainable, auditable, and recoverable.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">First Clarify Permissions for Three Task Types\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If permission boundaries are not clearly written before launch, problems will appear during operation: the model suggests refunds, customer service directly modifies orders, or external actions are submitted without user confirmation. At minimum, divide tasks into three types:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Can be executed automatically\u003C\u002Fstrong>: knowledge base retrieval, material summarization, initial ticket classification, field-extraction drafts, and read-only status queries. Outputs must be verifiable and must not directly change external state.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Suggestion only\u003C\u002Fstrong>: contract clause modifications, fee explanations, risk scripts, and cross-department process recommendations. Progress is allowed only after human confirmation.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Approval required\u003C\u002Fstrong>: refunds, account permission changes, data deletion, external sending, financial operations, and write actions that affect user rights. The approver, reason, and evidence must be retained.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Key Outputs Must Include Evidence Fields\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Without evidence fields, the group chat will only leave a sentence saying “something went wrong again.” Evidence should help people quickly answer: where the input came from, what the tool returned, whether validation passed, and which layer failed.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_source: ticket_id, doc_id, user_context_version\ntool_calls: retrieval_status, extraction_status, validation_status\noutput_check: schema_passed, safety_passed, business_rule_passed\nfailure_reason: plan_timeout, tool_error, permission_denied, user_correction\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">These fields do not necessarily all need to be shown to users, but they must be present in the logs.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— The state table is the control plane; evidence logs are the microscope.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkROdDNdbMQuptW8ogdgFPnJdzGnbEwAEXKeb6NmUrRcxgqkCV0g9eTkHU57XJbrodBxjX5DDzOvFeOvS5qxic8vAXcL9hTibLu7qk\u002F0?from=appmsg\" alt=\"Illustration · Warm Autumn Hues\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Warm Autumn Hues\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">2. Minimum Monitoring and Evidence Checklist\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not start monitoring by building a large dashboard. First capture five failure types; they are more actionable than overall accuracy.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Start with Five Failure Types\u003C\u002Fh3>\u003Cul>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Parsing failure\u003C\u002Fstrong>: fields are not extracted, JSON does not match the schema, or table rows and columns are misaligned. The problem is often in input format and schema constraints.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Planning timeout\u003C\u002Fstrong>: task decomposition is too long, multi-round retrieval jumps repeatedly, or tool selection is hesitant. The problem is often in task boundaries and loop control.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Tool error\u003C\u002Fstrong>: a dependent API returns an exception, permissions expire, or parameter mapping is wrong. The problem is often in the tool chain and error propagation.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Permission denial\u003C\u002Fstrong>: the AI agent calls an action it should not call, or accesses unauthorized data. The problem is often in policy configuration and the approval chain.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">User correction\u003C\u002Fstrong>: the user explicitly says the answer is irrelevant, the evidence is insufficient, or the suggestion cannot be executed. The problem is often in task value and matching real scenarios.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Build Task-Level Logs\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The goal of logs is replayability. Given a request_id, you should be able to reconstruct which tools were called, which version was used, where it got stuck, and whether human takeover occurred. Fields do not need to be exhaustive, but they need to be stable.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">request_id\ntask_type\nagent_version\nprompt_version\ntool_call_chain\nlatency\ncost\nvalidation_status\nfail_layer\nhuman_takeover_count\nuser_feedback_id\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Set Actionable Thresholds\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Thresholds do not need complex models; first cover changes that can cause incidents. You can use fixed windows or sliding windows, but every threshold must correspond to an action.\u003C\u002Fp>\u003Col>\u003Cli>If the same task fails consecutively up to a preset count, automatically switch to suggestion mode and pause write operations.\u003C\u002Fli>\u003Cli>If denials of sensitive tools surge, freeze related actions and check permission mapping and input sources.\u003C\u002Fli>\u003Cli>If costs rise abnormally compared with the historical baseline, enable caching, rate limiting, and short-answer mode.\u003C\u002Fli>\u003Cli>If human takeover and user corrections rise continuously, trigger regression evaluation and rollback assessment.\u003C\u002Fli>\u003C\u002Fol>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Minimum Monitoring and Evidence Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Each request has a unique request_id;\u003Cbr>☐ Each tool call has status, latency, and error code;\u003Cbr>☐ Each failure can be attributed to one of five layers: input, planning, tool, output, and permission;\u003Cbr>☐ Each degradation switch has a responsible owner;\u003Cbr>☐ Each human takeover has an approval record and evidence package;\u003Cbr>☐ Each version change has an old entry point and rollback path;\u003Cbr>☐ Each alert has a clear action, rather than only sending a notification.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— The meaning of a closed loop is to ensure that the next anomaly no longer relies on ad-hoc firefighting.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRP5WeeqJZQxz7mAH0BoQwyJMKcKP4sOtfgAIljZh9YTMqB0bVF1SKgREkNLmfricnn6MdyYpIeW1g6YRWhZOoflWgeuOuq5lXko\u002F0?from=appmsg\" alt=\"Illustration · Bamboo Shadows by the River\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Bamboo Shadows by the River\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",3491459,"2026-09-29T17:56:37+08:00","2026-09-29T18:08:27.848109+08:00","2026-09-29T21:35:02.049844+08:00",["Island",81],{"key":82,"result":83},"PostCard_a2w4ZbPC2Qsn0vOlO1ap62B7zycplcayE7leDB99dM",{"head":84},{"link":85,"style":86},[],[],["Island",88],{"key":89,"result":90},"PostCard_xcXmSSw9o0dI9EuzgsNuCIvIiJ0gbS8F5HgA2i5yw",{"head":91},{"link":92,"style":93},[],[],["Island",95],{"key":96,"result":97},"PostCard_bPu9Vhxv9zAEdfGtN97KY1CZNjWwniUaZrJPUPeQg",{"head":98},{"link":99,"style":100},[],[],1790693749789]