[{"data":1,"prerenderedAt":100},["ShallowReactive",2],{"PortalFooter_QX9Sj6L0jYa0f3sHh8puk8D6pkI69QhELaa1ZMuWxNo":3,"article-ko-大语言模型入门-transformer-上下文窗口与提示词基础":10,"article-body-v-0-0-0":32,"AppImage_I3ayrfwhQS5yb23tg1fq3L7RTlvzEW91Q3BFmgMM":33,"related-ko-article-大语言模型入门-transformer-上下文窗口与提示词基础":40,"PostCard_VOkjeTg8LIxAWmgdIcTKTu3ZI5z6tRcsvpCy9XWAtDM":72,"PortalBreadcrumb_VajXtfdcSg17UMt4NfCPqD6zumg3naTjnRywqDudA":79,"PostCard_DFOAF3wKLSwzefs0XRmebnZByN0Db5wudM9fADuKNA":86,"PostCard_zRDr9crDL9xzLTocHfUkInuXoyHTBCrIQn7CTKZN0":93},["Island",4],{"key":5,"result":6},"PortalFooter_QX9Sj6L0jYa0f3sHh8puk8D6pkI69QhELaa1ZMuWxNo",{"head":7},{"link":8,"style":9},[],[],{"id":11,"locale":12,"slug":13,"type":14,"title":15,"summary":16,"content_html":17,"video_url":18,"cover_url":19,"author":20,"category":21,"status":22,"published_at":23,"created_at":24,"updated_at":25,"alternates":26},168,"en","大语言模型入门-transformer-上下文窗口与提示词基础","article","Getting Started with Large Language Models: Transformers, Context Windows, and Prompt Basics","This article walks developers through the fundamentals of large language models: how Transformers process text with attention, how context windows shape input and output budgets, and how to craft prompts that express tasks consistently. It includes a practical checklist, common pitfalls, and directions for further reading.","\u003Ch2>Background and Problems\u003C\u002Fh2>\u003Cp>When many developers first encounter large language models, they can easily get confused by concepts such as parameter scale, tokens, context windows, temperature, and system prompts. In real implementations, the problem is often not whether the model is smart enough, but whether we have clearly defined the task boundaries: which text can the model see in a single call? Which information must be retained? How should the output format be constrained? How should errors be handled? From an engineering perspective, this article clarifies three fundamental but critical topics: how Transformers work, the limits of context windows, and basic methods for prompt design.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\u003Ch2>Core Concepts\u003C\u002Fh2>\u003Ch3>Transformer: Understanding Sequences with Attention\u003C\u002Fh3>\u003Cp>The Transformer is a deep learning architecture introduced in the 2017 paper \u003Cstrong>Attention Is All You Need\u003C\u002Fstrong>. It has a notable difference from earlier common RNNs and LSTMs: instead of forcing word-by-word processing in time steps, it uses self-attention to observe multiple positions in a sequence simultaneously and compute the relationships between them. This makes training easier to parallelize and better suited for modeling long-range dependencies.\u003C\u002Fp>\u003Cp>In large language models, text is first split into tokens. Tokens may be words, subwords, or symbols, depending on the tokenizer. Then each token is mapped to a vector and augmented with positional encoding so the model knows the order. After that, the vectors pass through multiple layers of attention and feed-forward networks, progressively refining contextual relationships. Many modern generative models adopt a decoder-only architecture, predicting the next token based on existing text. According to public materials, models such as GPT and LLaMA are based on or improve upon the Transformer architecture; refer to official documentation for specific implementations.\u003C\u002Fp>\u003Cp>When understanding Transformers, focus on several key components:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong>Token Embedding\u003C\u002Fstrong>: Converts discrete text into continuous vectors.\u003C\u002Fli>\u003Cli>\u003Cstrong>Positional Encoding\u003C\u002Fstrong>: Adds word-order information; otherwise the model only knows the set of contents, not their sequence.\u003C\u002Fli>\u003Cli>\u003Cstrong>Self-Attention\u003C\u002Fstrong>: Uses Query, Key, and Value to determine which positions are more relevant.\u003C\u002Fli>\u003Cli>\u003Cstrong>Multi-Head Attention\u003C\u002Fstrong>: Uses multiple attention groups to capture relationships from different dimensions.\u003C\u002Fli>\u003Cli>\u003Cstrong>Feed-Forward Networks, Residual Connections, and Layer Normalization\u003C\u002Fstrong>: Improve expressive power and help stabilize training.\u003C\u002Fli>\u003C\u002Ful>\u003Cpre>\u003Ccode>scores = Q @ K.T \u002F sqrt(d_k); weights = softmax(scores); output = weights @ V\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>This pseudocode is not a complete model; it only shows how attention is computed: first use Q and K to obtain relevance scores, then convert them into weights with softmax, and finally compute a weighted sum of V. For engineers, this abstraction is enough to explain why models can select important information based on context.\u003C\u002Fp>\u003Ch3>Context Window: The Range of Tokens a Model Can Reference at Once\u003C\u002Fh3>\u003Cp>The context window is often called context length, max tokens, or maximum context length. It limits the total number of tokens occupied by both input and output in a single request. For example, a window of 8,192 tokens does not mean you can blindly insert 8,192 Chinese characters or English words, because different tokenizers produce different splits. Chinese text, code, tables, and Markdown symbols can all affect token counts.\u003C\u002Fp>\u003Cp>A more robust engineering approach is to treat the context window as a budget rather than a capacity ceiling. System prompts, user input, retrieval results, conversation history, tool outputs, and expected responses should all be included in the budget. For long-document question answering, you usually need chunking, summarization, or retrieval augmentation instead of stuffing the entire document into the model at once.\u003C\u002Fp>\u003Ch3>Prompts: Writing Task Constraints as Model-Executable Input\u003C\u002Fh3>\u003Cp>A prompt is not just a simple question; it is the input specification for a reasoning task. A stable prompt usually includes role, task, context, constraints, output format, and necessary examples. For tasks such as classification, extraction, summarization, and code generation, the closer the prompt is to a clear, testable requirements document, the more stable the model output will be.\u003C\u002Fp>\u003Ch2>Practical Steps or Checklist\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Identify the task type\u003C\u002Fstrong>: First determine whether it is open-ended generation, structured extraction, classification, or code completion. Different tasks have different requirements for temperature and examples.\u003C\u002Fli>\u003Cli>\u003Cstrong>Provide the minimum necessary context\u003C\u002Fstrong>: Keep only the material directly related to the task, and avoid irrelevant text diluting key instructions.\u003C\u002Fli>\u003Cli>\u003Cstrong>Split long text\u003C\u002Fstrong>: If it exceeds the window or approaches the limit, chunk and summarize first, then aggregate or use retrieval augmentation.\u003C\u002Fli>\u003Cli>\u003Cstrong>Fix the output format\u003C\u002Fstrong>: Require output as a list, table, JSON, or fixed fields, and provide field descriptions.\u003C\u002Fli>\u003Cli>\u003Cstrong>Add few-shot examples\u003C\u002Fstrong>: Use one to three examples to calibrate tone, granularity, and edge cases.\u003C\u002Fli>\u003Cli>\u003Cstrong>Reserve space for output\u003C\u002Fstrong>: Do not fill the input completely; leave room for generation and formatting.\u003C\u002Fli>\u003Cli>\u003Cstrong>Run regression tests\u003C\u002Fstrong>: Prepare a set of typical cases and revalidate after every prompt or model version change.\u003C\u002Fli>\u003C\u002Ful>\u003Cp>A simple prompt template can be organized like this:\u003C\u002Fp>\u003Cpre>\u003Ccode>template = 'Role: Technical documentation assistant. Task: Generate a summary from the given text. Requirements: Output 3 key points, each no more than 20 words. Constraints: Do not fabricate facts, do not output irrelevant content. Text: {text}'\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>This example illustrates a structured way to write prompts: role, task, requirements, constraints, and input data are clearly separated, making it easier for programs to assemble and maintain later.\u003C\u002Fp>\u003Ch2>Common Pitfalls and Suggestions\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Treating the context window as permanent memory\u003C\u002Fstrong>: The model cannot see information outside the window. Long conversations require summarization, external storage, or retrieval mechanisms.\u003C\u002Fli>\u003Cli>\u003Cstrong>Confusing tokens with characters\u003C\u002Fstrong>: A Chinese character, an English word, or a code symbol is not necessarily equal to one token. Use the actual tokenizer to count.\u003C\u002Fli>\u003Cli>\u003Cstrong>Losing focus because the prompt is too long\u003C\u002Fstrong>: Place key constraints at the beginning and end, and present them as lists.\u003C\u002Fli>\u003Cli>\u003Cstrong>Only tuning parameters without fixing the input\u003C\u002Fstrong>: Sampling parameters such as temperature and top_p affect randomness, but when output is unstable, first check the prompt, context, and task boundaries.\u003C\u002Fli>\u003Cli>\u003Cstrong>Ignoring hallucinations and validation\u003C\u002Fstrong>: Models may generate fluent but incorrect content. For decisions involving facts, healthcare, legal matters, finance, or production environments, add retrieval, rule validation, human review, or refusal mechanisms.\u003C\u002Fli>\u003Cli>\u003Cstrong>Blindly pursuing larger windows\u003C\u002Fstrong>: Longer context is not always better; cost, latency, and information noise all increase. If a small context can solve the problem stably, it is usually more economical.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Directions for Further Reading\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>The original Transformer paper\u003C\u002Fstrong>: Understand self-attention, encoder, and decoder design.\u003C\u002Fli>\u003Cli>\u003Cstrong>Tokenizers and BPE\u003C\u002Fstrong>: Learn about token counts, vocabulary size, and multilingual processing differences.\u003C\u002Fli>\u003Cli>\u003Cstrong>RAG and long-document processing\u003C\u002Fstrong>: Study retrieval-augmented generation, chunking strategies, and reranking.\u003C\u002Fli>\u003Cli>\u003Cstrong>Structured output\u003C\u002Fstrong>: Explore JSON Schema, function calling, and constrained decoding.\u003C\u002Fli>\u003Cli>\u003Cstrong>Model evaluation\u003C\u002Fstrong>: Build comprehensive metrics covering accuracy, hallucination rate, latency, cost, and user feedback.\u003C\u002Fli>\u003C\u002Ful>\n\u003Csection class=\"portal-sources\" style=\"margin-top:2em;font-size:14px;line-height:1.7;color:#5c5850;\">\u003Cp style=\"margin:0 0 0.5em;font-weight:600;\">References\u003C\u002Fp>\u003Cul style=\"margin:0;padding-left:1.25em;\">\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F338817680\" rel=\"noopener noreferrer\" target=\"_blank\">Transformer模型详解（图解最完整版） - 知乎\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.runoob.com\u002Fpytorch\u002Ftransformer-model.html\" rel=\"noopener noreferrer\" target=\"_blank\">Transformer 模型 - 菜鸟教程\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fblog.csdn.net\u002Fweixin_42475060\u002Farticle\u002Fdetails\u002F121101749\" rel=\"noopener noreferrer\" target=\"_blank\">【超详细】【原理篇&amp;amp;实战篇】一文读懂Transformer-CSDN博客\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F47812375\" rel=\"noopener noreferrer\" target=\"_blank\">[整理] 聊聊 Transformer - 知乎\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.zhihu.com\u002Ftardis\u002Fzm\u002Fart\u002F600773858\" rel=\"noopener noreferrer\" target=\"_blank\">一文了解Transformer全貌（图解Transformer）\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>","","https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg","Guanshan Academy","ai","published","2026-09-29T14:05:48.096495+08:00","2026-09-29T14:10:06.126961+08:00","2026-09-29T22:15:37.517446+08:00",[27,29],{"locale":12,"path":28},"\u002Fen\u002Farticles\u002F大语言模型入门-transformer-上下文窗口与提示词基础",{"locale":30,"path":31},"zh","\u002Farticles\u002F大语言模型入门-transformer-上下文窗口与提示词基础","\u003Ch2>Background and Problems\u003C\u002Fh2>\u003Cp>When many developers first encounter large language models, they can easily get confused by concepts such as parameter scale, tokens, context windows, temperature, and system prompts. In real implementations, the problem is often not whether the model is smart enough, but whether we have clearly defined the task boundaries: which text can the model see in a single call? Which information must be retained? How should the output format be constrained? How should errors be handled? From an engineering perspective, this article clarifies three fundamental but critical topics: how Transformers work, the limits of context windows, and basic methods for prompt design.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\">\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\u003Ch2>Core Concepts\u003C\u002Fh2>\u003Ch3>Transformer: Understanding Sequences with Attention\u003C\u002Fh3>\u003Cp>The Transformer is a deep learning architecture introduced in the 2017 paper \u003Cstrong>Attention Is All You Need\u003C\u002Fstrong>. It has a notable difference from earlier common RNNs and LSTMs: instead of forcing word-by-word processing in time steps, it uses self-attention to observe multiple positions in a sequence simultaneously and compute the relationships between them. This makes training easier to parallelize and better suited for modeling long-range dependencies.\u003C\u002Fp>\u003Cp>In large language models, text is first split into tokens. Tokens may be words, subwords, or symbols, depending on the tokenizer. Then each token is mapped to a vector and augmented with positional encoding so the model knows the order. After that, the vectors pass through multiple layers of attention and feed-forward networks, progressively refining contextual relationships. Many modern generative models adopt a decoder-only architecture, predicting the next token based on existing text. According to public materials, models such as GPT and LLaMA are based on or improve upon the Transformer architecture; refer to official documentation for specific implementations.\u003C\u002Fp>\u003Cp>When understanding Transformers, focus on several key components:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong>Token Embedding\u003C\u002Fstrong>: Converts discrete text into continuous vectors.\u003C\u002Fli>\u003Cli>\u003Cstrong>Positional Encoding\u003C\u002Fstrong>: Adds word-order information; otherwise the model only knows the set of contents, not their sequence.\u003C\u002Fli>\u003Cli>\u003Cstrong>Self-Attention\u003C\u002Fstrong>: Uses Query, Key, and Value to determine which positions are more relevant.\u003C\u002Fli>\u003Cli>\u003Cstrong>Multi-Head Attention\u003C\u002Fstrong>: Uses multiple attention groups to capture relationships from different dimensions.\u003C\u002Fli>\u003Cli>\u003Cstrong>Feed-Forward Networks, Residual Connections, and Layer Normalization\u003C\u002Fstrong>: Improve expressive power and help stabilize training.\u003C\u002Fli>\u003C\u002Ful>\u003Cpre>\u003Ccode>scores = Q @ K.T \u002F sqrt(d_k); weights = softmax(scores); output = weights @ V\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>This pseudocode is not a complete model; it only shows how attention is computed: first use Q and K to obtain relevance scores, then convert them into weights with softmax, and finally compute a weighted sum of V. For engineers, this abstraction is enough to explain why models can select important information based on context.\u003C\u002Fp>\u003Ch3>Context Window: The Range of Tokens a Model Can Reference at Once\u003C\u002Fh3>\u003Cp>The context window is often called context length, max tokens, or maximum context length. It limits the total number of tokens occupied by both input and output in a single request. For example, a window of 8,192 tokens does not mean you can blindly insert 8,192 Chinese characters or English words, because different tokenizers produce different splits. Chinese text, code, tables, and Markdown symbols can all affect token counts.\u003C\u002Fp>\u003Cp>A more robust engineering approach is to treat the context window as a budget rather than a capacity ceiling. System prompts, user input, retrieval results, conversation history, tool outputs, and expected responses should all be included in the budget. For long-document question answering, you usually need chunking, summarization, or retrieval augmentation instead of stuffing the entire document into the model at once.\u003C\u002Fp>\u003Ch3>Prompts: Writing Task Constraints as Model-Executable Input\u003C\u002Fh3>\u003Cp>A prompt is not just a simple question; it is the input specification for a reasoning task. A stable prompt usually includes role, task, context, constraints, output format, and necessary examples. For tasks such as classification, extraction, summarization, and code generation, the closer the prompt is to a clear, testable requirements document, the more stable the model output will be.\u003C\u002Fp>\u003Ch2>Practical Steps or Checklist\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Identify the task type\u003C\u002Fstrong>: First determine whether it is open-ended generation, structured extraction, classification, or code completion. Different tasks have different requirements for temperature and examples.\u003C\u002Fli>\u003Cli>\u003Cstrong>Provide the minimum necessary context\u003C\u002Fstrong>: Keep only the material directly related to the task, and avoid irrelevant text diluting key instructions.\u003C\u002Fli>\u003Cli>\u003Cstrong>Split long text\u003C\u002Fstrong>: If it exceeds the window or approaches the limit, chunk and summarize first, then aggregate or use retrieval augmentation.\u003C\u002Fli>\u003Cli>\u003Cstrong>Fix the output format\u003C\u002Fstrong>: Require output as a list, table, JSON, or fixed fields, and provide field descriptions.\u003C\u002Fli>\u003Cli>\u003Cstrong>Add few-shot examples\u003C\u002Fstrong>: Use one to three examples to calibrate tone, granularity, and edge cases.\u003C\u002Fli>\u003Cli>\u003Cstrong>Reserve space for output\u003C\u002Fstrong>: Do not fill the input completely; leave room for generation and formatting.\u003C\u002Fli>\u003Cli>\u003Cstrong>Run regression tests\u003C\u002Fstrong>: Prepare a set of typical cases and revalidate after every prompt or model version change.\u003C\u002Fli>\u003C\u002Ful>\u003Cp>A simple prompt template can be organized like this:\u003C\u002Fp>\u003Cpre>\u003Ccode>template = 'Role: Technical documentation assistant. Task: Generate a summary from the given text. Requirements: Output 3 key points, each no more than 20 words. Constraints: Do not fabricate facts, do not output irrelevant content. Text: {text}'\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>This example illustrates a structured way to write prompts: role, task, requirements, constraints, and input data are clearly separated, making it easier for programs to assemble and maintain later.\u003C\u002Fp>\u003Ch2>Common Pitfalls and Suggestions\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>Treating the context window as permanent memory\u003C\u002Fstrong>: The model cannot see information outside the window. Long conversations require summarization, external storage, or retrieval mechanisms.\u003C\u002Fli>\u003Cli>\u003Cstrong>Confusing tokens with characters\u003C\u002Fstrong>: A Chinese character, an English word, or a code symbol is not necessarily equal to one token. Use the actual tokenizer to count.\u003C\u002Fli>\u003Cli>\u003Cstrong>Losing focus because the prompt is too long\u003C\u002Fstrong>: Place key constraints at the beginning and end, and present them as lists.\u003C\u002Fli>\u003Cli>\u003Cstrong>Only tuning parameters without fixing the input\u003C\u002Fstrong>: Sampling parameters such as temperature and top_p affect randomness, but when output is unstable, first check the prompt, context, and task boundaries.\u003C\u002Fli>\u003Cli>\u003Cstrong>Ignoring hallucinations and validation\u003C\u002Fstrong>: Models may generate fluent but incorrect content. For decisions involving facts, healthcare, legal matters, finance, or production environments, add retrieval, rule validation, human review, or refusal mechanisms.\u003C\u002Fli>\u003Cli>\u003Cstrong>Blindly pursuing larger windows\u003C\u002Fstrong>: Longer context is not always better; cost, latency, and information noise all increase. If a small context can solve the problem stably, it is usually more economical.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Directions for Further Reading\u003C\u002Fh2>\u003Cul>\u003Cli>\u003Cstrong>The original Transformer paper\u003C\u002Fstrong>: Understand self-attention, encoder, and decoder design.\u003C\u002Fli>\u003Cli>\u003Cstrong>Tokenizers and BPE\u003C\u002Fstrong>: Learn about token counts, vocabulary size, and multilingual processing differences.\u003C\u002Fli>\u003Cli>\u003Cstrong>RAG and long-document processing\u003C\u002Fstrong>: Study retrieval-augmented generation, chunking strategies, and reranking.\u003C\u002Fli>\u003Cli>\u003Cstrong>Structured output\u003C\u002Fstrong>: Explore JSON Schema, function calling, and constrained decoding.\u003C\u002Fli>\u003Cli>\u003Cstrong>Model evaluation\u003C\u002Fstrong>: Build comprehensive metrics covering accuracy, hallucination rate, latency, cost, and user feedback.\u003C\u002Fli>\u003C\u002Ful>\n\u003Csection class=\"portal-sources\" style=\"margin-top:2em;font-size:14px;line-height:1.7;color:#5c5850;\">\u003Cp style=\"margin:0 0 0.5em;font-weight:600;\">References\u003C\u002Fp>\u003Cul style=\"margin:0;padding-left:1.25em;\">\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F338817680\" rel=\"noopener noreferrer\">Transformer模型详解（图解最完整版） - 知乎\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.runoob.com\u002Fpytorch\u002Ftransformer-model.html\" rel=\"noopener noreferrer\">Transformer 模型 - 菜鸟教程\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fblog.csdn.net\u002Fweixin_42475060\u002Farticle\u002Fdetails\u002F121101749\" rel=\"noopener noreferrer\">【超详细】【原理篇&amp;amp;实战篇】一文读懂Transformer-CSDN博客\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fzhuanlan.zhihu.com\u002Fp\u002F47812375\" rel=\"noopener noreferrer\">[整理] 聊聊 Transformer - 知乎\u003C\u002Fa>\u003C\u002Fli>\u003Cli style=\"margin:0.25em 0;\">\u003Ca href=\"https:\u002F\u002Fwww.zhihu.com\u002Ftardis\u002Fzm\u002Fart\u002F600773858\" rel=\"noopener noreferrer\">一文了解Transformer全貌（图解Transformer）\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",["Island",34],{"key":35,"result":36},"AppImage_I3ayrfwhQS5yb23tg1fq3L7RTlvzEW91Q3BFmgMM",{"head":37},{"link":38,"style":39},[],[],[41,52,62],{"id":42,"locale":12,"slug":43,"type":14,"title":44,"summary":45,"content_html":46,"video_url":18,"cover_url":19,"author":47,"category":21,"status":22,"source_job_id":48,"published_at":49,"created_at":50,"updated_at":51},4592,"agent试点验收-七类失败复盘表","Agent Pilot Acceptance: A Postmortem Table for Seven Failure Types","After reading, you will get a seven-category failure attribution table, three acceptance checklists, and autonomy thresholds. Follow the steps to run a read-only pilot first, then decide which actions to automate.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will get a seven-category failure attribution table, three acceptance checklists, and autonomy thresholds. Follow the steps to run a read-only pilot first, then decide which actions should be executed automatically.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text: about 3,633 characters · estimated reading time: 13 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, many teams have placed AI agents into real workflows: customer support tickets, operations inspections, data organization, internal approvals. The demo runs smoothly, but problems appear once multiple people, multiple systems, and long-running execution are involved. On the surface, performance seems unstable; the real blocker is that failure signals have not been broken down.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article provides an acceptance method: a seven-category failure postmortem table, three checklists, a three-day read-only pilot, and staged delegation. You can follow the steps to run a read-only pilot first, then decide which actions to automate.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">I. Make Failure Explicit First: Seven Signals Determine Whether to Continue\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">1. Task Understanding Drift: It Completes a Different Task\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If the user asks the agent to organize customer complaints and generate reply drafts, it may only classify them. If the user asks it to fill in fields, it may rewrite the entire copy. The root cause of drift usually lies in goal decomposition, not the model itself.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can monitor first-pass success rate, human intervention rate, and the proportion of low-quality outputs. If the same intent frequently goes off track, first improve the task template and acceptance fields before considering prompts.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">2. Context Loss: Key Evidence Does Not Reach the Next Step\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">In multi-turn tasks, earlier constraints, customer tier, budget definitions, and historical handling conclusions may be dropped by the time tools are called. The result may look reasonable, but it cannot withstand review.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Observe context reference count, key-field accuracy, and task duration. If repeated queries and repeated confirmations increase, the pipeline is not passing evidence downstream.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">3. Tool Overreach: Actions That Should Not Be Taken Are Taken\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">A read API becomes a write API, a draft becomes a sent message, and a query becomes a deletion. This kind of issue is more dangerous than a wrong answer because it directly changes external state.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Record the number of sensitive-action hits, tool-call input parameters, caller identity, and target system. When overreach is found, first narrow the tool whitelist rather than adding a sentence saying “do not delete.”\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">4. Unverifiable Results: The Output Looks Good, but Cannot Be Traced\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The agent gives a conclusion but leaves no source of evidence, calculation definition, tool return value, or confidence indication. Business colleagues do not dare to sign off, and engineering colleagues cannot reproduce it.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Acceptance should examine key-field accuracy, completeness of the evidence chain, and whether users accept the result. Without verifiable fields, actions should not be handed to automated processes.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">5. Cost Runaway: One Task Turns into a Loop\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Repeated retrieval, repeated summarization, multiple model calls, and multiple tool requests cause the cost per task to keep rising. If the team only watches tokens, it can miss latency, retries, and manual review.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Track repeated call count, task duration, error recovery time, and per-task cost trend. Cost anomalies often mean the workflow has a loop; break the loop before scaling.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">6. Missing Approval: Automated Actions Bypass Humans\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">High-risk actions have no approval point, or approval is merely a formality. When something goes wrong, the chain of responsibility breaks, and the postmortem cannot recover why it was approved.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Check the missing-approval rate, number of external write actions, and retention period for approval records. Approval is not adding a button; it is binding each high-risk action to a person and a reason.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">7. Missing Rollback: Changes Cannot Be Reverted\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Editing documents, sending messages, updating tickets, and adjusting configurations—after failure, there is no rollback ID, no compensating action, and no fallback template. One success may conceal the next incident.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Record the rollback owner, error recovery time, idempotency, and audit ID. If write operations cannot be rolled back, allow them only in low-impact pilot scopes.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— First collect runtime logs, then determine whether the problem lies in the model, prompt, tools, or workflow.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not immediately switch frameworks, expand permissions, or add prompts. First make failures observable; only then does the team have a basis for discussion.\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkROMdibe5hjUkfGygfhy7nUib2bicXAV0oicfMrpdCPHs1NaYib56OA6H166X6IVxU9xib3DlnNuy5icy5bxiaONJ9iaQ9yGpAoR7bgTjfJ0\u002F0?from=appmsg\" alt=\"Illustration · Lakeside Dusk\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Lakeside Dusk\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">II. Define Acceptance Criteria: Replace “Looks Usable” with Three Checklists\u003C\u002Fh2>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">1. Business Acceptance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The business side asks only whether the result can be used. You can fix five items: task goal completion rate, key-field accuracy, whether users accept the result, proportion of low-quality outputs, and manual review ratio. Each item must specify its definition.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For example, key customer-complaint fields include customer identity, request type, responsible team, and handling deadline. If one is missing, the task is not complete.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">2. Engineering Acceptance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The engineering side must be able to reproduce, track, and stop losses. You can fix: timeout rate, retry count, idempotency, log-field completeness, dependency-service error codes, and call-chain latency distribution.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Logs must include at least \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">run_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">agent_version\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_name\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_hash\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">output_hash\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">duration\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">error_code\u003C\u002Fcode>.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">event: agent.plan\u003Cbr>task_id: T-1024\u003Cbr>run_id: R-77\u003Cbr>user_intent: Organize customer complaints and generate reply drafts\u003Cbr>planned_steps: retrieve, summarize, draft\u003Cbr>tool_call: search_tickets status=open\u003Cbr>risk_level: read_only\u003Cbr>approval_required: false\u003Cbr>idempotency_key: T-1024-R-77-search-001\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This log is not for appearance; it is for attribution after the three-day pilot. If there are only results and no process, the postmortem becomes guesswork.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">3. Security and Compliance Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The security focus is not model capability, but action boundaries. You need to check: exposure of sensitive data, external write actions, approvals for deletion\u002Fpayment\u002Fpublishing operations, retention period for audit records, and rollback owner.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Customer-facing outbound messages, fund changes, production configuration changes, and record deletion should not be automatically executed by default. Even if the business is pressing, run them first in a sandbox or shadow system.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Actionable Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Establish a three-key log of \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">run_id\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">agent_version\u003C\u002Fcode> to ensure every execution is traceable.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Record an input summary, output summary, duration, error code, and idempotency key for each tool call.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Place external writes, deletions, payments, publishing, and customer-facing outbound actions on a default block list.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Write clear definitions, criteria, and reviewers for business acceptance fields; do not accept vague evaluations.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Set up daily … for failure samples.\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>","进学斋",3491491,"2026-09-29T20:53:49+08:00","2026-09-29T21:12:09.724901+08:00","2026-09-29T21:35:02.881325+08:00",{"id":53,"locale":12,"slug":54,"type":14,"title":55,"summary":56,"content_html":57,"video_url":18,"cover_url":19,"author":47,"category":21,"status":22,"source_job_id":58,"published_at":59,"created_at":60,"updated_at":61},4590,"agent观测回滚-落地实战清单","Agent Observability and Rollback: A Practical Implementation Checklist","After reading, you will get a checklist of runtime observability fields for agents and an error-tier rollback table: start with low-traffic logging, then gradually grant more authority.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will get a checklist of runtime observability fields for agents and an error-tier rollback table: start with low-traffic logging, then gradually grant more authority.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text: about 3,580 characters · Reading time: about 12 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Several recent agent projects have entered the runtime phase, and the problems are becoming more complex: the model's answers look good, but the workflow is unstable. Issues range from missing prompts to excessive permissions, high costs, and unnoticed errors.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article focuses only on post-launch runtime data governance. You will get four observability dashboards, three levels of rollback actions, and a seven-day implementation plan. Start by recording low-volume traffic, then gradually grant more authority.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">1. Implementation Bottlenecks: From 'The Model Can Answer' to 'The System Is Manageable'\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">After connecting agents to tools, many teams assume that 'being able to call a tool' means 'being ready for production.' In reality, failures rarely remain in the answer text; they more often appear in broken tool-call chains, permission overreach, uncontrolled costs, and unmanaged human takeover.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For example, a query tool returns an empty value, and the agent continues to generate a conclusion; a write-back API times out, causing a task to be executed repeatedly; a sensitive operation is performed without approval and directly enters a production account. These problems cannot be fundamentally solved by ad hoc prompt edits.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">What must be done during runtime is to turn every execution into a traceable object. First, define stable execution: tasks must be traceable, errors must be stoppable, and costs must be rollback-capable.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Traceable: Every Task Must Have an Identity\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Without \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode> and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">trace_id\u003C\u002Fcode>, agent calls are like anonymous requests. You do not know which role, workflow, or user batch they belong to.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Each call should record at least: task identity, scene, input digest, prompt version, model version, start time, and status.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Stoppable: Errors Must Be Tiered\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Not all errors should be retried. Parameter validation failures can prompt the user to make corrections; repeated tool failures should trigger circuit breaking; operations without approved permissions should stop immediately.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Stop conditions should be written into the system, not left to operations staff watching dashboards.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Rollbackable: Actions Need Exit Paths\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If an agent only generates text, rollback is simple. Once it writes to a database, sends messages, or modifies tickets, you must know where the rollback point is.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Being rollback-capable does not mean undoing everything every time; it means ensuring that errors do not spread.\u003C\u002Fp>\u003Cp align=\"center\" style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Only when it is visible can it be managed.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkRNicSJdZqVvCiasq28Q9qpTaIzccpTIdMFMV6CG1VdagjPZLT84B1Yicz5ppSlAZAShzH8OIJJ570JSEorQEyRGSLus3EPRJJYnLM\u002F0?from=appmsg\" alt=\"配图·山谷柔绿\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Valley in Soft Green\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">2. Observability Model: Four Data Dashboards\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not pile all logs into a single table. During agent runtime, you should split monitoring into at least four dashboards: requests, tools, cost and quality, and audit. Each dashboard answers a different question.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Request Dashboard: Whether a Task Was Triggered Correctly\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It records who the request came from, what task it belongs to, and why it was triggered. It is useful for investigating why this agent started running.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">You can start by collecting a set of fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">task_id\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">scene\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_digest\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">prompt_version\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">status\u003C\u002Fcode>. The number of fields matters less than being able to link them together.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">trace_id=agent-ticket-001\ntask_id=ticket-8827\nscene=ticket_first_reply\ninput_digest=order_delay_customer_question\nprompt_version=ticket_reply_v3\nstart_time=2026-09-29T10:20:00\nstatus=success\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Tool Dashboard: Where the Chain Breaks\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The tool dashboard is more important than the answer dashboard. It records call order, return codes, latency, and failure points.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It is best to break it down by step: step 1 queries the order, step 2 reads the knowledge base, step 3 generates a draft, and step 4 writes back to the ticket. If step 3 has no write-back, the problem is in the tool layer, not the model layer.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">step_no\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_name\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">args_digest\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">return_code\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">latency_ms\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">retry_count\u003C\u002Fcode>.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Cost and Quality Dashboard: Whether It Is Worth Continuing to Grant Authority\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Agents are not better simply because they are more automated; you need to calculate the costs. Token usage, tool API fees, manual takeover rate, and repeat-question rate all determine whether traffic can be expanded.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">For quality metrics, do not look only at accuracy. You can also examine first-time completion rate, number of manual corrections, number of repeated triggers, and whether the result ultimately enters the business system.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Fields: \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tokens_in\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tokens_out\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">tool_cost_estimate\u003C\u002Fcode>, \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">manual_takeover_reason\u003C\u002Fcode>, and \u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">repeat_question_count\u003C\u002Fcode>.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Audit Dashboard: Who Approved Which Action\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">When actions involve ticket closure, customer notifications, refund requests, or inventory adjustments, auditing must be a separate dashboard.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It records sensitive operations, approvers, execution results, and whether rollback occurred. Audit is not for the model; it is for the process.\u003C\u002Fp>\u003Cp align=\"center\" style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— Rollback is not a step backward; it preserves the right to continue running.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRMBOBM1DicBljibhC3e83iaaT74NwgoO06223ibzZsRbHtgsXfCgp4pEpRIZX9A2WOh9PY4pW622DbAYVMfFWvjXzznsnYqkMvPwGA\u002F0?from=appmsg\" alt=\"配图·静湖云烟\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Still Lake and Misty Clouds\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">3. Rollback Mechanism: Three Action Levels\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Many teams rely on gut feeling for rollback, which is dangerous. Real production deployment requires three levels of action: read-only observation, shadow execution, and minimal closed loop. Each level has permission boundaries and trigger conditions.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Read-Only Observation: Record First, Do Not Touch the Business\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Read-only observation is suitable for the first week of a new scenario. The agent can plan, query, and generate drafts, but it cannot write to production databases or send external messages.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The goal is not to replace humans, but to obtain evidence of the execution chain. You can...\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",3491490,"2026-09-29T20:49:00+08:00","2026-09-29T21:00:41.908764+08:00","2026-09-29T21:35:01.807382+08:00",{"id":63,"locale":12,"slug":64,"type":14,"title":65,"summary":66,"content_html":67,"video_url":18,"cover_url":19,"author":47,"category":21,"status":22,"source_job_id":68,"published_at":69,"created_at":70,"updated_at":71},186,"agent落地-运行闭环七步法","AI Agent Deployment: The Seven-Step Operational Closed Loop","After reading, you will be able to build a post-launch closed loop for AI agent monitoring, alerting, human takeover, rollback, and evaluation, and use a two-week pilot to decide whether to continue, degrade, or stop.","\u003Csection data-gzh-readable=\"1\" style=\"margin:0;padding:0;\">\u003Cdiv data-gzh-header=\"1\">\u003Cp style=\"font-size:15px;color:#888;text-align:center;margin-top:8px;\">After reading, you will be able to build a post-launch closed loop for AI agent monitoring, alerting, human takeover, rollback, and evaluation, and use a two-week pilot to decide whether to continue, degrade, or stop.\u003C\u002Fp>\n\u003Cfigure class=\"portal-article-figure\">\u003Cimg src=\"https:\u002F\u002Fguanshanshuyuan.cn\u002Fimages\u002Farticles\u002Freading_books.jpg\" alt=\"读书明理\" loading=\"lazy\" decoding=\"async\" \u002F>\u003Cfigcaption>读书明理\u003C\u002Ffigcaption>\u003C\u002Ffigure>\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\u003Cp style=\"font-size:14px;color:#999;text-align:center;margin-top:4px;\">Author: Jinxuezhai · Guanshan Academy\u003C\u002Fp>\u003Cp style=\"font-size:13px;color:#999;text-align:center;margin:0 0 6px 0;\">Full text about 3,943 characters · Estimated reading time 14 minutes\u003C\u002Fp>\u003Chr style=\"border:none;border-top:1px solid #eee;margin:24px 0;\" \u002F>\u003C\u002Fdiv>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Recently, while reviewing several teams that deliver AI agent implementations, I found that their on-site demos were mostly impressive: they could summarize, query databases, and fill in fields. But once connected to real requests, problems appeared: tool timeouts, mistaken permission denials, output drift, and no clear owner for human takeover. This is not because the prompts are not flashy enough; it is because the operational closed loop has not been established.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">This article focuses only on the post-launch operational closed loop: how to define normal, degraded, and circuit-breaking states; how to replay failures; what minimum viable monitoring looks like; how to prepare human takeover, regression evaluation, and rollback packages; and finally how to use a two-week pilot to decide whether to continue, degrade, or stop.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">According to industry observations, most AI agent incidents are not caused by a single model error, but by operational boundaries that were not clearly defined in advance. Only when the closed loop is built can teams rely less on ad-hoc firefighting.\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">1. First Define an Operational State Table: Normal, Degraded, and Circuit-Broken\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Many teams reduce AI agent states to success or failure. In production, failures often expand gradually. A parsing error may first cause missing fields, then trigger tool retries, and finally drive up costs. The meaning of a state table is to let the system know when to continue, when to reduce privileges, and when to stop.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">The State Table Should Answer Four Things\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">It must define entry conditions, system actions, recovery conditions, and responsible owners. Without these four items, on-duty personnel can only rely on intuition. The table below can be directly adapted into your team’s version.\u003C\u002Fp>\u003Ctable style=\"width:100%;border-collapse:collapse;font-size:13px;margin:16px 0;\">\u003Ctr>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">State\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Entry Condition\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">System Action\u003C\u002Fth>\u003Cth style=\"border:1px solid #ddd;padding:8px 10px;background:#f5f8fc;color:#1a1a1a;font-weight:bold;text-align:left;\">Recovery Condition\u003C\u002Fth>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Normal\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Core tools are available, output validation passes, costs are within baseline, and sensitive operations are authorized.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Low-risk tasks are executed automatically; medium-risk tasks generate suggestions; high-risk tasks are routed for approval.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Evaluation and monitoring continue to pass, and the agent enters the next review cycle.\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Degraded\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Parsing failures, tool errors, or user corrections rise continuously; quality is unstable for a certain type of task; or costs grow abnormally.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">High-risk automatic execution is disabled, while read-only queries, draft suggestions, and human confirmation are retained.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The failure rate falls within the continuous observation window, regression tests pass, and the responsible person approves restoration.\u003C\u002Ftd>\u003C\u002Ftr>\u003Ctr>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Circuit-Broken\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">Permission boundaries are crossed, critical dependencies are unavailable, output causes business impact, or cost or data risk becomes uncontrollable.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The entry point is disabled, evidence is preserved, the process is switched back to the previous workflow, and the responsible person and affected users are notified.\u003C\u002Ftd>\u003Ctd style=\"border:1px solid #ddd;padding:8px 10px;vertical-align:top;\">The root cause has been fixed, the incident retrospective is complete, and a small-scale pilot passes.\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftable>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">There is a practical principle here: degradation is not turning the AI agent into a decoration, but narrowing its capabilities to a range that is explainable, auditable, and recoverable.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">First Clarify Permissions for Three Task Types\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">If permission boundaries are not clearly written before launch, problems will appear during operation: the model suggests refunds, customer service directly modifies orders, or external actions are submitted without user confirmation. At minimum, divide tasks into three types:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Can be executed automatically\u003C\u002Fstrong>: knowledge base retrieval, material summarization, initial ticket classification, field-extraction drafts, and read-only status queries. Outputs must be verifiable and must not directly change external state.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Suggestion only\u003C\u002Fstrong>: contract clause modifications, fee explanations, risk scripts, and cross-department process recommendations. Progress is allowed only after human confirmation.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Approval required\u003C\u002Fstrong>: refunds, account permission changes, data deletion, external sending, financial operations, and write actions that affect user rights. The approver, reason, and evidence must be retained.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Key Outputs Must Include Evidence Fields\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Without evidence fields, the group chat will only leave a sentence saying “something went wrong again.” Evidence should help people quickly answer: where the input came from, what the tool returned, whether validation passed, and which layer failed.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">input_source: ticket_id, doc_id, user_context_version\ntool_calls: retrieval_status, extraction_status, validation_status\noutput_check: schema_passed, safety_passed, business_rule_passed\nfailure_reason: plan_timeout, tool_error, permission_denied, user_correction\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">These fields do not necessarily all need to be shown to users, but they must be present in the logs.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— The state table is the control plane; evidence logs are the microscope.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002F4WyfSFibfkROdDNdbMQuptW8ogdgFPnJdzGnbEwAEXKeb6NmUrRcxgqkCV0g9eTkHU57XJbrodBxjX5DDzOvFeOvS5qxic8vAXcL9hTibLu7qk\u002F0?from=appmsg\" alt=\"Illustration · Warm Autumn Hues\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Warm Autumn Hues\u003C\u002Fp>\u003Ch2 style=\"font-size:20px;font-weight:bold;color:#1a1a1a;border-bottom:2px solid #4a90d9;padding-bottom:8px;margin:2em 0 1em;\">2. Minimum Monitoring and Evidence Checklist\u003C\u002Fh2>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Do not start monitoring by building a large dashboard. First capture five failure types; they are more actionable than overall accuracy.\u003C\u002Fp>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Start with Five Failure Types\u003C\u002Fh3>\u003Cul>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Parsing failure\u003C\u002Fstrong>: fields are not extracted, JSON does not match the schema, or table rows and columns are misaligned. The problem is often in input format and schema constraints.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Planning timeout\u003C\u002Fstrong>: task decomposition is too long, multi-round retrieval jumps repeatedly, or tool selection is hesitant. The problem is often in task boundaries and loop control.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Tool error\u003C\u002Fstrong>: a dependent API returns an exception, permissions expire, or parameter mapping is wrong. The problem is often in the tool chain and error propagation.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">Permission denial\u003C\u002Fstrong>: the AI agent calls an action it should not call, or accesses unauthorized data. The problem is often in policy configuration and the approval chain.\u003C\u002Fli>\u003Cli>\u003Cstrong style=\"font-weight:bold;color:#1a1a1a;\">User correction\u003C\u002Fstrong>: the user explicitly says the answer is irrelevant, the evidence is insufficient, or the suggestion cannot be executed. The problem is often in task value and matching real scenarios.\u003C\u002Fli>\u003C\u002Ful>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Build Task-Level Logs\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">The goal of logs is replayability. Given a request_id, you should be able to reconstruct which tools were called, which version was used, where it got stuck, and whether human takeover occurred. Fields do not need to be exhaustive, but they need to be stable.\u003C\u002Fp>\u003Cpre style=\"background:#f5f7fa;padding:15px;border-radius:6px;font-size:14px;line-height:1.6;overflow-x:auto;margin:1em 0;\">\u003Ccode style=\"font-family:'SF Mono',Menlo,Consolas,monospace;color:#333;background:#f5f5f5;padding:2px 6px;font-size:14px;border-radius:3px;\">request_id\ntask_type\nagent_version\nprompt_version\ntool_call_chain\nlatency\ncost\nvalidation_status\nfail_layer\nhuman_takeover_count\nuser_feedback_id\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Set Actionable Thresholds\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">Thresholds do not need complex models; first cover changes that can cause incidents. You can use fixed windows or sliding windows, but every threshold must correspond to an action.\u003C\u002Fp>\u003Col>\u003Cli>If the same task fails consecutively up to a preset count, automatically switch to suggestion mode and pause write operations.\u003C\u002Fli>\u003Cli>If denials of sensitive tools surge, freeze related actions and check permission mapping and input sources.\u003C\u002Fli>\u003Cli>If costs rise abnormally compared with the historical baseline, enable caching, rate limiting, and short-answer mode.\u003C\u002Fli>\u003Cli>If human takeover and user corrections rise continuously, trigger regression evaluation and rollback assessment.\u003C\u002Fli>\u003C\u002Fol>\u003Ch3 style=\"font-size:18px;font-weight:bold;color:#333;margin:1.6em 0 0.8em;\">Minimum Monitoring and Evidence Checklist\u003C\u002Fh3>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">☐ Each request has a unique request_id;\u003Cbr>☐ Each tool call has status, latency, and error code;\u003Cbr>☐ Each failure can be attributed to one of five layers: input, planning, tool, output, and permission;\u003Cbr>☐ Each degradation switch has a responsible owner;\u003Cbr>☐ Each human takeover has an approval record and evidence package;\u003Cbr>☐ Each version change has an old entry point and rollback path;\u003Cbr>☐ Each alert has a clear action, rather than only sending a notification.\u003C\u002Fp>\u003Cp style=\"font-size:16px;line-height:1.85;color:#333;margin:0 0 1em;text-align:justify;\">\u003Cem>— The meaning of a closed loop is to ensure that the next anomaly no longer relies on ad-hoc firefighting.\u003C\u002Fem>\u003C\u002Fp>\u003Cp style=\"text-align:center;margin:16px 0 6px 0;\">\u003Cimg src=\"http:\u002F\u002Fmmbiz.qpic.cn\u002Fsz_mmbiz_jpg\u002F4WyfSFibfkRP5WeeqJZQxz7mAH0BoQwyJMKcKP4sOtfgAIljZh9YTMqB0bVF1SKgREkNLmfricnn6MdyYpIeW1g6YRWhZOoflWgeuOuq5lXko\u002F0?from=appmsg\" alt=\"Illustration · Bamboo Shadows by the River\" style=\"max-width:100%;border-radius:6px;display:block;margin:0 auto;\" \u002F>\u003C\u002Fp>\u003Cp style=\"text-align:center;font-size:12px;color:#bbb;margin:0 0 18px 0;\">Illustration · Bamboo Shadows by the River\u003C\u002Fp>\u003C\u002Fsection>\n\u003Csection class=\"portal-disclaimer\" data-portal-disclaimer=\"1\" style=\"margin-top:2.5em;padding-top:1.25em;border-top:1px solid #e8e4dc;font-size:14px;line-height:1.7;color:#7a756c;\">\u003Cp style=\"margin:0;\">Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.\u003C\u002Fp>\u003Cp style=\"margin:0.75em 0 0;\">Contact: \u003Ca href=\"mailto:chenxj.g@gmail.com\">chenxj.g@gmail.com\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fsection>",3491459,"2026-09-29T17:56:37+08:00","2026-09-29T18:08:27.848109+08:00","2026-09-29T21:35:02.049844+08:00",["Island",73],{"key":74,"result":75},"PostCard_VOkjeTg8LIxAWmgdIcTKTu3ZI5z6tRcsvpCy9XWAtDM",{"head":76},{"link":77,"style":78},[],[],["Island",80],{"key":81,"result":82},"PortalBreadcrumb_VajXtfdcSg17UMt4NfCPqD6zumg3naTjnRywqDudA",{"head":83},{"link":84,"style":85},[],[],["Island",87],{"key":88,"result":89},"PostCard_DFOAF3wKLSwzefs0XRmebnZByN0Db5wudM9fADuKNA",{"head":90},{"link":91,"style":92},[],[],["Island",94],{"key":95,"result":96},"PostCard_zRDr9crDL9xzLTocHfUkInuXoyHTBCrIQn7CTKZN0",{"head":97},{"link":98,"style":99},[],[],1790694219379]