After reading, you will get a seven-category failure attribution table, three acceptance checklists, and autonomy thresholds. Follow the steps to run a read-only pilot first, then decide which actions should be executed automatically.

读书明理
读书明理

Author: Jinxuezhai · Guanshan Academy

Full text: about 3,633 characters · estimated reading time: 13 minutes


Recently, many teams have placed AI agents into real workflows: customer support tickets, operations inspections, data organization, internal approvals. The demo runs smoothly, but problems appear once multiple people, multiple systems, and long-running execution are involved. On the surface, performance seems unstable; the real blocker is that failure signals have not been broken down.

This article provides an acceptance method: a seven-category failure postmortem table, three checklists, a three-day read-only pilot, and staged delegation. You can follow the steps to run a read-only pilot first, then decide which actions to automate.

I. Make Failure Explicit First: Seven Signals Determine Whether to Continue

1. Task Understanding Drift: It Completes a Different Task

If the user asks the agent to organize customer complaints and generate reply drafts, it may only classify them. If the user asks it to fill in fields, it may rewrite the entire copy. The root cause of drift usually lies in goal decomposition, not the model itself.

You can monitor first-pass success rate, human intervention rate, and the proportion of low-quality outputs. If the same intent frequently goes off track, first improve the task template and acceptance fields before considering prompts.

2. Context Loss: Key Evidence Does Not Reach the Next Step

In multi-turn tasks, earlier constraints, customer tier, budget definitions, and historical handling conclusions may be dropped by the time tools are called. The result may look reasonable, but it cannot withstand review.

Observe context reference count, key-field accuracy, and task duration. If repeated queries and repeated confirmations increase, the pipeline is not passing evidence downstream.

3. Tool Overreach: Actions That Should Not Be Taken Are Taken

A read API becomes a write API, a draft becomes a sent message, and a query becomes a deletion. This kind of issue is more dangerous than a wrong answer because it directly changes external state.

Record the number of sensitive-action hits, tool-call input parameters, caller identity, and target system. When overreach is found, first narrow the tool whitelist rather than adding a sentence saying “do not delete.”

4. Unverifiable Results: The Output Looks Good, but Cannot Be Traced

The agent gives a conclusion but leaves no source of evidence, calculation definition, tool return value, or confidence indication. Business colleagues do not dare to sign off, and engineering colleagues cannot reproduce it.

Acceptance should examine key-field accuracy, completeness of the evidence chain, and whether users accept the result. Without verifiable fields, actions should not be handed to automated processes.

5. Cost Runaway: One Task Turns into a Loop

Repeated retrieval, repeated summarization, multiple model calls, and multiple tool requests cause the cost per task to keep rising. If the team only watches tokens, it can miss latency, retries, and manual review.

Track repeated call count, task duration, error recovery time, and per-task cost trend. Cost anomalies often mean the workflow has a loop; break the loop before scaling.

6. Missing Approval: Automated Actions Bypass Humans

High-risk actions have no approval point, or approval is merely a formality. When something goes wrong, the chain of responsibility breaks, and the postmortem cannot recover why it was approved.

Check the missing-approval rate, number of external write actions, and retention period for approval records. Approval is not adding a button; it is binding each high-risk action to a person and a reason.

7. Missing Rollback: Changes Cannot Be Reverted

Editing documents, sending messages, updating tickets, and adjusting configurations—after failure, there is no rollback ID, no compensating action, and no fallback template. One success may conceal the next incident.

Record the rollback owner, error recovery time, idempotency, and audit ID. If write operations cannot be rolled back, allow them only in low-impact pilot scopes.

— First collect runtime logs, then determine whether the problem lies in the model, prompt, tools, or workflow.

Do not immediately switch frameworks, expand permissions, or add prompts. First make failures observable; only then does the team have a basis for discussion.

Illustration · Lakeside Dusk

Illustration · Lakeside Dusk

II. Define Acceptance Criteria: Replace “Looks Usable” with Three Checklists

1. Business Acceptance Checklist

The business side asks only whether the result can be used. You can fix five items: task goal completion rate, key-field accuracy, whether users accept the result, proportion of low-quality outputs, and manual review ratio. Each item must specify its definition.

For example, key customer-complaint fields include customer identity, request type, responsible team, and handling deadline. If one is missing, the task is not complete.

2. Engineering Acceptance Checklist

The engineering side must be able to reproduce, track, and stop losses. You can fix: timeout rate, retry count, idempotency, log-field completeness, dependency-service error codes, and call-chain latency distribution.

Logs must include at least task_id, run_id, agent_version, tool_name, input_hash, output_hash, duration, and error_code.

event: agent.plan
task_id: T-1024
run_id: R-77
user_intent: Organize customer complaints and generate reply drafts
planned_steps: retrieve, summarize, draft
tool_call: search_tickets status=open
risk_level: read_only
approval_required: false
idempotency_key: T-1024-R-77-search-001

This log is not for appearance; it is for attribution after the three-day pilot. If there are only results and no process, the postmortem becomes guesswork.

3. Security and Compliance Checklist

The security focus is not model capability, but action boundaries. You need to check: exposure of sensitive data, external write actions, approvals for deletion/payment/publishing operations, retention period for audit records, and rollback owner.

Customer-facing outbound messages, fund changes, production configuration changes, and record deletion should not be automatically executed by default. Even if the business is pressing, run them first in a sandbox or shadow system.

Actionable Checklist

☐ Establish a three-key log of task_id, run_id, and agent_version to ensure every execution is traceable.

☐ Record an input summary, output summary, duration, error code, and idempotency key for each tool call.

☐ Place external writes, deletions, payments, publishing, and customer-facing outbound actions on a default block list.

☐ Write clear definitions, criteria, and reviewers for business acceptance fields; do not accept vague evaluations.

☐ Set up daily … for failure samples.

Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.

Contact: chenxj.g@gmail.com