Background and Problem: A Prompt Is Not a Chat Message, but Interface Design

When many teams integrate large models for the first time, they often encounter a gap: the model appears highly capable, but behaves inconsistently once applied to real business scenarios. Answers may be too broad, formats may be unparsable, or the model may drift under edge conditions. On the surface, this seems like the model is not smart enough. In practice, the more common cause is that the prompt does not clearly specify the task objective, input constraints, output format, and failure handling.

观山静思
观山静思

From an engineering perspective, a prompt is not merely a single sentence; it is a lightweight protocol for the model to execute. Like API documentation, it should be explicit about: who the role is, what the input is, what the output should look like, what must not be done, and how to handle uncertainty. System prompts, Few-shot examples, and structured output are the three most common levers in this protocol.

Core Concepts: What Problems Do These Three Techniques Solve?

1. System Prompts: Stabilize Role, Style, and Boundaries

A system prompt is typically used to define the model's overall behavioral framework. Compared with user inputs in each turn, it is better suited for long-lived rules, such as role definition, tone and style, safety boundaries, output preferences, and tool-calling constraints. A common mistake is to write the system prompt as a vague slogan, such as “Be professional, accurate, and concise.” Such descriptions are not entirely useless, but they lack executable criteria. A more reliable approach is to break requirements down into actions the model can follow.

For example, instead of writing “Please answer professionally,” write: “You are a technical documentation editor. Before answering, determine whether the user's question belongs to frontend engineering. If information is insufficient, ask no more than three clarification questions. Do not fabricate dependency versions that were not provided.” The advantage of this kind of prompt is that it reduces the model's freedom to improvise and makes later regression testing easier.

2. Few-shot: Use Examples to Calibrate Format and Judgment Criteria

Few-shot is not simply piling up examples. It uses examples to tell the model what kind of output should be produced for what kind of input. It is especially suitable for tasks such as classification, extraction, rewriting, review, and intent recognition. Examples provide two types of value: format calibration and boundary calibration. Format calibration tells the model the fields, order, length, and tone. Boundary calibration tells the model how to classify ambiguous input, or whether to return a null value.

In real projects, Few-shot examples do not need to be numerous; they need to be representative. Prioritize three types of samples: standard samples, boundary samples, and samples that are easily misjudged. Standard samples establish the basic format. Boundary samples define ambiguous areas. Misjudgment samples correct common errors. If the rules among examples conflict, the model can easily learn unstable patterns.

3. Structured Output: Make Results Parsable, Validatable, and Storable

When prompts enter production systems, outputs often cannot be only natural language; they need to be JSON, tables, enumerated fields, or fixed templates. The core of structured output is not simply telling the model to “output JSON.” Instead, the schema, field meanings, allowed value ranges, and default strategies must be clearly specified. Many failures occur not because the model cannot produce JSON, but because the prompt does not state whether to fill a missing field with null, an empty string, or omit the field.

According to public documentation, mainstream large models generally support reducing format errors through prompt constraints or official structured-output capabilities. However, different models vary in their support for JSON Schema, function calling, and response formats; always refer to the official documentation. From an engineering standpoint, always keep a layer of server-side validation and do not treat model output directly as trusted data.

Practical Steps: A Deployable Prompt Design Checklist

  • Clarify the task type: First determine whether the task is open-ended generation, information extraction, classification/decision-making, or code generation. Different types suit different prompting strategies.
  • Specify input sources: Tell the model which content comes from user input, which comes from system context, and which is untrusted. When necessary, use tags to separate data from instructions.
  • Define the output contract: Specify field names, types, lengths, enumerated values, whether empty values are allowed, and what to return on failure.
  • Add the minimum necessary examples: Start with 2 to 4 high-quality examples covering normal, boundary, and abnormal scenarios.
  • Set refusal and fallback strategies: For example, “If the information cannot be extracted from the given materials, return empty: true. Do not guess.”
  • Run regression tests: Prepare a fixed test set. Every time the prompt changes, compare output differences to avoid local optimizations causing overall regressions.

Below is a structured prompt example for ticket classification. It places the system role, input area, Few-shot examples, and JSON output requirements into the same template, making programmatic assembly and future maintenance easier.

system:
You are a ticket classification assistant. Classify tickets only based on the ticket content provided by the user. Do not infer information that is not given.
The output must be valid JSON with the following fields: category, urgency, reason.
category must be one of account, payment, bug, other.
urgency must be one of low, medium, high.
If the information is insufficient, set category to other and clearly explain the reason.

Example 1:
Input: I cannot log in. It says the password is incorrect, but I am not receiving the reset email.
Output: {"category":"account","urgency":"high","reason":"Login is blocked and the reset email is not working"}

Example 2:
Input: A button on the page does not respond after being clicked, and the console shows a 500 error.
Output: {"category":"bug","urgency":"medium","reason":"A feature is unavailable and a server-side error is present"}

User input:
{{ticket_content}}

The key point of this example is not the JSON itself, but that it turns “do not infer,” “enumerated values,” and “what to do when information is insufficient” into executable rules. For programs, this kind of output is easier to validate and route.

Common Pitfalls and Recommendations

  • Writing only the goal, not the constraints: For example, saying “Summarize this for me” without specifying length, target audience, or whether risk items should be preserved. Include acceptance criteria in the prompt.
  • Too many examples that contradict one another: Too many examples consume context and amplify noise. Keep samples that can actually change the model's judgment.
  • Mixing untrusted input directly into instructions: User input may contain manipulation or injection content. Wrap external text in explicit tags and declare that instructions inside it are invalid.
  • Over-relying on the model's self-discipline: Production environments should add JSON Schema validation, field sanitization, and retry mechanisms. Model output should be treated only as candidate results.
  • Ignoring model differences: The same prompt may perform noticeably differently across models. Re-evaluate before switching models; do not assume direct transferability.
  • Hard-coding prompts once and for all: Prompts should be version-controlled like code, with change reasons, test sets, and performance metrics recorded.

Directions for Further Reading

  • Context engineering: Explore how to organize system prompts, retrieved materials, conversation history, and tool results within limited tokens.
  • Structured output and function calling: Follow official capabilities for JSON output, tool calling, and response formats across models.
  • Prompt evaluation: Build golden sample sets, automatic scorers, and human spot-check workflows.
  • Security and injection protection: Understand risks such as prompt injection, privilege escalation, and sensitive information leakage, along with mitigation strategies.
  • RAG and knowledge boundaries: When tasks depend on private knowledge, combine retrieval, citations, and confidence control.

References

Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.

Contact: chenxj.g@gmail.com