Background and Problems
When many developers first encounter large language models, they can easily get confused by concepts such as parameter scale, tokens, context windows, temperature, and system prompts. In real implementations, the problem is often not whether the model is smart enough, but whether we have clearly defined the task boundaries: which text can the model see in a single call? Which information must be retained? How should the output format be constrained? How should errors be handled? From an engineering perspective, this article clarifies three fundamental but critical topics: how Transformers work, the limits of context windows, and basic methods for prompt design.

Core Concepts
Transformer: Understanding Sequences with Attention
The Transformer is a deep learning architecture introduced in the 2017 paper Attention Is All You Need. It has a notable difference from earlier common RNNs and LSTMs: instead of forcing word-by-word processing in time steps, it uses self-attention to observe multiple positions in a sequence simultaneously and compute the relationships between them. This makes training easier to parallelize and better suited for modeling long-range dependencies.
In large language models, text is first split into tokens. Tokens may be words, subwords, or symbols, depending on the tokenizer. Then each token is mapped to a vector and augmented with positional encoding so the model knows the order. After that, the vectors pass through multiple layers of attention and feed-forward networks, progressively refining contextual relationships. Many modern generative models adopt a decoder-only architecture, predicting the next token based on existing text. According to public materials, models such as GPT and LLaMA are based on or improve upon the Transformer architecture; refer to official documentation for specific implementations.
When understanding Transformers, focus on several key components:
- Token Embedding: Converts discrete text into continuous vectors.
- Positional Encoding: Adds word-order information; otherwise the model only knows the set of contents, not their sequence.
- Self-Attention: Uses Query, Key, and Value to determine which positions are more relevant.
- Multi-Head Attention: Uses multiple attention groups to capture relationships from different dimensions.
- Feed-Forward Networks, Residual Connections, and Layer Normalization: Improve expressive power and help stabilize training.
scores = Q @ K.T / sqrt(d_k); weights = softmax(scores); output = weights @ VThis pseudocode is not a complete model; it only shows how attention is computed: first use Q and K to obtain relevance scores, then convert them into weights with softmax, and finally compute a weighted sum of V. For engineers, this abstraction is enough to explain why models can select important information based on context.
Context Window: The Range of Tokens a Model Can Reference at Once
The context window is often called context length, max tokens, or maximum context length. It limits the total number of tokens occupied by both input and output in a single request. For example, a window of 8,192 tokens does not mean you can blindly insert 8,192 Chinese characters or English words, because different tokenizers produce different splits. Chinese text, code, tables, and Markdown symbols can all affect token counts.
A more robust engineering approach is to treat the context window as a budget rather than a capacity ceiling. System prompts, user input, retrieval results, conversation history, tool outputs, and expected responses should all be included in the budget. For long-document question answering, you usually need chunking, summarization, or retrieval augmentation instead of stuffing the entire document into the model at once.
Prompts: Writing Task Constraints as Model-Executable Input
A prompt is not just a simple question; it is the input specification for a reasoning task. A stable prompt usually includes role, task, context, constraints, output format, and necessary examples. For tasks such as classification, extraction, summarization, and code generation, the closer the prompt is to a clear, testable requirements document, the more stable the model output will be.
Practical Steps or Checklist
- Identify the task type: First determine whether it is open-ended generation, structured extraction, classification, or code completion. Different tasks have different requirements for temperature and examples.
- Provide the minimum necessary context: Keep only the material directly related to the task, and avoid irrelevant text diluting key instructions.
- Split long text: If it exceeds the window or approaches the limit, chunk and summarize first, then aggregate or use retrieval augmentation.
- Fix the output format: Require output as a list, table, JSON, or fixed fields, and provide field descriptions.
- Add few-shot examples: Use one to three examples to calibrate tone, granularity, and edge cases.
- Reserve space for output: Do not fill the input completely; leave room for generation and formatting.
- Run regression tests: Prepare a set of typical cases and revalidate after every prompt or model version change.
A simple prompt template can be organized like this:
template = 'Role: Technical documentation assistant. Task: Generate a summary from the given text. Requirements: Output 3 key points, each no more than 20 words. Constraints: Do not fabricate facts, do not output irrelevant content. Text: {text}'This example illustrates a structured way to write prompts: role, task, requirements, constraints, and input data are clearly separated, making it easier for programs to assemble and maintain later.
Common Pitfalls and Suggestions
- Treating the context window as permanent memory: The model cannot see information outside the window. Long conversations require summarization, external storage, or retrieval mechanisms.
- Confusing tokens with characters: A Chinese character, an English word, or a code symbol is not necessarily equal to one token. Use the actual tokenizer to count.
- Losing focus because the prompt is too long: Place key constraints at the beginning and end, and present them as lists.
- Only tuning parameters without fixing the input: Sampling parameters such as temperature and top_p affect randomness, but when output is unstable, first check the prompt, context, and task boundaries.
- Ignoring hallucinations and validation: Models may generate fluent but incorrect content. For decisions involving facts, healthcare, legal matters, finance, or production environments, add retrieval, rule validation, human review, or refusal mechanisms.
- Blindly pursuing larger windows: Longer context is not always better; cost, latency, and information noise all increase. If a small context can solve the problem stably, it is usually more economical.
Directions for Further Reading
- The original Transformer paper: Understand self-attention, encoder, and decoder design.
- Tokenizers and BPE: Learn about token counts, vocabulary size, and multilingual processing differences.
- RAG and long-document processing: Study retrieval-augmented generation, chunking strategies, and reranking.
- Structured output: Explore JSON Schema, function calling, and constrained decoding.
- Model evaluation: Build comprehensive metrics covering accuracy, hallucination rate, latency, cost, and user feedback.
References
Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.
Contact: chenxj.g@gmail.com
