Background and Problems

More and more teams want to place large models in their own data centers or dedicated clouds for reasons such as data compliance, intranet integration, controllable costs, and customized inference pipelines. However, private deployment is not as simple as downloading weights and starting a service. In real projects, the most common failure points are often not the model itself, but underestimation of hardware resources, chaotic container environment dependencies, unstable GPU scheduling, and lack of observability after launch.

观山静思
观山静思

From an engineering perspective, a privately deployed large-model service involves at least three layers: the underlying compute layer, including GPUs, CPUs, memory, storage, and networking; the middle runtime layer, including drivers, CUDA, container runtimes, and inference frameworks; and the upper service layer, including API gateways, authentication, rate limiting, monitoring, and alerting. A weakness in any layer can lead to timeouts, GPU memory overflows, or node drift.

Core Concepts

1. GPU Memory Is Not the Only Metric

Many people's first reaction is to look at GPU memory. GPU memory determines whether the model weights can fit, but the inference experience is also affected by memory bandwidth, compute unit performance, PCIe or NVLink interconnects, and CPU preprocessing capability. According to public materials, in long-context scenarios, the KV cache can significantly increase GPU memory usage.

A rough estimate: the GPU memory required for weights is approximately the number of parameters multiplied by the number of bytes per parameter. For example, a 7B model in FP16 requires about 14GB for weights. Adding KV cache, runtime buffers, and fragmentation, 24GB of GPU memory provides a more comfortable margin. If INT8 or INT4 quantization is used, weight memory decreases, but quality loss and framework compatibility should be evaluated. Please refer to official documentation.

2. Inference Frameworks and Service Deployment

In production environments, it is not recommended to load models directly with scripts. Instead, choose mature inference serving frameworks, such as vLLM, Text Generation Inference, Ollama, or vendor inference services. They usually provide batching, continuous batching, streaming output, OpenAI-compatible interfaces, and basic metrics. Interface standardization is more important than peak performance, because gateways, monitoring, and clients all rely on stable protocols.

3. The Value of Container Orchestration

Containerization isolates dependencies, and orchestration systems can handle replicas, health checks, resource limits, and node scheduling. For GPU workloads, NVIDIA drivers, Container Toolkit, GPU Operator, or device plugins must also work together. Kubernetes is suitable for multi-team scenarios; for single-machine or small-scale environments, Docker Compose can reduce complexity.

4. Monitoring and Alerting Should Cover Both GPUs and Business Metrics

Looking only at GPU utilization is not enough. High utilization does not necessarily mean the service is healthy; it may simply indicate request backlog. Low utilization does not necessarily indicate a failure; it may be due to batching policies or client-side rate limiting. A more reasonable approach is to collect hardware, runtime, and business metrics at the same time: GPU memory, temperature, power consumption, request latency, time to first token, generation speed, queue length, error rate, and replica restart count.

Practical Steps and Checklist

Step 1: Define the Business Profile

  • Model scale: 7B, 13B, 30B, or larger.
  • Context length: short conversations, long documents, or code completion.
  • Concurrency targets: peak QPS, maximum tokens per request, and whether queuing is allowed.
  • Security boundaries: whether it must be fully offline, and whether auditing and permission control are required.

Step 2: Hardware Selection

  • GPU: prioritize GPU memory capacity, memory bandwidth, and driver ecosystem. For long-context workloads, reserve at least 30% GPU memory headroom.
  • CPU: there is no need to blindly pursue top-tier models, but ensure that tokenization, data loading, logging, and sidecars do not become bottlenecks.
  • Memory: it is generally recommended to have at least twice the size of the model weights, while accounting for system and container overhead.
  • Storage: model files are large, so high-speed SSDs are recommended; separate the image registry from the model cache.
  • Networking: for multi-machine deployments, pay attention to high-speed Ethernet or RDMA; for single-machine deployments, pay attention to PCIe topology and cooling.

Step 3: Prepare the Runtime Environment

It is recommended to pin the versions of drivers, CUDA, container runtime, and inference framework. Do not upgrade drivers on production nodes on the fly. Check items include:

  • nvidia-smi can stably detect GPUs.
  • Device nodes are accessible inside containers.
  • Model files are mounted read-only or cached from object storage.
  • Service ports, health check paths, and timeout values are configured.

Step 4: Containerize the Inference Service

Below is a simplified docker run example for starting an inference service compatible with the OpenAI interface. In actual deployments, replace the image, model path, and resource parameters, and refer to official documentation.

docker run --gpus all -v /data/models:/models -p 8000:8000 vllm/vllm-openai:latest --model /models/your-model --max-model-len 8192

This example mounts the model directory into the container and exposes an HTTP interface. In production, also add restart policies, resource limits, log rotation, and health checks. If using Kubernetes, you can manage replicas and scheduling through Deployment, Service, HPA, and node affinity.

Step 5: Integrate Monitoring and Alerting

A common approach is to scrape metrics with Prometheus, visualize them in Grafana, and send alerts with Alertmanager. GPU metrics can be exposed through DCGM Exporter or vendor exporters. Business metrics can be provided by the inference framework or aggregated at the gateway layer. It is recommended to configure at least the following alerts: GPU memory continuously approaching the limit, abnormal GPU temperature, rising request error rate, time to first token exceeding a threshold, and frequent replica restarts.

Pre-launch Checklist

  • Load testing: cover short requests, long requests, concurrency spikes, and streaming output.
  • Failure drills: simulate GPU device loss, container OOM, and node restarts.
  • Rate limiting: set maximum concurrency and request body size at the gateway layer.
  • Logging: retain request IDs, durations, token counts, and error codes, but do not log sensitive prompts.
  • Capacity: reserve growth headroom for GPU memory, disk, and memory.

Common Pitfalls and Recommendations

  • Estimating GPU memory based only on weights: Ignoring KV cache and batching can cause direct OOM in long-context or high-concurrency scenarios. It is recommended to run load tests with real business samples.
  • Blindly pursuing multiple GPUs: If the model can run on a single GPU, multi-GPU parallelism may not improve throughput and can increase communication and scheduling complexity. First clarify whether you need tensor parallelism, pipeline parallelism, or replica scaling.
  • Oversized container images and version drift: Separate model files from runtime images, use fixed version tags, and avoid using latest directly in production.
  • GPU scheduling affinity issues in Kubernetes: Ensure that device plugins, node labels, taints, and resource requests are consistent; otherwise Pods may remain Pending for a long time.
  • Alerting only on GPU utilization: Combine time to first token, queue length, and error rate. Otherwise, metrics may look normal while user experience is poor.

Directions for Further Reading

  • Batching, PagedAttention, continuous batching, and quantization strategies in inference frameworks.
  • Kubernetes GPU Operator, Device Plugin, and node affinity design.
  • Metric modeling for AI services with DCGM, Prometheus, and OpenTelemetry.
  • Model gateway design: authentication, quotas, caching, content auditing, and multi-model routing.

Overall, privately deploying large models is a systems engineering effort. Hardware determines the upper bound, orchestration determines stability, and monitoring determines maintainability. Teams should first run through the full pipeline in a small-scale, reproducible environment, then gradually increase concurrency and model scale. Avoid piling up expensive hardware from the start while losing points on engineering and operations details.

References

Disclaimer: This article is compiled from publicly available online sources for educational purposes only and does not represent the official position of Guanshan Academy. If you believe your rights have been infringed, please contact us for removal.

Contact: chenxj.g@gmail.com