A walkthrough is presented for deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod. The deployment utilizes vLLM. The process includes cluster provisioning. NVFP4 quantization is used to reduce model size. An OpenAI-compatible endpoint is created. This endpoint includes built-in reasoning, tool calling, and native MTP speculative decoding. This configuration provides an inference endpoint suitable for a range of applications.
