chat_template_kwargs isn't being respected by the vLLM deployment. Qwen3 models support /no_think as an inline suffix in the user message to disable thinking mode. This is the most reliable method across all serving backends (vLLM, Ollama, SGLang).
chat_template_kwargs isn't being respected by the vLLM deployment. Qwen3 models support /no_think as an inline suffix in the user message to disable thinking mode. This is the most reliable method across all serving backends (vLLM, Ollama, SGLang).