# Vllm Xpu Run

> Serve a Hugging Face safetensors model on an Intel GPU with upstream vLLM-XPU's OpenAI-compatible API. Covers image choice, container launch, the right vllm serve flags (dtype, enforce-eager, model-impl fallback, attention backend, quant + KV-cache pairing), and the transformers-backend fallback for unsupported architectures. Use for /v1/chat/completions or /v1/completions on an Intel GPU. Not for pure PyTorch without a server (use torch-xpu-run), throughput numbers (use vllm-xpu-bench), or NVID

## Facts
- Page: https://tashan.sh/capability/plugin-intel-gpu-ai-skills-vllm-xpu-run
- tashan id: plugin:intel/gpu-ai-skills/vllm-xpu-run
- Source: https://github.com/intel/gpu-ai-skills
- Type: plugin
- Category: ai
- tashan score: 45.0 / 100
- Adoption: 7.0
- Upkeep: 100.0
- Freshness: 99.0
- Evidence coverage: 84% of the inputs this score can use
- Health: active
- Instruction depth: not yet graded
- License: Apache-2.0
- Official: no

## Install

```sh
/plugin marketplace add intel/gpu-ai-skills
/plugin install vllm-xpu-run@intel-model-skillpack
```

## Security audit
Not scanned. We audit npm-published capabilities; this one has no npm package we can resolve, or has not reached the queue. This is not a clean bill of health.

---
Measured 2026-09-13 by tashan (https://tashan.sh) from public evidence. Scorer s5.
