# Gpu Autoscaling Engineer

> - Author and validate autoscaling for GPU workloads on Kubernetes: KEDA ScaledObjects with DCGM + application-metric dual triggers, scale-to-zero, asymmetric scale-up/scale-down behavior for expensive GPU nodes, warm pools via WhenEmpty consolidation, GPU node-pool sizing and taints. Use whenever the user mentions KEDA, ScaledObject, HPA for GPU or LLM inference, scale to zero, DCGMFIDEVGPUUTIL, autoscaling triggers, TTFT or tokens-per-second or p95 latency as scaling signals, GPU nodes not scaling down, cold starts on GPU nodes, warm GPU capacity, Karpenter/cluster-autoscaler for GPUs, or "pods scale but nodes don't". Also use to review existing ScaledObject/HPA YAML for GPU services. For choosing how to share one GPU across pods use sibling gpu-sharing-advisor; for Pending pods and nodes that never join use gpu-workload-troubleshooter; for the money view use gpu-cost-optimizer.

## Facts
- Page: https://tashan.sh/capability/skill-cloud-byte-consulting-gpu-autoscaling-engineer
- tashan id: skill:Cloud-Byte-Consulting/gpu-autoscaling-engineer
- Source: https://github.com/Cloud-Byte-Consulting/plugins
- Type: skill
- Category: cloud
- tashan score: not scored (catalogued only — too little public evidence)
- Adoption: 9.0
- Upkeep: 95.0
- Freshness: 90.0
- Evidence coverage: 84% of the inputs this score can use
- Health: active
- Instruction depth: not yet graded
- License: Apache-2.0
- Official: no

## Install

```sh
cp -r gpu-autoscaling-engineer ~/.claude/skills/
```

## Security audit
Not scanned. We audit npm-published capabilities; this one has no npm package we can resolve, or has not reached the queue. This is not a clean bill of health.

---
Measured 2026-08-22 by tashan (https://tashan.sh) from public evidence. Scorer s5.
