dstack
dstack is a free, open-source orchestration layer for AI workloads across GPU clouds, Kubernetes, virtual machines, and bare-metal clusters. YAML configurations define fleets, development environments, tasks, services, presets, and volumes, while dstack handles infrastructure provisioning and job scheduling, including auto-scaling, port forwarding, and ingress. Supported accelerators include NVIDIA, AMD, TPU, and Tenstorrent. Documented backends include AWS, Azure, GCP, Kubernetes, GPU cloud providers, remote SSH hosts, and experimental Slurm. Tasks can use frameworks such as accelerate, torchrun, Ray, and Spark. Users manage resources with the CLI or HTTP API, and the server can run on a laptop or another environment able to access the clusters in use. Services can publish inference endpoints through gateways with HTTPS, custom domains, auto-scaling, and rate limits. The open-source plan is self-hosted. Server data and project secrets are stored in plaintext by default unless an administrator configures AES-256-GCM encryption. TPU support is limited to single-host instances of up to eight cores.
Who it is for
dstack suits teams orchestrating AI workloads across cloud and on-premises infrastructure, including GPU clouds, Kubernetes, and bare-metal clusters. It is relevant to users comfortable managing a self-hosted stack through YAML, a CLI, or an HTTP API.
What is good
- Free, open-source self-hosted orchestration stack
- Supports NVIDIA, AMD, TPU, and Tenstorrent accelerators
- Works with Kubernetes, VMs, GPU clouds, and bare metal
- CLI and HTTP API management options
- Inference gateways support HTTPS and custom domains
What to know first
- Server data is plaintext by default
- Project secrets are plaintext by default
- TPU support is limited to single-host instances
- Single-host TPU instances are capped at eight cores
Verdict
dstack provides a free orchestration layer for AI workloads across heterogeneous infrastructure, with YAML configuration and CLI or API management. Administrators should account for its plaintext-by-default data and secrets, and for the stated single-host TPU limit.
dstack plans and pricing
All plansCompared on GPU cluster management software
- Free plan
- Yes
- Deployment model
- hybrid
- Workload scheduling
- both
- Kubernetes support
- Yes
- Quota controls
- No
- GPU utilization metrics
- Yes
- Cloud GPU support
- Yes


