Skip to main content
Raw GPU capacity is not much use until something can see inside it and act on it. That is the job of the control plane: it turns capacity into hardware you can observe per GPU, divide into clusters, and deploy models onto without leaving the browser. The control plane is required on every Marketplace rental and is installed automatically at checkout — there is nothing to configure. Bare metal customers hold root access to their machines and run their own stack on them. Without the control plane, Acasia Cloud reports those deployments at the node level rather than per GPU. To discuss running the control plane on a bare metal deployment, contact your Acasia team.

What the control plane provides

Telemetry

Without the control plane, Acasia Cloud reports at the node level: machine health, node utilization, and alerting when a node needs attention. The control plane reports per-GPU state back to Cloud, so utilization, memory, and health are visible for each individual GPU rather than as a machine-level aggregate. Use it to confirm a workload is saturating the hardware you are paying for, and to catch a single GPU that has stopped contributing. On a Marketplace rental, Cloud scopes this to the GPUs you rented. On a bare metal reservation with the control plane installed, it covers the machines reserved for your organization.

Model deployment

The model library in Acasia Cloud lists models that are compatible with Acasia hardware. Selecting one deploys it onto a cluster and exposes it as an inference endpoint. Because the runtime is already in place, capacity can begin serving inference almost immediately after it is rented — which is usually the gating step when building a software product or service on top of a model. See Inference Models for the deployment workflow.

Cluster sizing

A cluster is an allocation of GPU, CPU, RAM, and storage that your workload runs on. The control plane lets you:
  • Create one or more clusters from the capacity you hold
  • Group GPUs within a node, or across several nodes, into a single cluster
  • Resize a cluster as the workload grows or shrinks
  • Wind a cluster down when the work is finished
This means the shape of your compute follows the workload rather than the other way around. See Clusters for lifecycle states and field-level detail.

Design principles

Installed, not assembled

On a Marketplace rental the control plane arrives with the capacity. No driver stack, runtime, or setup is expected of the customer.

Report from the GPU, not around it

Telemetry is collected at the GPU level so that Cloud reflects what each GPU is actually doing, rather than what the machine reports on their behalf.

Capacity follows the workload

Clusters are created, resized, and retired on demand, so allocated compute can match what the workload needs at that moment.

Next steps

Acasia Cloud

The platform the control plane reports into.

Clusters

Create, inspect, and retire clusters.

Inference Models

Deploy a model from the library onto a cluster.