Skip to main content
Inference Models Deployments page showing the Active tab with one Nemotron deployment, REST API banner, and Create New Inference button

Page layout

Deployments are organized by status tab:
  • Active — Currently running endpoints
  • Terminated — Stopped or removed deployments
  • All — Complete deployment history

Deployment fields

Endpoint lifecycle

Creating an inference endpoint

1

Click Create New Inference

Opens the deployment configuration panel.
2

Select a model

Choose from the available models in the catalog. Model availability depends on your organization’s catalog configuration.
3

Select a cluster

Choose an Active cluster with sufficient GPU capacity for the model. Large models may require multiple GPUs.
4

Configure Hugging Face token (if required)

Some models require a Hugging Face access token for download. Enter your token if prompted. Store this token securely — it is used at deployment time to pull the model weights.
5

Deploy

Submit the deployment. The endpoint enters Deploying state. Once provisioning completes, it transitions to Active and the endpoint URL becomes available.

Calling an endpoint

Once Active, the endpoint URL is available in the deployment table. Use this URL with an API key to call the model from your application.
Inference endpoints use the same API key authentication as other Acasia services. Generate an API key from the API Keys page and include it in the Authorization header of your requests.

Terminating an endpoint

To stop a deployment, select Terminate from the endpoint’s actions menu. The endpoint transitions to Stopping, then Terminated.
Terminating an endpoint immediately stops it from accepting requests. Confirm no critical workloads depend on the endpoint before terminating.

Common issues

Endpoint stuck in Deploying — Confirm the selected cluster is Active and has sufficient GPU capacity. Large models may take several minutes to load. Endpoint URL not available — The endpoint must reach Active state before the URL is usable. Hugging Face token error — Confirm the token has access to the requested model and has not expired. Some models require accepting a license agreement on Hugging Face before the token grants access. Failed deployment — Check that the cluster has sufficient GPU memory for the model. Review the model catalog for hardware requirements.