> ## Documentation Index
> Fetch the complete documentation index at: https://docs.acasia.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Inference Models

> Put a model from the library onto a cluster and serve it behind a REST endpoint.

![Inference Models Deployments page showing the Active tab with one Nemotron deployment, REST API banner, and Create New Inference button](https://docs.acasia.com/assets/inference-models-overview-ygr934BL.png)

## Page layout

Deployments are organized by status tab:

* **Active** — Currently running endpoints
* **Terminated** — Stopped or removed deployments
* **All** — Complete deployment history

## Deployment fields

| Field        | Description                                 |
| ------------ | ------------------------------------------- |
| Model name   | Name and provider of the deployed model     |
| Status       | Current state of the deployment             |
| Cluster      | Cluster the endpoint is running on          |
| Endpoint URL | REST API URL for calling the deployed model |
| Created      | When the deployment was created             |

## Endpoint lifecycle

| State      | Meaning                                                       |
| ---------- | ------------------------------------------------------------- |
| Deploying  | The endpoint is being provisioned on the cluster              |
| Active     | The endpoint is running and accepting requests                |
| Stopping   | The endpoint is being shut down                               |
| Terminated | The endpoint has been stopped — it no longer accepts requests |
| Failed     | The deployment did not complete successfully                  |

## Creating an inference endpoint

<Steps>
  <Step title="Click Create New Inference">
    Opens the deployment configuration panel.
  </Step>

  <Step title="Select a model">
    Choose from the available models in the catalog. Model availability depends on your organization's catalog configuration.
  </Step>

  <Step title="Select a cluster">
    Choose an Active cluster with sufficient GPU capacity for the model. Large models may require multiple GPUs.
  </Step>

  <Step title="Configure Hugging Face token (if required)">
    Some models require a Hugging Face access token for download. Enter your token if prompted. Store this token securely — it is used at deployment time to pull the model weights.
  </Step>

  <Step title="Deploy">
    Submit the deployment. The endpoint enters Deploying state. Once provisioning completes, it transitions to Active and the endpoint URL becomes available.
  </Step>
</Steps>

## Calling an endpoint

Once Active, the endpoint URL is available in the deployment table. Use this URL with an API key to call the model from your application.

<Note>
  Inference endpoints use the same API key authentication as other Acasia services. Generate an API key from the API Keys page and include it in the Authorization header of your requests.
</Note>

```bash theme={null}
curl https://<your-endpoint-url>/v1/chat/completions \
  -H "Authorization: Bearer <your-api-key>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model-name>",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
```

## Terminating an endpoint

To stop a deployment, select **Terminate** from the endpoint's actions menu. The endpoint transitions to Stopping, then Terminated.

<Warning>
  Terminating an endpoint immediately stops it from accepting requests. Confirm no critical workloads depend on the endpoint before terminating.
</Warning>

## Common issues

**Endpoint stuck in Deploying** — Confirm the selected cluster is Active and has sufficient GPU capacity. Large models may take several minutes to load.

**Endpoint URL not available** — The endpoint must reach Active state before the URL is usable.

**Hugging Face token error** — Confirm the token has access to the requested model and has not expired. Some models require accepting a license agreement on Hugging Face before the token grants access.

**Failed deployment** — Check that the cluster has sufficient GPU memory for the model. Review the model catalog for hardware requirements.
