> ## Documentation Index
> Fetch the complete documentation index at: https://docs.acasia.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Deploy your first inference endpoint

> Pick a model, pick a cluster, and have a callable REST endpoint in about five minutes.

**In this guide:**

* Confirm your organization context and review available clusters.
* Pick a model from the Inference Models catalog and a target cluster with enough capacity.
* Deploy the endpoint, wait for Active status, then connect your application through the Acasia API.

## Before you begin

| Requirement                     | Description                                                               |
| ------------------------------- | ------------------------------------------------------------------------- |
| Acasia account access           | You can log in to the Acasia Developer Portal.                            |
| Correct organization context    | You are operating inside the intended organization.                       |
| Active cluster                  | Your organization has at least one active cluster available.              |
| Available resources             | The cluster has enough GPU, CPU, RAM, and storage for the selected model. |
| Model access                    | The model is available in the Inference Models catalog.                   |
| Hugging Face token, if required | Some models may require authenticated Hugging Face access.                |

<Warning>
  Inference endpoints, clusters, and API keys are scoped per organization. Use the organization switcher to confirm you're in the right one before deploying.
</Warning>

## Step 1 — Review your dashboard

| Dashboard area | What to check                                               |
| -------------- | ----------------------------------------------------------- |
| Clusters       | Confirm that at least one cluster is available.             |
| Cluster status | Confirm the cluster is marked Active.                       |
| GPU model      | Confirm the cluster has the GPU type needed for your model. |
| Live models    | Check whether models are already running on the cluster.    |

## Step 2 — Inspect cluster resources

| Resource  | Why it matters                                                               |
| --------- | ---------------------------------------------------------------------------- |
| GPU count | Determines whether the model can run with the required accelerator capacity. |
| GPU model | Confirms hardware compatibility.                                             |
| CPU cores | Supports runtime services, preprocessing, and orchestration.                 |
| RAM       | Supports model loading and serving processes.                                |
| Storage   | Stores model files, cache, logs, and temporary outputs.                      |
| Status    | The cluster must be Active before deployment.                                |

## Step 3 — Open Inference Models

From the left navigation in the Developer Portal, select **Inference Models**.

| Page element         | Purpose                                               |
| -------------------- | ----------------------------------------------------- |
| Create New Inference | Starts the deployment workflow.                       |
| Hugging Face Token   | Adds or manages Hugging Face access for gated models. |
| Active tab           | Shows running inference deployments.                  |
| Terminated tab       | Shows stopped or removed deployments.                 |

## Step 4 — Choose a model

Select **Create New Inference** and pick a model from the catalog.

| Consideration       | Guidance                                                                                        |
| ------------------- | ----------------------------------------------------------------------------------------------- |
| Task                | Select a model that matches your use case — text, image, video, audio, or multimodal inference. |
| Model size          | Larger models usually require more GPU memory and compute.                                      |
| Access requirements | Some models may require a Hugging Face token.                                                   |

<Note>
  If the model is gated, add a Hugging Face token from the Inference Models page before deployment. Without it, provisioning will fail when the runtime tries to pull model weights.
</Note>

## Step 5 — Choose a cluster and deploy

<Steps>
  <Step title="Select the target cluster">
    Pick an Active cluster that has enough GPU, CPU, RAM, and storage for the selected model.
  </Step>

  <Step title="Review the deployment summary">
    Confirm the model, cluster, and any optional parameters.
  </Step>

  <Step title="Click Deploy Endpoint">
    Acasia Cloud records the desired deployment state. The control plane then provisions the model on the target cluster.
  </Step>
</Steps>

## Step 6 — Verify the endpoint

Return to the **Active** tab on Inference Models and watch the endpoint move through its lifecycle states.

| State        | Meaning                                       |
| ------------ | --------------------------------------------- |
| Active       | The endpoint is running and available.        |
| Provisioning | The endpoint is being created or prepared.    |
| Failed       | The deployment did not complete successfully. |
| Terminated   | The endpoint has been stopped or removed.     |

## Step 7 — Call the endpoint

```bash theme={null}
curl --request POST \
  --url https://endpoint.acasia.ai/v1/completions \
  --header "Content-Type: application/json" \
  --header "Authorization: Bearer $ACASIA_API_KEY" \
  --data '{
    "model": "your-model-id",
    "prompt": "Summarize the safety risks in a warehouse aisle.",
    "max_tokens": 300
  }'
```

<Note>
  Endpoint paths, model identifiers, and payload formats vary by deployment. Always confirm against the endpoint detail view in the Developer Portal.
</Note>

## Pre-deployment checklist

| Check                | Confirmation                                            |
| -------------------- | ------------------------------------------------------- |
| Organization context | You are in the correct organization.                    |
| Cluster status       | The target cluster is active.                           |
| GPU capacity         | The cluster has enough GPUs for the model.              |
| Memory and storage   | The cluster has enough RAM and storage.                 |
| Model selection      | The model matches your use case.                        |
| Hugging Face access  | Token is configured if the model requires gated access. |

## Summary

1. Confirm the correct organization context
2. Review available clusters
3. Open Inference Models → Create New Inference
4. Choose a model and an active cluster
5. Click Deploy Endpoint
6. Verify the endpoint appears under Active
7. Connect your application through the Acasia API

## Next steps

<CardGroup cols={3}>
  <Card title="Create an API key" icon="key" href="/get-started/create-api-key">
    Generate and store a credential for production use.
  </Card>

  <Card title="Connect with SSH" icon="terminal" href="/get-started/connect-cluster-ssh">
    Open a secure shell to your cluster.
  </Card>

  <Card title="Connect your IDE" icon="code" href="/get-started/connect-ide-to-cluster">
    Develop against a remote GPU cluster.
  </Card>
</CardGroup>
