Skip to main content
In this guide:
  • Confirm your organization context and review available clusters.
  • Pick a model from the Inference Models catalog and a target cluster with enough capacity.
  • Deploy the endpoint, wait for Active status, then connect your application through the Acasia API.

Before you begin

Inference endpoints, clusters, and API keys are scoped per organization. Use the organization switcher to confirm you’re in the right one before deploying.

Step 1 — Review your dashboard

Step 2 — Inspect cluster resources

Step 3 — Open Inference Models

From the left navigation in the Developer Portal, select Inference Models.

Step 4 — Choose a model

Select Create New Inference and pick a model from the catalog.
If the model is gated, add a Hugging Face token from the Inference Models page before deployment. Without it, provisioning will fail when the runtime tries to pull model weights.

Step 5 — Choose a cluster and deploy

1

Select the target cluster

Pick an Active cluster that has enough GPU, CPU, RAM, and storage for the selected model.
2

Review the deployment summary

Confirm the model, cluster, and any optional parameters.
3

Click Deploy Endpoint

Acasia Cloud records the desired deployment state. The control plane then provisions the model on the target cluster.

Step 6 — Verify the endpoint

Return to the Active tab on Inference Models and watch the endpoint move through its lifecycle states.

Step 7 — Call the endpoint

Endpoint paths, model identifiers, and payload formats vary by deployment. Always confirm against the endpoint detail view in the Developer Portal.

Pre-deployment checklist

Summary

  1. Confirm the correct organization context
  2. Review available clusters
  3. Open Inference Models → Create New Inference
  4. Choose a model and an active cluster
  5. Click Deploy Endpoint
  6. Verify the endpoint appears under Active
  7. Connect your application through the Acasia API

Next steps

Create an API key

Generate and store a credential for production use.

Connect with SSH

Open a secure shell to your cluster.

Connect your IDE

Develop against a remote GPU cluster.