- Confirm your organization context and review available clusters.
- Pick a model from the Inference Models catalog and a target cluster with enough capacity.
- Deploy the endpoint, wait for Active status, then connect your application through the Acasia API.
Before you begin
Step 1 — Review your dashboard
Step 2 — Inspect cluster resources
Step 3 — Open Inference Models
From the left navigation in the Developer Portal, select Inference Models.Step 4 — Choose a model
Select Create New Inference and pick a model from the catalog.If the model is gated, add a Hugging Face token from the Inference Models page before deployment. Without it, provisioning will fail when the runtime tries to pull model weights.
Step 5 — Choose a cluster and deploy
1
Select the target cluster
Pick an Active cluster that has enough GPU, CPU, RAM, and storage for the selected model.
2
Review the deployment summary
Confirm the model, cluster, and any optional parameters.
3
Click Deploy Endpoint
Acasia Cloud records the desired deployment state. The control plane then provisions the model on the target cluster.
Step 6 — Verify the endpoint
Return to the Active tab on Inference Models and watch the endpoint move through its lifecycle states.Step 7 — Call the endpoint
Endpoint paths, model identifiers, and payload formats vary by deployment. Always confirm against the endpoint detail view in the Developer Portal.
Pre-deployment checklist
Summary
- Confirm the correct organization context
- Review available clusters
- Open Inference Models → Create New Inference
- Choose a model and an active cluster
- Click Deploy Endpoint
- Verify the endpoint appears under Active
- Connect your application through the Acasia API
Next steps
Create an API key
Generate and store a credential for production use.
Connect with SSH
Open a secure shell to your cluster.
Connect your IDE
Develop against a remote GPU cluster.
.png?fit=max&auto=format&n=UdgXGI9RjuY0SENf&q=85&s=0fb4ba5917d74ae9edf418bef89524c3)
