Page layout
Deployments are organized by status tab:- Active — Currently running endpoints
- Terminated — Stopped or removed deployments
- All — Complete deployment history
Deployment fields
Endpoint lifecycle
Creating an inference endpoint
1
Click Create New Inference
Opens the deployment configuration panel.
2
Select a model
Choose from the available models in the catalog. Model availability depends on your organization’s catalog configuration.
3
Select a cluster
Choose an Active cluster with sufficient GPU capacity for the model. Large models may require multiple GPUs.
4
Configure Hugging Face token (if required)
Some models require a Hugging Face access token for download. Enter your token if prompted. Store this token securely — it is used at deployment time to pull the model weights.
5
Deploy
Submit the deployment. The endpoint enters Deploying state. Once provisioning completes, it transitions to Active and the endpoint URL becomes available.
Calling an endpoint
Once Active, the endpoint URL is available in the deployment table. Use this URL with an API key to call the model from your application.Inference endpoints use the same API key authentication as other Acasia services. Generate an API key from the API Keys page and include it in the Authorization header of your requests.
.png?fit=max&auto=format&n=UdgXGI9RjuY0SENf&q=85&s=0fb4ba5917d74ae9edf418bef89524c3)
