Skip to main content

Serverless Endpoints

Pre-configured instances of popular models hosted for free, priced per 1M tokens used. The below models are available through our inference API as serverless endpoints. Request for a model to be added to serverless endpoints or for dedicated instance or capacity for these models.

Chat Models

Use our Chat Completions endpoint for Chat Models.

Language Models

Use our Completions endpoint for Language Models.

Code Models

Use our Completions endpoint for Code Models.

Image Models

Use our Completions endpoint for Image Models.

Moderation Models

Use our Completions endpoint to run a moderation model as a standalone classifier, or use it alongside any of the other models above as a filter to safeguard responses from 100+ models, by specifying the parameter "safety_model": "MODEL_API_STRING"

Genomic Models

Use our Completions endpoint for Genomic Models. * Evo-1 models can handle up to 4096 input tokens, while output sequences can extend up to the difference between the context length and the input sequence length.

Model Request

Don’t see a model you want to use? Go to our contact page and add or upvote the model(s) you’d like to use on our API!

Dedicated Instances

Customizable on-demand deployable model instances, priced by hour hosted. All models in the serverless endpoints are available for hosting as private dedicated instances. Additionally, the below models are also available for hosting as private dedicated instances. Request an instance.

Chat Models

Language Models

Code Models

Request a model Whats Next