# Welcome on Sesterce Cloud

## About Sesterce

### Our Company

Sesterce shapes the future of AI with high-performance GPU clusters ranging from 100 to 15,000 GPUs, offering unmatched scalability and efficiency. Our flexible cloud and cluster solutions provide an optimal experience for AI builders.

### Our Mission

We're building the future of AI on the largest computing platform available, ensuring sustainability through renewable energy, and providing AI builders the best experience.

## Our offer

### Our GPUs

Sesterce On-Demand cloud platform provides its users with a wide choice of GPUs, including A100 80G, H100, L40S, L40, A6000, A5000 and RTX6000ADA.

Find [**here our tutorial to choose the instance that suits your AI workload use-case!**](/tutorials/which-compute-instance-for-ai-models-training-and-inference)

<table data-view="cards"><thead><tr><th>GPU</th><th>Hourly Price</th><th data-type="files"></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td>B200</td><td>$6,05</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=B200&#x26;numGpus=8">https://cloud.sesterce.com/compute/new?gpuType=B200&#x26;numGpus=8</a></td></tr><tr><td>H200</td><td>$3.58</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=1">https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=1</a></td></tr><tr><td>H100</td><td>$2.48</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=1">https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=1</a></td></tr><tr><td>A100 80G</td><td>$1.65</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=A100_80G&#x26;numGpus=1">https://cloud.sesterce.com/compute/new?gpuType=A100_80G&#x26;numGpus=1</a></td></tr><tr><td>A100</td><td>$1.29</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=A100&#x26;numGpus=1">https://cloud.sesterce.com/compute/new?gpuType=A100&#x26;numGpus=1</a></td></tr><tr><td>A6000</td><td>$0.57</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=A6000&#x26;numGpus=1">https://cloud.sesterce.com/compute/new?gpuType=A6000&#x26;numGpus=1</a></td></tr><tr><td>RTX4090</td><td>$0.66</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=RTX4090&#x26;numGpus=1">https://cloud.sesterce.com/compute/new?gpuType=RTX4090&#x26;numGpus=1</a></td></tr><tr><td>L40s</td><td>$1.10</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=L40S&#x26;numGpus=1">https://cloud.sesterce.com/compute/new?gpuType=L40S&#x26;numGpus=1</a></td></tr><tr><td>L40</td><td>$0.99</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=L40&#x26;numGpus=1">https://cloud.sesterce.com/compute/new?gpuType=L40&#x26;numGpus=1</a></td></tr><tr><td>A5000</td><td>$0.41</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=A5000&#x26;numGpus=1">https://cloud.sesterce.com/compute/new?gpuType=A5000&#x26;numGpus=1</a></td></tr><tr><td>RTX6000Ada</td><td>$0.97</td><td></td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=A6000&#x26;numGpus=1">https://cloud.sesterce.com/compute/new?gpuType=A6000&#x26;numGpus=1</a></td></tr></tbody></table>

### Our platform

We offer a **smooth** and **secured** GPU cloud on-demand platform, allowing our users to launch an instance in just a few clicks, with secure access to the Pod via SSH key. An all-in-one tool for training your AI model, with the ability to reuse your dataset ad infinitum thanks to the storage volumes they can create at the click of a button.


# Glossary

### Datacenter

A datacenter is a secured location where servers and GPUs are hosted. Sesterce is present [**within a large number of world class and ultra-safe datacenters**](https://www.sesterce.com/datacenters) to ensure the data security of its customers and ensure performances of its Pods.

### GPUs

A GPU (Graphics Processing Unit) is a specialized processor designed to accelerate graphics rendering and data-intensive computations. It excels in parallel processing, making it ideal for tasks like gaming, graphics design, and deep learning. 👉[**Discover how to choose your GPUs within Sesterce Cloud!**](/compute-instances#how-to-display-gpus-available)

### Pod, instance

Pods are composed by a certain number of GPUs, plus memory (RAM), a non-persistent storage space (container disk) and possibly a Volume (persistent storage space).

### Secured cloud

Secured cloud refers to cloud computing environments that are protected by robust cybersecurity measures which ensure that data and applications hosted in the cloud are safeguarded from unauthorized access and threats.

### vRAM

VRAM, or Video Random Access Memory, is a type of specialized memory used by GPUs to store image data for processing. In a way, it's your gpu's memory.

### Dataset

A dataset refers to a structured collection of data that is processed and analyzed using GPUs. These datasets are typically large and complex, making them ideal for machine learning, deep learning, and other data-intensive tasks that benefit from the parallel processing capabilities of GPUs.

### Image Template

Refers to a pre-configured virtual machine image that includes a specific operating system and installed software, optimized for GPU tasks. These templates provide a ready-to-deploy environment. 👉 [**Learn more about image templates in Sesterce Cloud!**](/compute-instances#image-configuration)

### Volume

A volume is a block storage unit attached to cloud instances for storing data. Volumes are scalable, persistent storage resources that can handle the high-throughput and low-latency requirements essential for GPU-intensive tasks, such as machine learning and data processing. 👉[ **Learn more about volumes in Sesterce Cloud!**](/compute-instances#get-a-volume)

### SSH Key

Cryptographic key used for secure access to cloud instances equipped with GPUs. It ensures a secure, password-less login to manage and operate GPU-intensive tasks remotely, enhancing security by eliminating the risks associated with password-based authentication. 👉 [**Learn more about SSH keys in Sesterce Cloud!**](/compute-instances#how-to-add-a-ssh-key)

### Top Up balance

The top-up balance is the credit balance allocated to the user's account 👉 [**How to refill credits on Sesterce Cloud?**](/welcome-on-sesterce-cloud/payment-and-billing)


# Account creation

## Create an account

To create your first Compute or Inference instance, as well as a Persistent Storage, you need to Sign Up by clicking on the bottom-left button.

<figure><img src="/files/bfbID9F47e5h00CdTryW" alt=""><figcaption></figcaption></figure>

You can create an account via email or oAuth (Github or Google)

<figure><img src="/files/aVngv9ijf9QBvg2qCiVz" alt=""><figcaption></figcaption></figure>

## Data privacy & security

Our platform strictly adheres to the General Data Protection Regulation (GDPR) to ensure the privacy and protection of our users' data. We implement all mandatory measures, including obtaining explicit user consent, providing clear data processing information, and ensuring the right to access, correct, and delete personal data. **All sensitive information is encrypted**, and we maintain robust security protocols to prevent unauthorized access. Our commitment to user privacy is paramount, and we continuously review and update our practices to remain compliant with GDPR requirements.


# Manage your account

## Update your account & billing information

You can modify your account information at any time in the "[Settings](https://cloud.sesterce.com/dashboard/settings)" tab on the platform. You can fill your first name, family name and personal adress directly from this tab.

You can also switch your account from personal to professional, and fill the name and adress of your company for billing purposes.![](/files/M2PQ6QvCTz3ZIuwcYnXT)

<figure><img src="/files/D45O50422W4pzzarSM8Y" alt=""><figcaption></figcaption></figure>


# Team accounts

### How to invite people to join my team?

to get people to join your team, share your invitation link with them from the top-left section or the [Settings page](https://cloud.sesterce.com/settings/).

<figure><img src="/files/Hhp0TsZWoSHMG8cg5UDh" alt=""><figcaption><p>Invite member from drop-down section</p></figcaption></figure>

<figure><img src="/files/1CZkq2nNnXJuazuHbs1H" alt=""><figcaption><p>Invite member from Settings page</p></figcaption></figure>

If you do not want to share access to your personal team, you can **create a new** one as well (see picture below). Define the name and the description for your new team, and share its invite link via the [method described above](#how-to-invite-people-to-join-my-team).

<figure><img src="/files/R5hA6JAMdjHrvGRjpKop" alt=""><figcaption></figcaption></figure>

### What are the different roles?

If you are sharing an invitation to your personal team or a new team you created, it means you are the team owner. Here are the differences between Owner and Member roles:

<table><thead><tr><th width="502.8828125">Action</th><th>Who can do this?</th></tr></thead><tbody><tr><td><strong>Billing Management</strong>: Credit balance and see balance amount, download invoices.</td><td>Owners and Co-Owners</td></tr><tr><td><strong>Settings Management</strong>: See and update settings linked to the team (Billing address, Account type), <a href="/pages/AEkuVDl7csz5b91OKHsk">create and share API Keys</a>, see user logs. </td><td>Owners only</td></tr><tr><td><strong>Resources Management</strong>: Launch compute and inference instances and create SSH Keys to access it, create Volumes, see running elements (Name, IP, etc.)</td><td>Owners, Co-Owners and Members</td></tr></tbody></table>

### What happens if my team does not have enough credits?

When you want to launch an instance or create a volume, you can select the team associated to this deployment in Checkout section. By default, the team selected top left is also selected here.

If the team selected does not have enough credits for deploying the resource, you'll receive an error log accordingly.

<figure><img src="/files/9QyED80JbdyDTq0x12kD" alt=""><figcaption></figcaption></figure>


# Payment & Billing

## What is the type of pricing supported on Sesterce Cloud?

Currently, Sesterce Cloud offers “**On-Demand**” pricing. It means that you’ll pay “as you go”, according to your subscription(s), **without the risk of your pod being interrupted**.

## How to refill credits?&#x20;

### Manual Top-Up

You can fill or refill your credits balance through “[Billing](https://cloud.sesterce.com/dashboard/billing)” tab, with your credit card. Our Payment Service Provider is Stripe, the ultra-secure market leader.

<figure><img src="/files/ga5NOgin9RO0fLujkFNF" alt=""><figcaption></figcaption></figure>

### Auto Top-Up

You want to launch instances for a long time, and you don't want to worry about recharging? That's why we've set up the auto-top up!

Activate the option through the dedicated toggle on Billing page, and set-up your Automatic Top-Up: the balance will be automatically recharge when your credit amount will be lower than the limit indicated (see below).

<figure><img src="/files/L741H7Yror0pyC6rf7Sa" alt=""><figcaption></figcaption></figure>

## What are the payments methods available?

Following **payment methods** are available:

* Credit Card
* Bank Transfer
* Cash App Pay
* Google Pay
* Link

### How to pay via Bank Transfer?

It is now possible to pay on Sesterce Cloud via Bank Transfer, whatever your location and your bank. When clicked on "Add Credit" button from Billing page, or on the bottom-left balance section, you'll open the following modal. Please click on "Bank Transfer" option.

<figure><img src="/files/zDSWY35EqPfXHZV66vHR" alt=""><figcaption></figcaption></figure>

Then, proceed to the Bank Transfer with the following bank account information. Upload a screen record of your transfer, and click "Place Order".

{% hint style="info" %}
Please note that your credit balance will be updated **once Sesterce will receive the funds**. It usually takes 1 to 3 business days. We recommend this payment method to our customers with high consumption (over $1,000 per month). You will receive **an email notification once your balance is credited**.
{% endhint %}

<figure><img src="/files/IlPKxiSrCzwC5Bf676yG" alt=""><figcaption></figcaption></figure>

## How is my balance debited?

Your credit balance is debited in real time, every minute, according to your instances and volumes utilization.&#x20;

## What happens if I do not have any credits left?

Having credits on your account is essential to ensure the smooth running of your proceedings and the storage of your dataset. Sesterce will notify you by email if your credit balance becomes too low.

If your credit balance drops to 0, our team will be forced to **stop your current instances** and **delete your storage volumes**, if you have any.


# Invoicing

## Where can I find my invoices?

You can find your invoices within “[**Billing**](https://cloud.sesterce.com/dashboard/billing)” tab. A button is displayed, allowing you to download your invoices. Two types of invoices are available:

{% tabs %}
{% tab title="Top Up Invoice" %}
These invoices are generated when you refill your top up balance.
{% endtab %}

{% tab title="Monthly Invoice" %}
These invoices compile all subscriptions you performed during the past month
{% endtab %}
{% endtabs %}

<figure><img src="/files/18IYpAg5zQ09foxYifsI" alt=""><figcaption></figcaption></figure>

## Is there a way to find my Purchase History?

At Sesterce, we want you to know exactly what you are paying for. This is why we implemented the Billing History feature, which allow you to see all invoices that have been generated for your account.

This is perfect if you need to use Sesterce Cloud for your company!

<figure><img src="/files/MlDC9sDHwunCfMVfgsly" alt=""><figcaption></figcaption></figure>


# Compute instances

## How to display GPUs available?

On the Compute page, you'll find a list of all our available instances, sorted by model and number of GPUs. Filters are available to help you find the right offer for your needs: [**read this tutorial to know which offer you should choose according to your workload and use case!**](/tutorials/which-compute-instance-for-ai-models-training-and-inference)

<figure><img src="/files/Iz1gNwy9UKrrLzuVcTtq" alt=""><figcaption></figcaption></figure>

## Pods characteristics

<table><thead><tr><th width="263">Characteristics</th><th>Definition</th></tr></thead><tbody><tr><td>vRAM</td><td>Virtual Random Access Memory: The amount of memory allocated to a virtual machine for its operations, functioning similarly to physical RAM in a physical machine.</td></tr><tr><td>Memory in GB</td><td>The total amount of memory available for use in a system, typically measured in gigabytes (GB), which can be used for running applications and processes.</td></tr><tr><td>vCPU</td><td>Virtual Central Processing Unit: An abstraction of physical CPU cores, allocated to virtual machines to handle computational tasks.</td></tr><tr><td>Storage</td><td>The space available for storing data, applications, and system files, measured in gigabytes (GB) or terabytes (TB).</td></tr><tr><td>NVLink</td><td><p>A high-speed interconnect technology developed by NVIDIA to enable fast communication between GPUs, and between GPUs and CPUs, for improved performance and efficiency in </p><p>computational tasks.</p></td></tr></tbody></table>

<details>

<summary>⚠️ The Storage of your Pod is non persistent. </summary>

It means that it will only host your data during the instance. Then, if you haven't subscribed to a volume, your dataset will be deleted. Learn [**how to get a volume in Sesterce Cloud!**](#how-to-get-a-volume)

</details>

## How to select my Pod?

Through “GPU” and “Specs” tabs, you can display information corresponding to each Pod. Please find above the corresponding definitions. Then, you can click “launch”!

## I need a persistent storage!

If you need a persistent storage to store your dataset at the end of your instance, you can thick the appropriate checkbox to display only Pods for which the persistent storage is available.

<figure><img src="/files/7xM9EopcijOEyqvmlvVp" alt=""><figcaption></figcaption></figure>

<details>

<summary>Please consider those rules for Volumes management. <a href="/pages/DanaahABEKYGhBuMui5q">See dedicated documentation here. </a></summary>

* Volumes should be hosted **in the same Cloud** and **region** as Pod selected
* Once created, volume size can be increased but not decreased, to preserve your data
* A single volume can contain several datasets, but **it is not possible to link a single Pod with more than one Volume**.
* It is **not possible** to link a persistent storage to an instance **already created**.

</details>

{% hint style="success" %}
According to the Pod you'll choose, you will have the possibility to create a volume directly from the launching path, hosted in the appropriate cloud & region.
{% endhint %}


# Configure your Compute Instance

## Fill Machine Name

The first step in configuration is to choose the instance name, using the text field at the top of the page.

## Select Region

Next, you can choose the region in which the instance will be deployed. Bear in mind that different specifications (RAM and Storage non-persistent) are available for different regions.

<figure><img src="/files/pcWETlkqn2FKMKqVl8qF" alt=""><figcaption></figcaption></figure>

## Link Persistent Volume

Depending on the region you've chosen, you can create a Volume (storage persistent) or link to a previously created Volume. To filter instances that accept persistent volumes, [please select the corresponding option on the Compute page](/compute-instances#i-need-a-persistent-storage).

{% hint style="danger" %}
You'll not be able to link a Volume to the instance after its creation: this must be done at the time of launch. Find out more details about how to [create and connect persistent storage to an instance](/compute-instances/configure-your-compute-instance/persistent-storage-volumes).
{% endhint %}

## Choose OS and Images

The OS availables are displayed in the dedicated section (see below). Note that you can also set-up your VM to be launched with an specific application, like TensorFlow or Pytorch for training, as well as ML Models for inference (see below)

<figure><img src="/files/JKvKtklLYA40zqDWDSrP" alt=""><figcaption></figcaption></figure>

## Set-up SSH Key

To connect your instance, you'll need to generate a SSH Key through the section below.&#x20;

<figure><img src="/files/uCvuIjXZNMjGOqR1xqi7" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
To generate your public SSH key from your laptop, open your terminal and type the following command: \
`ssh-keygen -t ed25519 -C "youremail@example.com"` . Then, type enter to select by-default saving place and use the following command to expose your public key: `cat ~/.ssh/id_ed25519.pub`
{% endhint %}


# SSH Keys management

To launch an instance on Sesterce Cloud, you need to link a SSH Key to it. You can paste your public key directly from the launching path, or [**via the Settings page**](https://cloud.sesterce.com/settings).

{% hint style="danger" %}
Currently, it is **not possible** to add a SSH Key to an instance already live. Make sure to add all keys you need before launching your compute instance. [You can add several keys to a single instance](#how-to-add-multiple-ssh-key-in-a-compute-instance).
{% endhint %}

<figure><img src="/files/QOrX1J1VUyyODRO6cVXV" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
To generate your public SSH key from your laptop, open your terminal and type the following command: \
`ssh-keygen -t ed25519 -C "youremail@example.com"` . Then, type enter to select by-default saving place and use the following command to expose your public key: `cat ~/.ssh/id_ed25519.pub`
{% endhint %}

### How to mark a SSH Key as default?

To mark a SSH Key as default, click on the 3 dots icon an select "Make Default". This will automatically select the SSH Key when you'll create a new Compute instance.

<figure><img src="/files/BVkdFgja7GoUCQraU21b" alt=""><figcaption></figcaption></figure>

### How to add multiple SSH Key in a Compute instance?

You can merge several SSH Keys by pasting the keys one after the other in the text field as below (make sure to brake line between the keys):

<figure><img src="/files/FT3vuUAEd7Vedr0aOzrh" alt=""><figcaption></figcaption></figure>

Then, select the SSH Key created to launch your instance. Both users will be able to connect to the Compute instance.

<figure><img src="/files/FlhNax1uiyktnVuFPmqs" alt=""><figcaption></figcaption></figure>


# Persistent storage (volumes)

## How to link a persistent storage (volume) to my Compute instance?

If you need to store your dataset after the end of your instance, you need to add a volume to your Pod. There are two ways to create a volume on Sesterce Cloud:

* From "Storage" tab
* Directly from the instance creation path

If you don't have volume created yet, we suggest the second option. All you need to do is to make sure that the offers displayed are [supporting volumes by ticking the appropriate box](/compute-instances#i-need-a-persistent-storage).

To create a volume, select your volume name and the size you need (in GB).&#x20;

### How to create a persistent volume from Storage page?

First, click on "Create volume" or "New volume" button (see below)

<figure><img src="/files/DJSBvh7Y8aEW1sCIyDrE" alt=""><figcaption></figcaption></figure>

#### Volume name and region

Then, choose Volume name via the corresponding text field and Region where the volume will be hosted.

<figure><img src="/files/YlZV5QXcDNYaRO6sjTWU" alt=""><figcaption></figcaption></figure>

#### Availability Zone

According to the Region selected, several AZ will be available or note. Please note that you'll be able to **link your Volume only to Compute instances that are launched in the same Regions and Availability Zone.**

### How to link the Volume to my Compute instance?

To simplify the process to its maximum, we implemented the following path: once your Volume is created, you can launch an instance via a dedicated CTA from volume page directly.&#x20;

#### Click on Volume Card

<figure><img src="/files/4tRoQfkkU0Wd1nUvSBfs" alt=""><figcaption></figcaption></figure>

#### Click "Attach GPU Instance"

<figure><img src="/files/hDYFaR0wr8thd5UDEo2U" alt=""><figcaption></figcaption></figure>

You will be redirected to the Compute page, displaying all instances offers that match with your Persistent Volume location.

#### **Select the Volume**

In the instance creation page, you'll find your Volume displayed in the dedicated section. Select it to link your instance to this persistent storage.

<figure><img src="/files/ISvkAH2feGUvaY2c2B7L" alt=""><figcaption></figcaption></figure>

### How to verify the Volume on the instance?

Once the instance is active, connect into it [through SSH command](/compute-instances/terminal-connection). Then, run the following command:

```
# List all block devices attached to the instance
lsblk
```

You should have an output similar to this:

```
NAME    MAJ:MIN RM   SIZE RO TYPE MOUNTPOINTS
loop0     7:0    0  63.9M  1 loop /snap/core20/2318
loop1     7:1    0    87M  1 loop /snap/lxd/28373
loop2     7:2    0  38.8M  1 loop /snap/snapd/21759
loop3     7:3    0  44.3M  1 loop /snap/snapd/23258
loop4     7:4    0  63.7M  1 loop /snap/core20/2434
vda     252:0    0   128G  0 disk 
├─vda1  252:1    0 127.9G  0 part /
├─vda14 252:14   0     4M  0 part 
└─vda15 252:15   0   106M  0 part /boot/efi
vdb     252:16   0    50G  0 disk # this is the volume I mounted
nvme0n1 259:0    0 894.3G  0 disk
```

This command will display all block devices attached to the instance. You should see your volume listed as a new device, such as `vdb`, `sbd`, or similar.

### How to mount the Volume to a filesystem?

To make the storage space accessible and usable by your operating system and applications, you need to mount the volume to a filesystem using the following command.

```
sudo mkdir /mnt/vdb
sudo mkfs.ext4 /dev/vdb
sudo mount -t ext4 /dev/vdb /mnt/vdb/
sudo chown -R sesterce:sesterce /mnt/vdb/
```

Now, the `/mnt/vdb` directory should be connected to your volume and fully accessible for you to write and read from.

<br>


# Terminal connection

When your instance is launched, you'll be able to find it on your Compute tab, above the list of offers available. Click the instance card to display details about terminal connection!

<figure><img src="/files/LseQEzfmnOEjyHjIa2hv" alt=""><figcaption></figcaption></figure>

You can then copy your SSH Command to connect the server!

<figure><img src="/files/OWMugUunPPQNRgQS9UTY" alt=""><figcaption></figcaption></figure>


# AI Inference instances

## What is Sesterce AI Inference service?

We built our inference feature to enable our users to bring their ML model to life by deploying it in a dedicated production environment accessible to all via an endpoint.

You can use our inference service to **deploy your own custom model** to make it accessible to your users, or to [**infer with the best models**](#pre-charged-public-models) on the market which you can then seamlessly integrate into your applications.

{% hint style="info" %}
In addition to the [classic compute instances](/compute-instances) that are very useful for building and training your model, the **AI inference feature** is an additional brick that **will allow you to manage also the deployment** of your ML Model **as closely as possible to your users**.
{% endhint %}

## Deploy your Model as closely as possible to your customers

Sesterce's AI inference feature allows you to deploy your model as close as possible to your end-users, to guarantee minimal latency, here's why:

{% stepper %}
{% step %}

### Edge inference nodes distributed around the world

Processes data locally at the network's edge, minimizing latency and bandwidth usage for real-time applications.
{% endstep %}

{% step %}

### Anycast Endpoint setupped automatically

Directs user requests to the nearest instance of a service, optimizing performance and reducing response time.
{% endstep %}

{% step %}

### Smart Routing technology to 180 points of presence worldwide&#x20;

End users' queries are routed to the closest active model, ensuring low latency and an improved user experience.
{% endstep %}
{% endstepper %}

## Pre-charged public models

Here is a non exhaustive list of models available on Sesterce Cloud. [Click here to discover the entire model catalog](https://cloud.sesterce.com/ai-inference)!

<table data-full-width="false"><thead><tr><th width="202">Model</th><th width="162.33333333333331">Type</th><th>Description</th></tr></thead><tbody><tr><td>distilbert-base</td><td>Text processing</td><td>A smaller, faster version of BERT used for natural language tasks.</td></tr><tr><td>stable-diffusion</td><td>Text-to-image</td><td>Generates images from text descriptions using deep learning techniques.</td></tr><tr><td>stable-cascade</td><td>Text-to-image</td><td>Enhances image generation with multiple refinement steps.</td></tr><tr><td>sdxl-lightning</td><td>Text-to-image</td><td>Optimized for fast image generation from text inputs.</td></tr><tr><td>ResNet-50</td><td>Image classification</td><td>A convolutional neural network designed for image recognition tasks.</td></tr><tr><td>Llama-Pro-8b</td><td>Text generation</td><td>A large language model designed for generating human-like text.</td></tr><tr><td>Llama-3.2-3B-Instruct</td><td>Text generation</td><td>An instruction-tuned model for generating text with specific guidelines.</td></tr><tr><td>Mistral-Nemo-Instruct-2407</td><td>Text generation</td><td>Tailored for creating text based on given instructions.</td></tr><tr><td>Llama-3.1-8B-Instruct</td><td>Text generation</td><td>An advanced model for generating text with detailed instructions.</td></tr><tr><td>Pixtral-12B-2409</td><td>Text-to-image</td><td>Produces high-quality images from text prompts using a large model.</td></tr><tr><td>Llama-3.2-1B-Instruct</td><td>Text generation</td><td>Focused on generating text according to user-provided instructions.</td></tr><tr><td>Mistral-7B-Instruct-v.0.3</td><td>Text generation</td><td>Designed for generating guided text outputs with minimal latency.</td></tr><tr><td>Whisper-large-V3-turbo</td><td>Audio-to-text</td><td>Quickly transcribes audio into text with high accuracy.</td></tr><tr><td>Whisper-large-V3</td><td>Audio-to-text</td><td>Transcribes spoken language into written text using deep learning.</td></tr></tbody></table>


# Inference Instance configuration

## Select your model

### How to select Public model?

#### Model catalog

If you want to infere with one of the best-known models, you can select it in our catalog of pre-charged models list.

<figure><img src="/files/pFtCCZCakH256rMBYu7v" alt=""><figcaption></figcaption></figure>

#### Public custom model

If you want to select a custom model that is publicly hosted on Docker Hub, for example, click "New Deployment" and fill your docker tag in the "Public Model" text field.

<figure><img src="/files/9YCxgg1zSur8Hec74BYE" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/2caIZ9sy9qGaHCxnahs1" alt=""><figcaption></figcaption></figure>

### How to use Private custom model?

If you want to deploy a private custom template, you'll need to create a registry. Click on “Private Model", and add your Registry informations.

<figure><img src="/files/jmzdypH2nmooQ8sGQa7u" alt=""><figcaption></figcaption></figure>

To create a registry, perform the **following steps**:

{% stepper %}
{% step %}

### Registry name

Give your registry a name consisting of lowercase Latin characters, which can be separated by dashes.
{% endstep %}

{% step %}

### Location link

Provide the link to the location where your AI model is stored. We’ll use this URL to retrieve the model during deployment.
{% endstep %}

{% step %}

### Username

Specify the username you use to access the storage location of your AI model.
{% endstep %}

{% step %}

### Password

Enter the password required to access the model.
{% endstep %}
{% endstepper %}


# Select your Flavor

### When to choose GPU or CPU flavors?&#x20;

According to your needs, you can choose from two options: CPU and GPU Flavors.&#x20;

{% tabs %}
{% tab title="When to choose GPU Flavor?" %}
GPU Flavor is ideal for **inference of complex deep learning models**, image processing, or multimedia content generation.
{% endtab %}

{% tab title="When to choose GPU flavor?" %}
CPU Flavor is more **dedicated to light tasks** as simple text processing algorithms or structured data processing, that do not require a very low latency time.
{% endtab %}
{% endtabs %}

### Which GPU Flavors are available for AI Inference?&#x20;

<table data-card-size="large" data-view="cards" data-full-width="false"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-cover data-type="files"></th></tr></thead><tbody><tr><td><strong>NVIDIA L40S</strong></td><td>GPU specialized for inference tasks able to accelerate multiple workloads.</td><td></td><td><a href="/files/i2CeSBskEZt7sRGVoWF1">/files/i2CeSBskEZt7sRGVoWF1</a></td></tr><tr><td><strong>NVIDIA H100 TensorCore</strong></td><td>Up to 30 times acceleration of LLM processing. Ideal for complete models with up to 30 billion parameters.</td><td></td><td><a href="/files/GUHRlyhTe5hqsFyMrfrD">/files/GUHRlyhTe5hqsFyMrfrD</a></td></tr><tr><td><strong>NVIDIA A100 TensorCore</strong></td><td>A100 provide up to 20X higher performance over the NVIDIA Volta with zero code changes and an additional 2X boost with automatic mixed precision and FP16</td><td></td><td><a href="/files/KnDn69dTaCQlRBEBSUTZ">/files/KnDn69dTaCQlRBEBSUTZ</a></td></tr></tbody></table>


# Select your regions

Sesterce AI inference allows you to deploy your model close to your final users. According to the flavor chosen, you'll be able to select one or several regions.

<figure><img src="/files/J8H97TWbEHu6EDyGlMrP" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Our edge nodes spread all over the world work as pick-up points to redirect your final users to the closest inference instance.
{% endhint %}


# Autoscaling limits

With Sesterce Cloud's AI inference feature, you can define autoscaling limits to manage the scale of your model while keeping your expenses under control.

To define your autoscaling limit, you can enter the following main two elements:

1. The minimum and maximum number of pods to be mobilized&#x20;
2. The triggers to scale automatically your ressources

### How to define autoscaling limits?

From advanced configuration settings, define the minimum and maximum number of pods to be mobilized to infere your model.&#x20;

If you want to deploy your model in several regions, you can choose to set the same limits in each of them, or specify specific limits for each region.

{% hint style="info" %}
This is particularly useful if you expect to have a peak of users on your endpoint in a specific region.
{% endhint %}

<figure><img src="/files/KMJyyfkD67PNUolO7xVR" alt=""><figcaption></figcaption></figure>

### What is the Cooldown Period?&#x20;

The cooldown period corresponds to intervals (in seconds) the between trigger executions. This setting helps avoid frequent and unnecessary scaling adjustments. You can select a value between 1 and 3600 seconds.

### How to set autoscaling triggers?

From the dedicated form (see below), choose among the 3 available triggers type:

* GPU memory utilization
* Memory utilization
* RAM usage
* HTTP requests

<figure><img src="/files/IuolWUkIDUqx376vRBLt" alt=""><figcaption></figcaption></figure>

For each of these triggers types, you can set a threshold value in percents. When this value is reached, then pod numbers increase within limits and decrease to the initial level when values drop, ensuring stable operation and high performance.

{% hint style="info" %}
You can configure up to 3 threshold limits corresponding to each of the 3 triggers types.
{% endhint %}

The minimum setting is **1% of the resource capacity**, and **only the HTTP requests trigger can scale pods to and from 0.** The maximum setting is 100% of the resource capacity.


# Edit an inference instance

### How to edit an inference instance?

Even after launching, it is possible to update the settings of your inference instances. This is particularly relevant if you need to edit your autoscalling limits, according to the usage of your endpoint.

From Inference, select "Edit". Then, fill up the form with the new parameters to want to pass into your instance.

<figure><img src="/files/vwKuIj3ZiWppAW2kBOfL" alt=""><figcaption></figcaption></figure>

### What can I edit from a running inference instance?

Here are the element that can be customized from an Running inference instance:

* GPU/CPU Flavor
* Region
* Startup Command
* Containers and autoscaling triggers
* Environment variables


# Chat with Endpoint

## From Chat Playground

When your inference instance turns active, a button "Open Playground" will appear if the model hosted is Open AI compatible.

Click the button to access Chat Playground! You'll be able to interact with your endpoint, and change parameters such as Temperature, Top-p, Top-k and repetition penalty.

<figure><img src="/files/ExpaTiQEp3KtZBKQYTek" alt=""><figcaption></figcaption></figure>

## From your Terminal

To interact with the endpoint directly from your terminal, you'll need to follow this process. Our endpoints follow OpenAI's specification, enabling seamless integration with existing tools and libraries.

### Prerequisites

* Your endpoint URL (format: `https://<id>-<hash>.ai.sesterce.dev/`)
* Your API Secret (provided upon deployment)
* Model ID (retrieved via API)
* OpenAI-compatible client or SDK

### Authentication

#### Verifying connection

First, list available models:

```
curl -H "x-api-key: <SECRET>" -X GET "<ENDPOINT>/v1/models"
```

### Model types and endpoints

{% hint style="info" %}
Ensure to replace \<SECRET> by your own secret (available from launched instance page), as well as \<MODEL\_ID> and \<ENDPOINT>, to be replaced by your own data.
{% endhint %}

#### 1. Text generation

Endpoint: `/v1/chat/completions`

```
curl -H "Content-Type: application/json" \
     -H "x-api-key: <SECRET>" \
     -X POST "<ENDPOINT>/v1/chat/completions" \
     -d '{
       "model": "<MODEL_ID>",
       "messages": [
         {
           "role": "user",
           "content": "Hello, how are you?"
         }
       ]
     }'
```

#### 2. Multimodel (text+image)

Endpoint: `/v1/chat/completions`

```
curl -H "Content-Type: application/json" \
     -H "x-api-key: <SECRET>" \
     -X POST "<ENDPOINT>/v1/chat/completions" \
     -d '{
       "model": "<MODEL_ID>",
       "messages": [
         {
           "role": "user",
           "content": [
             {
               "type": "text",
               "text": "What's in this image?"
             },
             {
               "type": "image_url",
               "image_url": {
                 "url": "https://example.com/image.jpg"
               }
             }
           ]
         }
       ]
     }'
```

#### 3. Audio-Speech Recognition (ASR)

Endpoint: `/v1/audio/transcriptions`

```
curl -H "x-api-key: <SECRET>" \
     -X POST "<ENDPOINT>/v1/audio/transcriptions" \
     -H "Content-Type: multipart/form-data" \
     -F file="@/path/to/audio.mp3" \
     -F model="<MODEL_ID>"
```

### Open AI SDK integration

#### Javascript/Typescript

```javascript
import OpenAI from "openai";

const openai = new OpenAI({
    apiKey: "<SECRET>",
    baseURL: "<ENDPOINT>/v1"
});

// Text Generation
async function generateText() {
    const completion = await openai.chat.completions.create({
        model: "<MODEL_ID>",
        messages: [
            {
                role: "user",
                content: "Hello, how are you?"
            }
        ]
    });
    console.log(completion.choices[0].message.content);
}

// Multimodal
async function analyzeImage() {
    const response = await openai.chat.completions.create({
        model: "<MODEL_ID>",
        messages: [
            {
                role: "user",
                content: [
                    {
                        type: "text",
                        text: "What's in this image?"
                    },
                    {
                        type: "image_url",
                        image_url: {
                            url: "https://example.com/image.jpg"
                        }
                    }
                ]
            }
        ]
    });
}
```

#### Python

```python
from openai import OpenAI

client = OpenAI(
    api_key="<SECRET>",
    base_url="<ENDPOINT>/v1"
)

# Text Generation
response = client.chat.completions.create(
    model="<MODEL_ID>",
    messages=[
        {"role": "user", "content": "Hello!"}
    ]
)

# Audio Transcription
with open("audio.mp3", "rb") as audio_file:
    transcript = client.audio.transcriptions.create(
        model="<MODEL_ID>", 
        file=audio_file
    )
```

### Best Practices

1. **Performance Optimization**
   * Use batch processing for multiple requests
   * Implement caching when possible
   * Compress files before upload
   * Use URLs for large files
2. **Security**
   * Never share your API Secret
   * Implement rate limiting
   * Monitor usage patterns
   * Secure stored credentials
3. **Integration Tips**
   * Always validate model availability
   * Implement proper error handling
   * Use the SDK for robust integration
   * Keep dependencies updated

### Supported Formats and Limitations

#### File Formats

* Images: JPEG, PNG, WEBP, GIF
* Audio: mp3, mp4, mpeg, mpga, m4a, wav, webm

#### Size Limits

* Images: Maximum 20MB
* Audio: Maximum 25MB
* Audio Duration: Up to 4 hours

### Error Handling

#### Common Error Codes

* 401: Invalid API Secret
* 404: Model not found
* 429: Too many requests
* 500: Server error

#### Error Handling example

```
try:
    response = client.chat.completions.create(...)
except Exception as e:
    if "file too large" in str(e):
        # Handle size error
    elif "unsupported file type" in str(e):
        # Handle format error
    else:
        # Handle other errors
```

If you need support, please reach us at <support@sesterce.com>.


# Manage your instances

## What could be the different status of my instance?

{% tabs %}
{% tab title="Pending" %}
When your instance is about to be launched, its status will be "Pending" until you are connected to your Pod and it is started.

💡 Your balance **is not debitted** when your Pod is under “Pending” status
{% endtab %}

{% tab title="PartiallyDeployed" %}
When your inference instance is PartiallyDeployed, it means the deployment is not finished yet. You'll need to wait few minutes before turning Active.

💡 Your balance **is not debitted** when your Pod is under “Pending” status
{% endtab %}

{% tab title="Active" %}
When your Inference instance is running, its status will be “Active”.

💡 Your balance **is debitted** when your Pod is under “Running” status
{% endtab %}

{% tab title="Deleted" %}
When you stop your Pod, it will automatically go through “Deleted” status. Your dataset will be stored (if you subscribed to a volume) or deleted.

💡 Your balance **is not debitted** when your Instance is under “Deleted” status
{% endtab %}
{% endtabs %}

## How to see my running instances?

You'll find all your running instances within the Inference section. You'll be able to fill your instances via its name, IP adress, Region and so on!

## How to stop a running instance?

You can stop your running instances at any time via the Running Instance section, by clicking "Stop" button through "Action" column.&#x20;


# API Reference

## What is the purpose of Sesterce Cloud API?

Our API has been built to allow our users to launch and manage their instances, persistent storages and SSH Keys via an API rather than our UI console.

### Prerequisites

To be able to use Sesterce Cloud API Key, please follow these 2 prerequisites:

1. [Create an account](/welcome-on-sesterce-cloud/account-creation) and top-up your balance through [Billing tab](https://cloud.sesterce.com/billing)
2. Generate and retrieve your API key through [Settings tab](https://cloud.sesterce.com/settings)


# Authentication

### How to generate my API Key?

From Setting page > API Key section, click "Generate API Key" and name your key. The key will be automatically generated.

{% hint style="danger" %}
Copy your API key and store it in a secured place. You'll not be able to see it again then, to avoid any security issue.
{% endhint %}

<figure><img src="/files/enMxNyY97hYMnRKYzF9t" alt=""><figcaption></figcaption></figure>

Then, use your API key as in the following request to authenticate:

{% tabs %}
{% tab title="cURL" %}

```json
curl --request GET \
--url https://api.cloud.sesterce.com/gpu-cloud/instances/offers \
--header 'X-API-KEY: <your-api-key-secret>'
```

{% endtab %}
{% endtabs %}


# GPU Cloud instances

The following endpoints allow to create and manage GPU Cloud instances from the API.

### Get the list of available offers

Use this endpoint when you want to explore the different GPU instance options available for your project. This is particularly useful when planning new deployments and needing to compare offers to find the best fit in terms of cost and performance.

{% hint style="info" %}
Double-check all inputs in the request body to ensure successful creation.
{% endhint %}

## GET /gpu-cloud/instances/offers

>

```json
{"openapi":"3.0.0","info":{"title":"Cloud Sesterce API","version":"1.0"},"paths":{"/gpu-cloud/instances/offers":{"get":{"operationId":"GPUCloudInstancesController_getInstancesOffers","parameters":[{"name":"x-api-key","in":"header","description":"The API Key secret should be sent through this header to authenticate the request.","required":true,"schema":{"type":"string"}},{"name":"region","required":false,"in":"query","schema":{"type":"string"}},{"name":"numGpus","required":false,"in":"query","schema":{"type":"string"}},{"name":"gpuType","required":false,"in":"query","schema":{"type":"string"}},{"name":"available","required":false,"in":"query","schema":{"type":"boolean"}},{"name":"sort","required":false,"in":"query","schema":{"enum":["price"],"type":"string"}},{"name":"deploymentType","required":false,"in":"query","schema":{"enum":["vm","container","baremetal"],"type":"string"}}],"responses":{"200":{"description":"Returns the list of available offers for instances","content":{"application/json":{"schema":{"type":"array","items":{"$ref":"#/components/schemas/InstanceOfferDto"}}}}},"403":{"description":"API key invalid","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}},"404":{"description":"Not found","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}}},"tags":["GPUCloudInstances"]}}},"components":{"schemas":{"InstanceOfferDto":{"type":"object","properties":{"gpuName":{"type":"string"},"gpuCount":{"type":"number"},"nvlink":{"type":"boolean"},"deploymentType":{"type":"string"},"instanceId":{"type":"string"},"cloudInitAvailable":{"type":"boolean"},"cloud":{"$ref":"#/components/schemas/CloudProviderDto"},"configuration":{"$ref":"#/components/schemas/ConfigurationInstanceDto"},"hourlyPrice":{"type":"number"},"availability":{"type":"array","items":{"$ref":"#/components/schemas/AvailabilityDto"}}},"required":["gpuName","gpuCount","nvlink","deploymentType","instanceId","cloudInitAvailable","cloud","configuration","hourlyPrice","availability"]},"CloudProviderDto":{"type":"object","properties":{"_id":{"type":"string"},"name":{"type":"string"}},"required":["_id","name"]},"ConfigurationInstanceDto":{"type":"object","properties":{"ramGB":{"type":"number"},"storageGB":{"type":"number"},"vCpu":{"type":"number"},"vRamGB":{"type":"number"},"os":{"type":"array","items":{"type":"string"}},"interconnect":{"type":"string"}},"required":["ramGB","storageGB","vCpu","vRamGB","os","interconnect"]},"AvailabilityDto":{"type":"object","properties":{"region":{"type":"string"},"name":{"type":"string"},"countryCode":{"type":"string"},"available":{"type":"boolean"}},"required":["region","name","countryCode","available"]}}}}
```

### Create a GPU Cloud instance

Use this endpoint when you're ready to launch a new GPU instance for a specific project or task. This is the crucial step to deploy new computing resources.

{% hint style="info" %}
You can use query parameters to fine-tune search results and find the best offer for your specific needs
{% endhint %}

## POST /gpu-cloud/instances

>

```json
{"openapi":"3.0.0","info":{"title":"Cloud Sesterce API","version":"1.0"},"paths":{"/gpu-cloud/instances":{"post":{"operationId":"GPUCloudInstancesController_create","parameters":[{"name":"x-api-key","in":"header","description":"The API Key secret should be sent through this header to authenticate the request.","required":true,"schema":{"type":"string"}}],"requestBody":{"required":true,"content":{"application/json":{"schema":{"$ref":"#/components/schemas/CreateInstanceDto"}}}},"responses":{"201":{"description":"Return the created instance","content":{"application/json":{"schema":{"$ref":"#/components/schemas/InstanceDto"}}}},"403":{"description":"API key invalid","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}},"404":{"description":"Not found","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}}},"tags":["GPUCloudInstances"]}}},"components":{"schemas":{"CreateInstanceDto":{"type":"object","properties":{"name":{"type":"string"},"cloudProvider":{"type":"string"},"instanceId":{"type":"string"},"volumeId":{"type":"string"},"region":{"type":"string"},"vm":{"$ref":"#/components/schemas/VirtualMachineDto"},"dockerContainer":{"$ref":"#/components/schemas/DockerContainerDto"},"sshKeyId":{"type":"string"},"sshKey":{"type":"string"}},"required":["name","cloudProvider","instanceId","region"]},"VirtualMachineDto":{"type":"object","properties":{"os":{"type":"string"},"base64CloudInitScript":{"type":"string","description":"Base64 encoded cloud-init script that will be executed during instance initialization"}},"required":["os"]},"DockerContainerDto":{"type":"object","properties":{"image":{"type":"string"}},"required":["image"]},"InstanceDto":{"type":"object","properties":{"_id":{"type":"string"},"name":{"type":"string"},"provider":{"type":"string"},"region":{"type":"object"},"volumes":{"type":"array","items":{"type":"array"}},"status":{"type":"string","enum":["pending","active","error","deleted","deleting"]},"gpuCount":{"type":"number"},"gpuModel":{"type":"string"},"ram":{"type":"number"},"storage":{"type":"number"},"vramPerGpu":{"type":"number"},"vcpus":{"type":"number"},"interconnect":{"type":"string"},"nvlink":{"type":"boolean"},"os":{"type":"string"},"ip":{"type":"string"},"sshUser":{"type":"string"},"sshPort":{"type":"number"},"dockerImage":{"type":"string"},"dockerCommand":{"type":"string"},"hourlyPrice":{"type":"number"},"deletedAt":{"format":"date-time","type":"string"},"createdAt":{"format":"date-time","type":"string"},"updatedAt":{"format":"date-time","type":"string"}},"required":["_id","name","provider","region","volumes","status","gpuCount","gpuModel","ram","storage","vramPerGpu","vcpus","interconnect","nvlink","os","ip","sshUser","sshPort","dockerImage","dockerCommand","hourlyPrice","deletedAt","createdAt","updatedAt"]}}}}
```

{% hint style="warning" %}
Make sure you filled the right OS name, such as `ubuntu24.04_cuda12.4_shade_os`. You can retrieve the os available with the `/gpu-cloud/instances/offers` endpoint.
{% endhint %}

### Get the list of instances created

This endpoint is essential for users who want an overview of all the instances they have created. It is ideal for managing and monitoring current resources.

{% hint style="info" %}
Ensure your API key is active and correctly entered to view your instances.
{% endhint %}

## GET /gpu-cloud/instances

>

```json
{"openapi":"3.0.0","info":{"title":"Cloud Sesterce API","version":"1.0"},"paths":{"/gpu-cloud/instances":{"get":{"operationId":"GPUCloudInstancesController_getUserInstances","parameters":[{"name":"x-api-key","in":"header","description":"The API Key secret should be sent through this header to authenticate the request.","required":true,"schema":{"type":"string"}},{"name":"deploymentType","required":false,"in":"query","schema":{"enum":["vm","container","baremetal"],"type":"string"}}],"responses":{"200":{"description":"Returns the list of user instances","content":{"application/json":{"schema":{"type":"array","items":{"$ref":"#/components/schemas/InstanceDto"}}}}},"403":{"description":"API key invalid","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}},"404":{"description":"API key not found","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}}},"tags":["GPUCloudInstances"]}}},"components":{"schemas":{"InstanceDto":{"type":"object","properties":{"_id":{"type":"string"},"name":{"type":"string"},"provider":{"type":"string"},"region":{"type":"object"},"volumes":{"type":"array","items":{"type":"array"}},"status":{"type":"string","enum":["pending","active","error","deleted","deleting"]},"gpuCount":{"type":"number"},"gpuModel":{"type":"string"},"ram":{"type":"number"},"storage":{"type":"number"},"vramPerGpu":{"type":"number"},"vcpus":{"type":"number"},"interconnect":{"type":"string"},"nvlink":{"type":"boolean"},"os":{"type":"string"},"ip":{"type":"string"},"sshUser":{"type":"string"},"sshPort":{"type":"number"},"dockerImage":{"type":"string"},"dockerCommand":{"type":"string"},"hourlyPrice":{"type":"number"},"deletedAt":{"format":"date-time","type":"string"},"createdAt":{"format":"date-time","type":"string"},"updatedAt":{"format":"date-time","type":"string"}},"required":["_id","name","provider","region","volumes","status","gpuCount","gpuModel","ram","storage","vramPerGpu","vcpus","interconnect","nvlink","os","ip","sshUser","sshPort","dockerImage","dockerCommand","hourlyPrice","deletedAt","createdAt","updatedAt"]}}}}
```

### Get details about a GPU Cloud instance created

This endpoint is useful when you need to check the details and status of a specific instance, for example, for troubleshooting or configuration verification.

{% hint style="info" %}
Use the instance ID from your list to quickly retrieve detailed information.
{% endhint %}

## GET /gpu-cloud/instances/{id}

>

```json
{"openapi":"3.0.0","info":{"title":"Cloud Sesterce API","version":"1.0"},"paths":{"/gpu-cloud/instances/{id}":{"get":{"operationId":"GPUCloudInstancesController_getInstanceById","parameters":[{"name":"x-api-key","in":"header","description":"The API Key secret should be sent through this header to authenticate the request.","required":true,"schema":{"type":"string"}},{"name":"id","required":true,"in":"path","schema":{"type":"string"}}],"responses":{"200":{"description":"Return the instance's details","content":{"application/json":{"schema":{"$ref":"#/components/schemas/InstanceDetailsDto"}}}},"403":{"description":"API key invalid","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}},"404":{"description":"Not found","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}}},"tags":["GPUCloudInstances"]}}},"components":{"schemas":{"InstanceDetailsDto":{"type":"object","properties":{"_id":{"type":"string"},"name":{"type":"string"},"provider":{"type":"string"},"status":{"type":"string","enum":["pending","active","error","deleted","deleting"]},"region":{"type":"object"},"volumes":{"type":"array","items":{"type":"array"}},"sshKey":{"type":"object"},"createdAt":{"format":"date-time","type":"string"},"deletedAt":{"format":"date-time","type":"string"},"isPending":{"type":"boolean"},"hourlyPrice":{"type":"number"},"ip":{"type":"string"},"sshUser":{"type":"string"},"sshPort":{"type":"number"},"portForwards":{"type":"array","items":{"type":"string"}},"gpuCount":{"type":"number"},"gpuModel":{"type":"string"},"ram":{"type":"number"},"storage":{"type":"number"},"vramPerGpu":{"type":"number"},"vcpus":{"type":"number"},"interconnect":{"type":"string"},"nvlink":{"type":"boolean"},"os":{"type":"string"},"dockerImage":{"type":"string"},"dockerCommand":{"type":"string"}},"required":["_id","name","provider","status","region","volumes","sshKey","createdAt","deletedAt","isPending","hourlyPrice","ip","sshUser","sshPort","portForwards","gpuCount","gpuModel","ram","storage","vramPerGpu","vcpus","interconnect","nvlink","os","dockerImage","dockerCommand"]}}}}
```

### Delete a GPU Cloud instance

{% hint style="info" %}
If needed, ensure data backup before deleting instances to prevent data loss. Discover how to create persistent storage [through the following endpoint](/api-reference/volumes#create-a-new-volume).
{% endhint %}

Use this endpoint when you want to free up resources by deleting an instance that is no longer needed, optimizing your resource usage and costs.

## DELETE /gpu-cloud/instances/{id}

>

```json
{"openapi":"3.0.0","info":{"title":"Cloud Sesterce API","version":"1.0"},"paths":{"/gpu-cloud/instances/{id}":{"delete":{"operationId":"GPUCloudInstancesController_deleteInstance","parameters":[{"name":"x-api-key","in":"header","description":"The API Key secret should be sent through this header to authenticate the request.","required":true,"schema":{"type":"string"}},{"name":"id","required":true,"in":"path","schema":{"type":"string"}}],"responses":{"204":{"description":"Instance deleted successfully"},"403":{"description":"API key invalid","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}},"404":{"description":"Not found","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}}},"tags":["GPUCloudInstances"]}}}}
```


# SSH Keys

### Create a SSH Key

Use this endpoint to add new SSH keys, especially when integrating new developers or setting up new instances that require secure access.

{% hint style="info" %}
Use descriptive names for your keys to easily manage multiple SSH credentials.
{% endhint %}

## POST /gpu-cloud/ssh-keys

>

```json
{"openapi":"3.0.0","info":{"title":"Cloud Sesterce API","version":"1.0"},"paths":{"/gpu-cloud/ssh-keys":{"post":{"operationId":"GPUCloudSSHKeysController_create","parameters":[{"name":"x-api-key","in":"header","description":"The API Key secret should be sent through this header to authenticate the request.","required":true,"schema":{"type":"string"}}],"requestBody":{"required":true,"content":{"application/json":{"schema":{"$ref":"#/components/schemas/CreateSSHKeyDto"}}}},"responses":{"201":{"description":"Return the created ssh key","content":{"application/json":{"schema":{"$ref":"#/components/schemas/SSHKeyDto"}}}},"403":{"description":"API key invalid","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}},"404":{"description":"API key not found","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}}},"tags":["GPUCloudSSHKeys"]}}},"components":{"schemas":{"CreateSSHKeyDto":{"type":"object","properties":{"name":{"type":"string"},"publicKey":{"type":"string"}},"required":["name","publicKey"]},"SSHKeyDto":{"type":"object","properties":{"_id":{"type":"string"},"name":{"type":"string"},"publicKey":{"type":"string"},"isDefault":{"type":"boolean"},"createdAt":{"type":"string"},"updatedAt":{"type":"string"}},"required":["_id","name","publicKey","isDefault","createdAt","updatedAt"]}}}}
```

### Retrieve the list of SSH Keys created

Access this endpoint to manage your SSH keys, which is essential for securing access to your instances and maintaining a safe environment.

## GET /gpu-cloud/ssh-keys

>

```json
{"openapi":"3.0.0","info":{"title":"Cloud Sesterce API","version":"1.0"},"paths":{"/gpu-cloud/ssh-keys":{"get":{"operationId":"GPUCloudSSHKeysController_getUserSSHKeys","parameters":[{"name":"x-api-key","in":"header","description":"The API Key secret should be sent through this header to authenticate the request.","required":true,"schema":{"type":"string"}}],"responses":{"200":{"description":"Returns the list of user ssh keys","content":{"application/json":{"schema":{"type":"array","items":{"$ref":"#/components/schemas/SSHKeyDto"}}}}},"403":{"description":"API key invalid","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}},"404":{"description":"API key not found","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}}},"tags":["GPUCloudSSHKeys"]}}},"components":{"schemas":{"SSHKeyDto":{"type":"object","properties":{"_id":{"type":"string"},"name":{"type":"string"},"publicKey":{"type":"string"},"isDefault":{"type":"boolean"},"createdAt":{"type":"string"},"updatedAt":{"type":"string"}},"required":["_id","name","publicKey","isDefault","createdAt","updatedAt"]}}}}
```

### Mark a SSH Key as default

This endpoint is handy when you want to change the default SSH key for your connections, for example, during regular key rotation for security reasons.

## PATCH /gpu-cloud/ssh-keys/{id}/makedefault

>

```json
{"openapi":"3.0.0","info":{"title":"Cloud Sesterce API","version":"1.0"},"paths":{"/gpu-cloud/ssh-keys/{id}/makedefault":{"patch":{"operationId":"GPUCloudSSHKeysController_makeDefault","parameters":[{"name":"x-api-key","in":"header","description":"The API Key secret should be sent through this header to authenticate the request.","required":true,"schema":{"type":"string"}},{"name":"id","required":true,"in":"path","schema":{"type":"string"}}],"responses":{"204":{"description":"SSH key marked as default"},"403":{"description":"API key invalid","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"timestamp":{"type":"string"},"path":{"type":"string"},"message":{"type":"string"}}}}}},"404":{"description":"Not found","content":{"application/json":{"schema":{"type":"object","properties":{"statusCode":{"type":"number"},"message":{"type":"string"}}}}}}},"tags":["GPUCloudSSHKeys"]}}}}
```

### Delete a SSH Key

Use this endpoint to clean up and remove obsolete or compromised SSH keys to maintain the security of your environment.

{% hint style="info" %}
Regularly review and clean up unused SSH keys to maintain security.
{% endhint %}

{% openapi src="/files/yOaAnvwplPHKY2I9hbUw" path="/gpu-cloud/ssh-keys/{id}" method="delete" %}
[V2\_docAPI.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FpzcBuompwnhSyFJWnZ0Q%2FV2_docAPI.json?alt=media\&token=8e494e06-cadb-419f-b509-08f84b8f3c10)
{% endopenapi %}


# Volumes

### Retrieve available Volume offers

{% hint style="danger" %}
To be attached to an instance, volumes should be linked to the same `CloudProvider` and the same `region` as the Pod selected. Get all informations about Volumes[ into this section](/compute-instances#i-need-a-persistent-storage)
{% endhint %}

Access this endpoint to explore available volume offers, which is crucial when planning storage needs for your instances.

{% openapi src="/files/yOaAnvwplPHKY2I9hbUw" path="/gpu-cloud/volumes/offers" method="get" %}
[V2\_docAPI.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FpzcBuompwnhSyFJWnZ0Q%2FV2_docAPI.json?alt=media\&token=8e494e06-cadb-419f-b509-08f84b8f3c10)
{% endopenapi %}

### Create a new Volume

This endpoint is used when you need to create new Volume to add a persistent storage solution to your instance, ensuring to keep your dataset stored even after the instance stops.

{% openapi src="/files/yOaAnvwplPHKY2I9hbUw" path="/gpu-cloud/volumes" method="post" %}
[V2\_docAPI.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FpzcBuompwnhSyFJWnZ0Q%2FV2_docAPI.json?alt=media\&token=8e494e06-cadb-419f-b509-08f84b8f3c10)
{% endopenapi %}

### Get the list of Volumes created

Use this endpoint to get an overview of your active volumes, useful for monitoring storage usage and planning expansions.

{% openapi src="/files/yOaAnvwplPHKY2I9hbUw" path="/gpu-cloud/volumes" method="get" %}
[V2\_docAPI.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FpzcBuompwnhSyFJWnZ0Q%2FV2_docAPI.json?alt=media\&token=8e494e06-cadb-419f-b509-08f84b8f3c10)
{% endopenapi %}

### Get details of a specific Volume

Get specific details about a particular volume when you need to verify its configuration or status.

{% openapi src="/files/yOaAnvwplPHKY2I9hbUw" path="/gpu-cloud/volumes/{id}" method="get" %}
[V2\_docAPI.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FpzcBuompwnhSyFJWnZ0Q%2FV2_docAPI.json?alt=media\&token=8e494e06-cadb-419f-b509-08f84b8f3c10)
{% endopenapi %}

### Delete a Volume

Use this endpoint to delete volumes that are no longer needed, optimizing storage usage and reducing costs.

{% hint style="danger" %}
Ensure data backup before deleting volumes to prevent data loss.
{% endhint %}

{% openapi src="/files/yOaAnvwplPHKY2I9hbUw" path="/gpu-cloud/volumes/{id}" method="delete" %}
[V2\_docAPI.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FpzcBuompwnhSyFJWnZ0Q%2FV2_docAPI.json?alt=media\&token=8e494e06-cadb-419f-b509-08f84b8f3c10)
{% endopenapi %}


# Inference Instances

### Get Inference models

This endpoint allows you to view available AI models for deployment, aiding in selecting the right model for your needs.

{% hint style="info" %}
Check model features to match your specific project requirements.
{% endhint %}

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/models" method="get" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Get inference hardware

You have several hardware options available through the AI inference feature of Sesterce Cloud (you can consult this section for more information). This endpoint allows you to explore options for deploying AI instances, which are crucial for planning resources and manage latency rate.

{% hint style="info" %}
Evaluate hardware capabilities to ensure optimal performance for your AI tasks.
{% endhint %}

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/hardwares" method="get" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Get Regions available for inference instances

Identify available regions for deploying AI instances, important for compliance and latency considerations.

{% hint style="info" %}
The region choice is a crucial parameter for your inference endpoint hosting. It will determine the latency rate for your final end-users. Choose regions that align with your data residency and latency needs.
{% endhint %}

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/regions" method="get" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Create a Registry

A registry is necessary if you need to infere your own custom model, which is not publicly available. [Click here to learn more about Registries](/ai-inference-instances/inference-instance-configuration#how-to-use-private-custom-model) on Sesterce Cloud AI Inference service!

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/registries" method="post" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Get the list of registries created

To manage your registries for storing and accessing AI models, use the following endpoint:

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/registries" method="get" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Update a registry

To modify registry details to ensure they meet current security and access needs, use the following endpoint:

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/registries/{id}" method="patch" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Delete a Registry

The following endpoint allows to remove outdated or unused registries to maintain a clean environment.

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/registries/{id}" method="delete" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Create an inference instance

Time has come! You can now deploy a new AI inference instance to scale your applications and services, or deploy in production an existing model! Use the following endpoint to perform this action.

{% hint style="warning" %}
To create an inference instance, check that your credit balance is filled. Please [check here ](/welcome-on-sesterce-cloud/payment-and-billing)our documentation to top up your balance.
{% endhint %}

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/instances" method="post" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Start an inference instance

This endpoint allows you to activate an AI inference instance to begin processing tasks and data.

{% hint style="info" %}
You can monitor startup times to assess performance efficiency.
{% endhint %}

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/instances/{id}/start" method="post" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Get the list of your Inference instances

Here is the endpoint to monitor your active AI instances to manage resources and performance.

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/instances" method="get" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Get details about a specific Inference Instance

Retrieve detailed information about a specific AI instance for management and troubleshooting.

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/instances/{id}" method="get" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Preview AI instance pricing

This endpoints allows you to estimate costs for your running AI instances, helping in budget planning.

{% hint style="info" %}
Sesterce Cloud AI inference service is based on an unlimited-token pricing. This means you are charged for a global hour price, whatever the use of your dedicated endpoint.
{% endhint %}

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/instances/pricing" method="post" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Update an inference instance

This endpoint allows you to modify existing AI instances to adapt to changing project needs. This is particularly useful is you need to update your hardware flavor and/or autoscaling limits according to the use of your dedicated endpoint :

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/instances/{id}" method="patch" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}

### Stop an inference instance

If you need to pause an AI instance to conserve resources and manage costs, use the following endpoint:

{% openapi src="/files/xUqTfYLJzwhLR3swwK50" path="/ai-inference/instances/{id}/stop" method="post" %}
[user-api-docs.json](https://3376774032-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FoExYATuECEyEKGJ4cD8X%2Fuploads%2FPJS3wkW5AlGxKDItJyuE%2Fuser-api-docs.json?alt=media\&token=7cdfae61-f77a-40df-8a75-197d1ad217a1)
{% endopenapi %}


# Which compute instance for AI models training and inference?

Choosing the right GPU instance is crucial to the success of your AI projects. An optimal configuration not only improves performance, but also keeps costs under control. This guide will help you navigate through the various options available to find the ideal solution for your needs.

## Why is the compute instance choice so important?

Your choice of GPU infrastructure has a direct impact on..:

* **Performance**: model training speed and inference latency
* **Cost**: optimizing your budget by avoiding over-sizing
* **Scalability**: ability to scale according to your needs
* **Reliability**: stability of your workloads in production

## Which compute instance to choose for model training?

### 1. Large Language Models (LLMs) models training

Fine-tuning Large Language Models represents one of the most resource-intensive tasks in modern AI development.&#x20;

The hardware requirements vary significantly based on model size, from smaller 7B parameter models to massive 70B+ architectures. This section will help you select the optimal configuration for your fine-tuning project, ensuring efficient resource utilization while maintaining performance.

<table><thead><tr><th width="147.28125">Model Size</th><th width="125.578125">Server type</th><th width="181.26171875">VRAM</th><th width="247.87890625">Recommended Offers</th></tr></thead><tbody><tr><td>Small (less than 7B parameters)</td><td>VM</td><td>From <strong>22 GB</strong> (for 1B parameters models) to <strong>140 GB</strong> (to fine-tune models such as DeepSeek-R1 7B)</td><td><p>1 to 3B models:<br><span data-gb-custom-inline data-tag="emoji" data-code="1f449">👉</span> The most cost-effective: <a href="https://cloud.sesterce.com/compute/new?gpuType=A100_80G&#x26;numGpus=1"><strong>1xA100 80G</strong></a></p><p><span data-gb-custom-inline data-tag="emoji" data-code="1f449">👉</span> The most efficient: </p><p><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=1"><strong>1xH100 80G</strong></a><br>7B models:<br><span data-gb-custom-inline data-tag="emoji" data-code="1f449">👉</span> <a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=1"><strong>1xH200 141 GB</strong></a></p></td></tr><tr><td>Medium (12B-32B)</td><td>VM or Bare-Metal</td><td>From <strong>200</strong> to <strong>500 GB</strong> (to fine-tune models such as DeepSeek-R1 32B)</td><td><p>12 to 14B models:<br><span data-gb-custom-inline data-tag="emoji" data-code="1f449">👉</span> <a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=2"><strong>2xH200 141 GB</strong></a><br>27 to 32B models:</p><p><span data-gb-custom-inline data-tag="emoji" data-code="1f449">👉</span><a href="https://cloud.sesterce.com/compute/new?gpuType=A100_80G&#x26;numGpus=8"> <strong>8xA100 80G Bare Metal</strong></a><br>or <a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=4"><strong>4xH200 141 GB</strong></a></p></td></tr><tr><td>Large (70B and more)</td><td>Bare-Metal</td><td>More than <strong>1000GB</strong></td><td><span data-gb-custom-inline data-tag="emoji" data-code="1f449">👉</span> <a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8xH200 141 GB Bare-Metal</strong></a><br><span data-gb-custom-inline data-tag="emoji" data-code="1f449">👉</span> <a href="https://cloud.sesterce.com/compute/new?gpuType=B200&#x26;numGpus=8"><strong>8xB200 192 GB Bare-Metal</strong></a></td></tr></tbody></table>

### 2. Computer Vision models training

Whether you're developing object detection systems, processing medical imagery, or creating next-generation AI art, selecting the right GPU infrastructure is crucial for your success.

The computational requirements for vision tasks vary significantly based on complexity and scale. Classification tasks might require modest GPU power, while advanced generative models demand substantial computational resources. This section outlines three primary categories of computer vision workloads - classification, segmentation, and generation - each with its unique hardware requirements and optimal configurations.

<table><thead><tr><th width="253.15625">Use Case</th><th width="148.96875">Models example</th><th width="156.046875">Resource intensity</th><th>Batch</th></tr></thead><tbody><tr><td><a href="#classification-models"><strong>Classification</strong></a> (images, videos): object detection, face recognition</td><td>ResNet, YOLO</td><td>⭐️⭐️⭐️</td><td>32-64</td></tr><tr><td><a href="#segmentation-models"><strong>Segmentation</strong></a> (pixel-level): medical imaging, satellite analysis</td><td>U-Net, DeepLab</td><td>⭐️⭐️⭐️⭐️</td><td>16-32</td></tr><tr><td><a href="#generative-models"><strong>Generative models</strong></a> (stable diffusion): image generation, AI art</td><td>Stable Diffusion</td><td>⭐️⭐️⭐️⭐️⭐️</td><td>8-16</td></tr></tbody></table>

#### Classification Models

<table><thead><tr><th width="114.4140625">Data Volume</th><th width="112.734375">Training Time</th><th width="176.16015625">Recommended Instance</th><th width="88.5">Type</th><th>Comment</th></tr></thead><tbody><tr><td>&#x3C;100GB</td><td>&#x3C;24h</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=L40S&#x26;numGpus=1"><strong>1xL40S 48 GB</strong></a></td><td>VM</td><td>Perfect for development and testing</td></tr><tr><td>100-500GB</td><td>1-3 days</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=1"><strong>1xH100 80GB</strong></a></td><td>VM</td><td>Parallel processing beneficial</td></tr><tr><td>500GB-1TB</td><td>3-7 days</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=L40S&#x26;numGpus=2"><strong>2xL40S 48 GB</strong></a></td><td>VM/BM</td><td>Higher throughput needed</td></tr><tr><td>>1TB</td><td>>1 week</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=8"><strong>8xH100 80 GB</strong></a></td><td>BM</td><td>Bare Metal for optimal performance</td></tr></tbody></table>

#### Segmentation Models

<table><thead><tr><th width="112.78515625">Data Volume</th><th width="111.54296875">Training Time</th><th width="180.3828125">Recommended Instance</th><th width="87.15234375">Type</th><th>Comment</th></tr></thead><tbody><tr><td>&#x3C; 200GB</td><td>&#x3C; 48h</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=2"><strong>2x H200 141 GB</strong></a></td><td>VM</td><td>High-res image processing</td></tr><tr><td>200GB-1TB</td><td>3-5 days</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=A100_80G&#x26;numGpus=4"><strong>4x A100 80GB</strong></a></td><td>VM</td><td>Multiple batch processing</td></tr><tr><td>1TB-5TB</td><td>1-2 weeks</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=8"><strong>8x H100 80 GB</strong></a></td><td>BM</td><td>Heavy data augmentation</td></tr><tr><td>> 5TB</td><td>> 2 weeks</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8x H200 80 GB</strong></a></td><td>BM</td><td>Maximum processing power</td></tr></tbody></table>

#### Generative Models

<table><thead><tr><th width="113.91796875">Data Volume</th><th width="110.5625">Training Time</th><th width="170.08203125">Recommended Instance</th><th width="101.39453125">Type</th><th>Notes</th></tr></thead><tbody><tr><td>&#x3C; 500GB</td><td>&#x3C; 3 days</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=4"><strong>4x H100 80 GB</strong></a></td><td>VM</td><td>Model fine-tuning</td></tr><tr><td>500GB-2TB</td><td>3-7 days</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=8"><strong>8x H100 80 GB</strong></a></td><td>BM</td><td>Full model training</td></tr><tr><td>2TB-10TB</td><td>1-3 weeks</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8x H200 141 GB</strong></a></td><td>BM</td><td>Large scale training</td></tr><tr><td>> 10TB</td><td>> 3 weeks</td><td><a href="https://www.sesterce.com/booking"><strong>16x H100 80GB</strong></a></td><td>BM</td><td>Distributed training</td></tr></tbody></table>

### 3. Audio/Speech models training

Whether you're developing voice recognition systems, building text-to-speech applications, or exploring the cutting edge of AI music generation, selecting the right GPU infrastructure is crucial for successful model training.

This section outlines three primary categories of audio ML workloads - speech recognition, text-to-speech synthesis, and audio generation - each requiring specific hardware configurations to achieve optimal performance.

| Use Case           | Input Data Type                               | Dataset Size | Model Examples                  | Resource Intensity |
| ------------------ | --------------------------------------------- | ------------ | ------------------------------- | ------------------ |
| Speech Recognition | Audio files (.wav, .mp3), Labeled transcripts | 100GB-1TB    | Whisper, DeepSpeech, Wav2Vec    | ⭐⭐⭐                |
| Text-to-Speech     | Text corpus, Audio pairs                      | 50-500GB     | Tacotron, FastSpeech, VALL-E    | ⭐⭐⭐⭐               |
| Audio Generation   | Audio samples, MIDI files                     | 1-2TB        | MusicLM, AudioLDM, Stable Audio | ⭐⭐⭐⭐⭐              |

#### Speech Recognition Models

<table><thead><tr><th>Model Size (param)</th><th width="138.01953125">VRAM/RAM Needed</th><th width="164.63671875">Recommended Instance</th><th width="101.12890625">Type</th><th width="156.2421875">Comment</th></tr></thead><tbody><tr><td>Small (&#x3C; 100M params)</td><td>24GB/64GB</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=A6000&#x26;numGpus=1"><strong>1x A6000 48GB</strong></a></td><td>VM</td><td>Development/testing</td></tr><tr><td>Medium (100M-500M)</td><td>48GB/128GB</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=L40S&#x26;numGpus=2"><strong>2x L40S 48GB</strong></a><br><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=1"><strong>1x H100 80GB</strong></a></td><td>VM</td><td>Production training</td></tr><tr><td>Large (500M-1B)</td><td>96GB/256GB</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=L40S&#x26;numGpus=4"><strong>4x L40S 48GB</strong></a></td><td>VM/BM</td><td>Large scale training</td></tr><tr><td>Very Large (>1B)</td><td>160GB/384GB</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=8"><strong>8xH100 80GB</strong></a></td><td>BM</td><td>Enterprise scale</td></tr></tbody></table>

#### Text-to-Speech Models

<table><thead><tr><th width="143.1796875">Model Size</th><th width="143.52734375">VRAM/RAM Needed</th><th width="160.80078125">Recommended Instance</th><th width="102.390625">Type</th><th>Notes</th></tr></thead><tbody><tr><td>Small (&#x3C; 200M params)</td><td>80GB/192GB</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=2"><strong>2x H100 80GB</strong></a></td><td>VM</td><td>Basic TTS</td></tr><tr><td>Medium (200M-500M)</td><td>160GB/384GB</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=4"><strong>4xH100 80GB</strong></a></td><td>VM</td><td>Multi-speaker</td></tr><tr><td>Large (500M-1B)</td><td>320GB/768GB</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=8"><strong>8x H100 80GB</strong></a></td><td>BM</td><td>High-quality TTS</td></tr><tr><td>Very Large (>1B)</td><td>640GB/1TB</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8x H200 141GB</strong></a></td><td>BM</td><td>Enterprise TTS</td></tr></tbody></table>

## Which compute instance to choose for model inference?

### 1. Large Language Models (LLMs) inference

Whether you're serving chatbots, content generation, or text analysis applications, choosing the right infrastructure is crucial for balancing performance, cost, and user experience.

The requirements for LLM inference vary significantly based on several key factors: model size (from 7B to 70B+ parameters), user load (from individual testing to thousands of concurrent users), and latency requirements (from real-time chat applications to batch processing). Each of these factors directly impacts your choice of infrastructure, from single GPU instances to distributed multi-GPU deployments.

#### LLM Inference Sizing - Small Scale (1-50 concurrent users)

<table><thead><tr><th width="113.4453125">Model Size</th><th width="143.8046875">Concurrent Users</th><th width="155.19140625">VRAM/RAM needed</th><th>Latency Target</th><th>Recommended Instance</th><th>Type</th><th>Estimated RPS*</th><th>Comment</th></tr></thead><tbody><tr><td>7B</td><td>1-10</td><td>16GB/32GB</td><td>&#x3C; 100ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=RTX4090&#x26;numGpus=1"><strong>1x RTX4090</strong></a></td><td>VM</td><td>15-20</td><td>Development/testing</td></tr><tr><td>7B</td><td>11-25</td><td>24GB/64GB</td><td>&#x3C; 100ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=L40S&#x26;numGpus=1"><strong>1x L40S</strong></a></td><td>VM</td><td>30-40</td><td>Small production</td></tr><tr><td>7B</td><td>26-50</td><td>48GB/128GB</td><td>&#x3C; 100ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=1"><strong>1x H100</strong></a></td><td>VM</td><td>60-80</td><td>Medium production</td></tr><tr><td>13B</td><td>1-10</td><td>24GB/64GB</td><td>&#x3C; 150ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=RTX4090&#x26;numGpus=2"><strong>1x RTX4090</strong></a></td><td>VM</td><td>10-15</td><td>Development/testing</td></tr><tr><td>13B</td><td>11-25</td><td>48GB/128GB</td><td>&#x3C; 150ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=1"><strong>1x H100</strong></a></td><td>VM</td><td>25-35</td><td>Small production</td></tr><tr><td>13B</td><td>26-50</td><td>80GB/192GB</td><td>&#x3C; 150ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=2"><strong>2x H100</strong></a></td><td>VM/BM</td><td>50-70</td><td>Medium production</td></tr><tr><td>70B</td><td>1-10</td><td>80GB/192GB</td><td>&#x3C; 200ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=2"><strong>2x H100</strong></a></td><td>VM/BM</td><td>5-8</td><td>Small production</td></tr><tr><td>70B</td><td>11-25</td><td>160GB/384GB</td><td>&#x3C; 200ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=4"><strong>4x H100</strong></a></td><td>BM</td><td>15-20</td><td>Medium production</td></tr><tr><td>70B</td><td>26-50</td><td>320GB/768GB</td><td>&#x3C; 200ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=4"><strong>8x H100</strong></a></td><td>BM</td><td>35-45</td><td>Large production</td></tr></tbody></table>

#### LLM inference Sizing - Medium Scale (51-200 concurrent users)

<table><thead><tr><th width="114.2421875">Model Size</th><th width="141.48828125">Concurrent Users</th><th width="160.5546875">VRAM/RAM</th><th>Latency Target</th><th>Recommended Instance</th><th>Type</th><th>Estimated RPS*</th><th>Notes</th></tr></thead><tbody><tr><td>7B</td><td>51-100</td><td>48GB/128GB</td><td>&#x3C; 100ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=L40S&#x26;numGpus=2"><strong>2x L40S</strong></a></td><td>VM</td><td>100-120</td><td>Production</td></tr><tr><td>7B</td><td>101-150</td><td>80GB/192GB</td><td>&#x3C; 100ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=2"><strong>2x H100</strong></a></td><td>VM/BM</td><td>150-180</td><td>High-performance</td></tr><tr><td>7B</td><td>151-200</td><td>160GB/384GB</td><td>&#x3C; 100ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=4"><strong>4x H100</strong></a></td><td>BM</td><td>200-240</td><td>Enterprise scale</td></tr><tr><td>13B</td><td>51-100</td><td>160GB/384GB</td><td>&#x3C; 150ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=4"><strong>4x H100</strong></a></td><td>BM</td><td>80-100</td><td>Production</td></tr><tr><td>13B</td><td>101-150</td><td>240GB/512GB</td><td>&#x3C; 150ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=4"><strong>4x H200</strong></a></td><td>BM</td><td>120-150</td><td>High-performance</td></tr><tr><td>13B</td><td>151-200</td><td>320GB/768GB</td><td>&#x3C; 150ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=8"><strong>8x H100</strong></a></td><td>BM</td><td>160-200</td><td>Enterprise scale</td></tr><tr><td>70B</td><td>51-100</td><td>480GB/1TB</td><td>&#x3C; 200ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8x H200</strong></a></td><td>BM</td><td>60-80</td><td>Production</td></tr><tr><td>70B</td><td>101-150</td><td>640GB/1.5TB</td><td>&#x3C; 200ms</td><td><a href="https://www.sesterce.com/booking"><strong>16x H100</strong></a></td><td>BM</td><td>90-120</td><td>High-performance</td></tr><tr><td>70B</td><td>151-200</td><td>800GB/2TB</td><td>&#x3C; 200ms</td><td><a href="https://www.sesterce.com/booking"><strong>16x H200</strong></a></td><td>BM</td><td>140-180</td><td>Enterprise scale</td></tr></tbody></table>

#### LLM Inference Sizing - Large Scale (201-1000+ concurrent users)

<table><thead><tr><th width="111.828125">Model Size</th><th>Concurrent Users</th><th width="173.234375">VRAM/RAM</th><th width="135.875">Latency Target</th><th>Recommended Instance</th><th>Type</th><th>Estimated RPS*</th><th>Notes</th></tr></thead><tbody><tr><td>7B</td><td>201-500</td><td>320GB/768GB</td><td>&#x3C; 100ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=8"><strong>8x H100</strong></a></td><td>BM</td><td>300-400</td><td>Enterprise scale</td></tr><tr><td>7B</td><td>501-1000</td><td>640GB/1.5TB</td><td>&#x3C; 100ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8x H200</strong></a></td><td>BM</td><td>600-800</td><td>High-scale production</td></tr><tr><td>7B</td><td>1000+</td><td>1.2TB/2.5TB</td><td>&#x3C; 100ms</td><td><a href="https://www.sesterce.com/booking"><strong>16xH200</strong></a></td><td>BM</td><td>1000+</td><td>Distributed clusters</td></tr><tr><td>13B</td><td>201-500</td><td>480GB/1TB</td><td>&#x3C; 150ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8x H200</strong></a></td><td>BM</td><td>250-350</td><td>Enterprise scale</td></tr><tr><td>13B</td><td>501-1000</td><td>800GB/2TB</td><td>&#x3C; 150ms</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8x H200</strong></a></td><td>BM</td><td>500-700</td><td>High-scale production</td></tr><tr><td>13B</td><td>1000+</td><td>1.6TB/3TB</td><td>&#x3C; 150ms</td><td><a href="https://www.sesterce.com/booking"><strong>16x H200</strong></a></td><td>BM</td><td>800+</td><td>Distributed clusters</td></tr><tr><td>70B</td><td>201-500</td><td>1.2TB/2.5TB</td><td>&#x3C; 200ms</td><td><a href="https://www.sesterce.com/booking"><strong>16x H100</strong></a></td><td>BM</td><td>200-300</td><td>Enterprise scale</td></tr><tr><td>70B</td><td>501-1000</td><td>2TB/4TB</td><td>&#x3C; 200ms</td><td><a href="https://www.sesterce.com/booking"><strong>24x H200</strong></a></td><td>BM</td><td>400-600</td><td>High-scale production</td></tr><tr><td>70B</td><td>1000+</td><td>3TB/6TB</td><td>&#x3C; 200ms</td><td><a href="https://www.sesterce.com/booking"><strong>32x H200</strong></a></td><td>BM</td><td>700+</td><td>Distributed clusters</td></tr></tbody></table>

### 2. Image Generation Inference Sizing

Deploying image generation models like Stable Diffusion for production introduces unique infrastructure challenges compared to traditional ML workload.

The hardware requirements vary significantly based on three key factors: model complexity (from base models to SDXL with refiners), concurrent user load (affecting batch processing and queue management), and image generation parameters (resolution, steps, and additional features like ControlNet or inpainting). Each of these factors directly impacts your choice of infrastructure and can significantly affect both performance and operational costs.

#### Small scale (1-50 concurrent users)

<table><thead><tr><th width="116.60546875">Model Type</th><th>Concurrent Users</th><th width="162.890625">VRAM/RAM</th><th>Latency Target*</th><th width="161.92578125">Recommended Instance</th><th>Type</th><th>Images/Minute**</th><th>Notes</th></tr></thead><tbody><tr><td>SD XL Base</td><td>1-10</td><td>16GB/32GB</td><td>&#x3C; 3s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=RTX4090&#x26;numGpus=1"><strong>1x RTX4090</strong></a></td><td>VM</td><td>15-20</td><td>Development/testing</td></tr><tr><td>SD XL Base</td><td>11-25</td><td>24GB/64GB</td><td>&#x3C; 3s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=L40S&#x26;numGpus=1"><strong>1x L40S</strong></a></td><td>VM</td><td>30-40</td><td>Small production</td></tr><tr><td>SD XL Base</td><td>26-50</td><td>48GB/128GB</td><td>&#x3C; 3s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=2"><strong>1x H100</strong></a></td><td>VM</td><td>60-80</td><td>Medium production</td></tr><tr><td>SD XL + Refiner</td><td>1-10</td><td>24GB/64GB</td><td>&#x3C; 5s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=L40S&#x26;numGpus=1"><strong>1x L40S</strong></a></td><td>VM</td><td>10-15</td><td>Development/testing</td></tr><tr><td>SD XL + Refiner</td><td>11-25</td><td>48GB/128GB</td><td>&#x3C; 5s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=2"><strong>1x H100</strong></a></td><td>VM</td><td>25-35</td><td>Small production</td></tr><tr><td>SD XL + Refiner</td><td>26-50</td><td>80GB/192GB</td><td>&#x3C; 5s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=2"><strong>2x H100</strong></a></td><td>VM or BM</td><td>50-70</td><td>Medium production</td></tr></tbody></table>

#### Medium scale (51-200 concurrent users)

<table><thead><tr><th>Model Type</th><th>Concurrent Users</th><th width="164.40234375">VRAM/RAM</th><th>Latency Target*</th><th>Recommended Instance</th><th>Type</th><th>Images/Minute**</th><th>Notes</th></tr></thead><tbody><tr><td>SD XL Base</td><td>51-100</td><td>160GB/384GB</td><td>&#x3C; 3s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=4"><strong>4x H100</strong></a></td><td>VM or BM</td><td>120-150</td><td>Production</td></tr><tr><td>SD XL Base</td><td>101-150</td><td>320GB/768GB</td><td>&#x3C; 3s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=8"><strong>8x H100</strong></a></td><td>BM</td><td>200-250</td><td>High-performance</td></tr><tr><td>SD XL Base</td><td>151-200</td><td>480GB/1TB</td><td>&#x3C; 3s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8x H200</strong></a></td><td>BM</td><td>300-350</td><td>Enterprise scale</td></tr><tr><td>SD XL + Refiner</td><td>51-100</td><td>320GB/768GB</td><td>&#x3C; 5s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H100&#x26;numGpus=8"><strong>8x H100</strong></a></td><td>BM</td><td>100-130</td><td>Production</td></tr><tr><td>SD XL + Refiner</td><td>101-150</td><td>480GB/1TB</td><td>&#x3C; 5s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8x H200</strong></a></td><td>BM</td><td>180-220</td><td>High-performance</td></tr><tr><td>SD XL + Refiner</td><td>151-200</td><td>640GB/1.5TB</td><td>&#x3C; 5s</td><td><a href="https://cloud.sesterce.com/compute/new?gpuType=H200&#x26;numGpus=8"><strong>8x H200</strong></a></td><td>BM</td><td>250-300</td><td>Enterprise scale</td></tr></tbody></table>

#### Large scale (201-1000+ concurrent users)

| Model Type      | Concurrent Users | VRAM/RAM  | Latency Target\* | Recommended Instance                             | Type | Images/Minute\*\* | Notes               |
| --------------- | ---------------- | --------- | ---------------- | ------------------------------------------------ | ---- | ----------------- | ------------------- |
| SD XL Base      | 201-500          | 800GB/2TB | < 3s             | [**16x H100**](https://www.sesterce.com/booking) | BM   | 400-500           | Multi-cluster       |
| SD XL Base      | 501-1000         | 1.6TB/4TB | < 3s             | [**16x H200**](https://www.sesterce.com/booking) | BM   | 800-1000          | Distributed system  |
| SD XL Base      | 1000+            | 2.4TB/6TB | < 3s             | [**24x H200**](https://www.sesterce.com/booking) | BM   | 1500+             | Global distribution |
| SD XL + Refiner | 201-500          | 1.2TB/3TB | < 5s             | [**24x H100**](https://www.sesterce.com/booking) | BM   | 350-450           | Multi-cluster       |
| SD XL + Refiner | 501-1000         | 2TB/5TB   | < 5s             | [**32x H200**](https://www.sesterce.com/booking) | BM   | 700-900           | Distributed system  |
| SD XL + Refiner | 1000+            | 3TB/8TB   | < 5s             | [**40x H200**](https://www.sesterce.com/booking) | BM   | 1200+             | Global distribution |


# Expose AI model from Hugging Face using vLLM

You want to expose to the world an API endpoint to allow your end-users to access an AI Model from Hugging Face? This tutorial is for you!

## Choose your VM offer

The first step is to choose the VM offer tailored to your needs on [**Sesterce Cloud**](https://cloud.sesterce.com/compute). The choice of the instance will depend on several factors, such as the size of your model and the number of end-users.

<figure><img src="/files/VMmUl4lWDRG7drh9g3jD" alt=""><figcaption></figcaption></figure>

You can consider following examples according to your use-case:

**Small Scale (1-10 simultaneous users)**

* 7B params: L4 (24GB VRAM)
* 13B params: L40S (48GB VRAM)
* Example: Mistral 7B with 8 users in 4K context = 16GB + (8 × 2GB) = 32GB VRAM

**Medium Scale (10-50 users)**

* 7B params : 2-4x L40S in parallel
* 13B params: 4x L40S or H100
* Recommended configuration: Load balancing between multiple GPUs

**Large Scale (50+ users)**

* 7B params: 8x L40S or 2x H100
* 13B+ params: H100 multi-GPU
* Use quantization techniques (INT8/INT4)

## Set-up your instance

Well done, you are now able to configure your VM! Please find here [**the steps required to do it**](/compute-instances/configure-your-compute-instance). In "Images" section, choose vLLM option. The instance launch usually takes around 5 minutes.

<figure><img src="/files/eLNEQiYjkLPaQSj6ghle" alt=""><figcaption></figcaption></figure>

## Connect to your instance

When your VM instance is launched, you'll be able to get ssh command to connect into it, like `ssh sesterce@<IP_MACHINE>`.&#x20;

{% hint style="info" %}
**Pay attention:** the docker pull is running! Wait until it's finished to type the following command.
{% endhint %}

When the docker pull running is finished, use the following command:

```
docker run -d -e HF_TOKEN=<HUGGING_FACE_TOKEN> --runtime nvidia --gpus all --net=host --ipc=host vllm/vllm-openai --model <MODEL_ID>
```

You can now fill your [Hugging Face Token](https://huggingface.co/settings/tokens) and Model ID.

<figure><img src="/files/g9rPpFtk3HrUk49f5mM6" alt=""><figcaption><p>Create token from Hugging Face</p></figcaption></figure>

<figure><img src="/files/Kq8vlo1DeeSwWahumuB8" alt=""><figcaption><p>Get your Model ID from Hugging Face</p></figcaption></figure>

## Use your model

Well done! Once ce container is running you'll be able to use the model by typing the following command (container run on port 8000). Make sure you replace the variable in the example with your Model ID.

```
curl -X POST http://<IP_MACHINE>:8000/v1/completions -H "Content-Type: application/json" -d '{"model": "<MODEL_ID>","prompt": "Hello world","max_tokens": 50,"temperature": 0}'
```

Well done! the model is accessible from the following endpoint: [http://\<IP\_MACHINE>:8000/v1/models](http://38.128.232.27:8000/v1/models) :rocket:&#x20;

<figure><img src="/files/WHEwAJAhc9qjypKGFcDA" alt=""><figcaption></figcaption></figure>

## Common errors

### Instance RAM insufficient

Make sure you choose an instance with sufficient RAM to run your model.

### Container not running yet

Wait before container is running to access the model. You can use following command to get container status:

```
// docker ps
```

* If the list is empty, it means the launching failed
* Otherwise, you should see the Container name, you can use it through the following command:

```
// docker logs <CONTAINER_NAME>
```

### Port used or unavailable

You can check port status with following command:

```
// ss -tuln | grep <PORT>
```


