Web & Application Development · Hosting
Top

A Cloud GPU server is an ordinary Linux server — your own vCPU, memory and NVMe storage, with full root access — that also has an NVIDIA card attached to it. That is the whole idea, and it is worth being plain about, because "GPU server" is often taken to mean a machine you submit jobs to rather than one you log into and run.

The cards are NVIDIA A16 and A40. The smaller plans give you a measured share of one — a slice of the silicon with its own dedicated VRAM — so you are not paying for a whole card to run a model that needs eight gigabytes. The larger plans scale up to a full card. The VRAM figure in the table is usually the number that decides it: a model that does not fit will not run at all, however many cores sit beside it.

Everything here is billed by the hour and invoiced afterwards, for the hours the server actually exists. A training run that takes an afternoon costs an afternoon. Nothing is charged when you place the order — you put a card on file, we meter the hours, and the invoice follows.

Pick a GPU, Pay for the Hours You Use

Deploys to 3 locations — choose from every location this plan offers at checkout; some cost more

Select Plan vCPU Memory Storage Bandwidth Location
2 8 GB 50 GB NVMe 1 TB Newark, US $0.12/hour
1 5 GB 90 GB NVMe 3 TB Newark, US $0.15/hour
2 16 GB 80 GB NVMe 2 TB Newark, US $0.22/hour
2 10 GB 180 GB NVMe 4 TB Newark, US $0.27/hour
3 32 GB 170 GB NVMe 3 TB Newark, US $0.44/hour
4 20 GB 360 GB NVMe 5 TB Newark, US $0.54/hour
12 128 GB 700 GB NVMe 10 TB Bangalore, IN $1.76/hour
24 120 GB 1400 GB 15 TB London, UK $3.20/hour

What a Cloud GPU Server Gives You

🎮

A Real Server, With a GPU In It

This is not a GPU appliance you submit jobs to. It is a complete Linux server — your own vCPU, memory, NVMe storage and bandwidth, with full root access — that happens to have an NVIDIA card attached. You install what you like and run it how you like.

🧠

NVIDIA A16 and A40 Cards

The A16 suits many smaller concurrent workloads and virtual desktops. The A40 is the heavier card, for training, rendering and larger inference. Both are listed with their real VRAM so you can match the card to the model you intend to load.

🔪

A Share of a Card, or All of It

The entry plans give you a measured slice of a GPU — one eighth of a card, with its own dedicated VRAM — so you are not paying for silicon you will not use. Larger plans scale up to whole cards. The plan table shows exactly which you are buying.

⏱️

Billed by the Hour, in Arrears

You are charged for the hours the server actually exists, invoiced afterwards rather than as a month up front. A job that runs for six hours costs six hours. Nothing is charged when you place the order.

Running in Minutes

Provisioning is automated. The server is built, the GPU drivers' hardware is present and the machine is reachable within minutes of authorising your order — not queued behind a human.

🌍

More Than One Location

GPU capacity is genuinely scarce and not every card is in every datacenter. Each plan lists the locations it can actually deploy to, and you choose from that list at checkout rather than discovering the limit afterwards.

Before You Order

Hours Are Counted, Not Estimated

A daily meter records how long each server has existed and the total appears as a line on your invoice. You can see the running count in your account rather than waiting to be surprised. Because it is measured in arrears, a server you destroy today stops adding to the bill today.

A Stopped Server Is Still a Server

Powering the machine off does not stop the charges — the GPU stays reserved for you, and our own supplier keeps billing for it. Only destroying the server ends the meter. If you are finished with a workload, destroy it rather than halting it, and take a snapshot first if you want the disk back later.

Choosing Between the Cards

If you are serving many light sessions, the A16 slices are usually the better value. If you are training, rendering, or loading a model that needs the VRAM in one piece, take the A40. The VRAM figure in the table is the number that decides it — a model that does not fit will not run faster on a cheaper card.

FAQs

Have A Question?

If you can't find the answer you are looking for our support is just an email away.

Contact Us
Is the GPU shared with other customers?

The VRAM and the slice of the card allocated to your plan are yours for as long as the server exists — they are not oversubscribed. On the smaller plans you are using a partition of a physical card rather than the whole card, which is why they cost a fraction of the price. The plan table states exactly what each one includes.

Can I add a GPU to a server I already have?

No, and it is worth being clear about why. A GPU is part of the machine a plan deploys, not an accessory that can be attached afterwards, and our provider offers no operation to add one to a running server. Moving an existing VPS onto a GPU plan is a new deployment and a migration rather than an upgrade.

If that is what you need, take a snapshot of your current server, deploy the GPU plan, and move your data across.

What does an hour actually cost me?

The rate shown on each plan is per hour, and you are charged for whole hours the server exists. A GPU server left running for a full month costs roughly 730 times the hourly rate. If you only need it for a training run or a render, you only pay for that.

The invoice arrives after the period, listing the hours recorded.

Do I pay anything when I place the order?

No. Because these plans are billed for the hours you use, there is nothing to collect up front. You put a card on file to authorise the order, and it is charged when the invoice for your actual usage is raised.

What happens if I forget and leave it running?

It keeps billing, because it keeps existing and keeps costing us. There is no automatic cut-off. If you are running something expensive, destroy the server when the job finishes rather than powering it off, and take a snapshot first if you want to come back to the same disk.

Which drivers and software are installed?

You choose the operating system at checkout and get a clean server with full root access. Installing CUDA, drivers and your own frameworks is yours to do, which means you get the versions your code actually needs rather than whatever was baked into an image.

Can I get more GPU later?

You can deploy a larger GPU plan at any time, but not resize an existing GPU server into a different card — the machine is built around the hardware it has. Plan a migration rather than an upgrade, and snapshot before you move.