Skip to content

AI training and inference

Rent GPUs by the day while you experiment, move to monthly inference once the model settles, and keep data and checkpoints in object storage rather than on one machine.

The challenge

  • 01Experimentation is bursty; monthly reservations sit idle.
  • 02Datasets and checkpoints are large and hard to share or roll back on local disks.
  • 03Production inference needs predictable cost and capacity.

Reference architecture

How it rolls out

  1. 01

    Experiment on daily GPUs

    Fine-tune and evaluate on daily RTX 4090 or L40S plans, with datasets and checkpoints in object storage.

  2. 02

    Settle the evaluation

    Fix your evaluation scripts and baseline data so every run is compared automatically and the release candidate is clear.

  3. 03

    Move inference to monthly

    Once the model settles, serve it on monthly GPUs or the inference service for predictable cost and capacity.

  4. 04

    Scale with traffic

    Put inference behind a load balancer and add daily GPUs at peaks.

How a customer uses it

A start-up building AI for legal documents

Before
They rented four cards monthly that sat idle outside experiment weeks, and kept checkpoints on local disks, re-uploading hundreds of gigabytes whenever they changed machines.
After
Daily cards for experiments, monthly for production, checkpoints in one bucket. With the same training rhythm, GPU spend fell sharply, and switching machines means attaching the same bucket.
  • −45%

    Monthly GPU spend

  • Minutes

    To resume on a new machine

  • 2 weeks

    From experiment to production

Design notes

Mix daily and monthly

NVIDIA RTX 4090, RTX 5090 and L40S offer daily plans for short training runs; A100 and H100 are monthly for sustained inference and large models.

Use the Ningxia GPU region

Ningxia is a GPU region; keep object storage in the same region to cut data-transfer cost. Users in the US or Europe can pick Texas or Munich.

Rather not manage GPUs?

Call models through the AI inference service API without deploying anything.