AI training and inference
Rent GPUs by the day while you experiment, move to monthly inference once the model settles, and keep data and checkpoints in object storage rather than on one machine.
The challenge
- 01Experimentation is bursty; monthly reservations sit idle.
- 02Datasets and checkpoints are large and hard to share or roll back on local disks.
- 03Production inference needs predictable cost and capacity.
Reference architecture
How it rolls out
- 01
Experiment on daily GPUs
Fine-tune and evaluate on daily RTX 4090 or L40S plans, with datasets and checkpoints in object storage.
- 02
Settle the evaluation
Fix your evaluation scripts and baseline data so every run is compared automatically and the release candidate is clear.
- 03
Move inference to monthly
Once the model settles, serve it on monthly GPUs or the inference service for predictable cost and capacity.
- 04
Scale with traffic
Put inference behind a load balancer and add daily GPUs at peaks.
How a customer uses it
A start-up building AI for legal documents
- Before
- They rented four cards monthly that sat idle outside experiment weeks, and kept checkpoints on local disks, re-uploading hundreds of gigabytes whenever they changed machines.
- After
- Daily cards for experiments, monthly for production, checkpoints in one bucket. With the same training rhythm, GPU spend fell sharply, and switching machines means attaching the same bucket.
−45%
Monthly GPU spend
Minutes
To resume on a new machine
2 weeks
From experiment to production
Design notes
Mix daily and monthly
NVIDIA RTX 4090, RTX 5090 and L40S offer daily plans for short training runs; A100 and H100 are monthly for sustained inference and large models.
Use the Ningxia GPU region
Ningxia is a GPU region; keep object storage in the same region to cut data-transfer cost. Users in the US or Europe can pick Texas or Munich.
Rather not manage GPUs?
Call models through the AI inference service API without deploying anything.