Less idle time.
More useful compute.

Run your container. Resize its resources. Let idle compute sleep.
Pay for allocated GPU, CPU and RAM, plus the workspace you keep.

Two meters.
One clear bill.

Compute stops billing after resources are released. Your persistent workspace stays available and billable while you keep it.

COMPUTE

GPU + CPU + RAM

Resource rate × allocated time

Metered by the second, including allocated startup and idle time.
+
WORKSPACE

Persistent storage

GB rate × size × retained time

Continues during sleep. Ends after storage release is confirmed.
EXPLORE A 30-DAY MONTH

The idle tail includes time waiting for the sleep threshold. A request wakes a sleeping container; model loading and available capacity affect cold starts.

COMPUTE ALLOCATED5 / 24h
Work Startup / idle Sleeping
WORKSPACE RETAINED24 / 24h

Illustrative daily pattern. Persistent connections and protected tasks can prevent automatic sleep. Sleep releases compute; it does not erase your files.

Read the billing details

Choose your
compute.

GPU rates are per device or MIG partition. CPU, system RAM and persistent storage are separate. A published rate is not a capacity reservation.

FETCHING CURRENT CATALOG

Scroll sideways to see every rate and status →

Current GPU prices for the selected placement
GPU modelMemoryAllocationUSD / hourCatalog status
PRO 6000 Blackwell24 GBMIG partition—Loading
PRO 6000 Blackwell48 GBMIG partition—Loading
PRO 6000 Blackwell96 GBFull GPU—Loading
H200141 GBFull GPU—Loading
B200180 GBFull GPU—Loading
B300288 GBFull GPU—Loading
vCPU— / vCPU-hour
RAMSystem memory— / GB-hour
Workspace storage— / GB-hour

“Not quoted” means no current price, never free. Availability depends on deployment, placement, balance and capacity. The console confirms your selected configuration before creation. CPU/RAM may need adjustment for your workload.

YOUR ILLUSTRATIVE CONFIGURATION

See the whole estimate.

1 GPU, 4 vCPU, 16 GB RAM, 20 GB workspace.
5 allocated hours/day · 30 days · storage retained throughout.

ESTIMATED MONTHLY TOTAL— / 30 days

Select a currently quoted profile in an enabled placement to calculate. Unavailable profiles cannot be estimated.

Review in the console

Match the spend
to the work.

These compare workload strategies, not competitor invoices. Time, utilization and performance inputs are illustrative. Monthly estimates use the live configuration above; no guaranteed savings or startup benchmark is implied.

CHOOSE A SCENARIO
01 / 05
SCENARIO 01

A production API with quiet hours.

A service can be reachable all day without keeping its GPU allocated all day—when the workload can tolerate waking from sleep.

A / ALWAYS ALLOCATED
—/ 30 days
B / OIY / IDLE SLEEP
—/ 30 days
Allocated compute720h / month150h / month
Request latencyWarm applicationCold-start delay after sleep
Useful work / allocated time17%80%
UsageManually retain capacityWake on authenticated request

Uses the work and startup/idle inputs above, with the same GPU, CPU, RAM and storage. Sleep follows the idle threshold, not the instant the last request ends. For strict latency targets or steady traffic, keep compute warm.

Your workload decides the right setup.

Benchmark a representative job. Include startup, idle thresholds, storage and your latency target. Then choose the smallest allocation that gets the work done.

Configure a container