Run your container. Resize its resources. Let idle compute sleep. Pay for allocated GPU, CPU and RAM, plus the workspace you keep.
01 / THE BILLING MODELMETERED RESOURCE USAGE
Two meters. One clear bill.
Compute stops billing after resources are released. Your persistent workspace stays available and billable while you keep it.
COMPUTE
GPU + CPU + RAM
Resource rate × allocated time
Metered by the second, including allocated startup and idle time.
+
WORKSPACE
Persistent storage
GB rate × size × retained time
Continues during sleep. Ends after storage release is confirmed.
EXPLORE A 30-DAY MONTH
The idle tail includes time waiting for the sleep threshold. A request wakes a sleeping container; model loading and available capacity affect cold starts.
COMPUTE ALLOCATED5 / 24h
Work Startup / idle Sleeping
WORKSPACE RETAINED24 / 24h
Illustrative daily pattern. Persistent connections and protected tasks can prevent automatic sleep. Sleep releases compute; it does not erase your files.
02 / CURRENT RESOURCE RATESDIRECT FROM THE OIY CATALOG
Choose your compute.
GPU rates are per device or MIG partition. CPU, system RAM and persistent storage are separate. A published rate is not a capacity reservation.
FETCHING CURRENT CATALOG
Scroll sideways to see every rate and status →
Current GPU prices for the selected placement
GPU model
Memory
Allocation
USD / hour
Catalog status
PRO 6000 Blackwell
24 GB
MIG partition
—
Loading
PRO 6000 Blackwell
48 GB
MIG partition
—
Loading
PRO 6000 Blackwell
96 GB
Full GPU
—
Loading
H200
141 GB
Full GPU
—
Loading
B200
180 GB
Full GPU
—
Loading
B300
288 GB
Full GPU
—
Loading
vCPU— / vCPU-hour
RAMSystem memory— / GB-hour
Workspace storage— / GB-hour
“Not quoted” means no current price, never free. Availability depends on deployment, placement, balance and capacity. The console confirms your selected configuration before creation. CPU/RAM may need adjustment for your workload.
03 / FIVE PRACTICAL COMPARISONSCHANGE THE ASSUMPTIONS. SEE THE TRADE-OFF.
Match the spend to the work.
These compare workload strategies, not competitor invoices. Time, utilization and performance inputs are illustrative. Monthly estimates use the live configuration above; no guaranteed savings or startup benchmark is implied.
CHOOSE A SCENARIO
01 / 05
SCENARIO 01
A production API with quiet hours.
A service can be reachable all day without keeping its GPU allocated all day—when the workload can tolerate waking from sleep.
A / ALWAYS ALLOCATED
—/ 30 days
B / OIY / IDLE SLEEP
—/ 30 days
Allocated compute720h / month150h / month
Request latencyWarm applicationCold-start delay after sleep
Useful work / allocated time17%80%
UsageManually retain capacityWake on authenticated request
Uses the work and startup/idle inputs above, with the same GPU, CPU, RAM and storage. Sleep follows the idle threshold, not the instant the last request ends. For strict latency targets or steady traffic, keep compute warm.
SCENARIO 02
Inference that doesn’t fill a GPU.
If your model fits a smaller memory slice and meets its latency target, a PRO 6000 MIG profile may be a better fit than a full 96 GB device.
FitHeadroom may sit unusedRoom for this working set
UsageDedicated full GPUOne isolated MIG device
A live full-GPU/MIG price pair is not available here. The price ratio below is an adjustable assumption, not an Oiy quote. CPU, RAM and storage are excluded from this GPU-only comparison. MIG divides resources; throughput and latency must be measured for your model.
SCENARIO 03
Research: time-to-result matters.
A cheaper hourly rate can become expensive when a job runs longer. Compare waiting plus execution time, and price × runtime, before choosing a larger GPU.
UsageWait for cluster admissionRequest capacity when available
Model: 48 execution hours on the baseline; on-demand capacity assumed available with no queue. Speedup and price multiples are assumptions, not measured Oiy results or hardware promises. At equal price and speedup multiples, compute cost is equal. Storage, setup, transfer and funding arrangements are excluded.
SCENARIO 04
Notebooks with an off switch.
Development happens in sessions. Keep your workspace between experiments without paying to keep the same compute allocated every night and weekend.
A / MONTH-LONG INSTANCE
—/ 30 days
B / OIY / WORK SESSIONS
—/ 30 days
Allocated compute720h / month176h / month
Performance while runningSelected GPU profileSame selected GPU profile
Between sessionsCompute stays allocatedWorkspace stays; compute sleeps
UsageKeep the instance runningSave, pause, resume
Assumes 22 workdays/month. Session duration includes all allocated startup and idle time. Same selected compute configuration and retained storage in A and B. Close active notebook connections or explicitly pause; live connections and task leases can keep compute awake.
SCENARIO 05
Batch work, without the empty shifts.
Nightly embeddings, rendering or evaluation often arrive in batches. Allocate for the batch and its startup tail, then release compute until the next run.
A / DEDICATED BATCH WORKER
—/ 30 days
B / OIY / PER-BATCH ALLOCATION
—/ 30 days
Allocated compute720h / month75h / month
Useful work / allocated time8%80%
Performance while runningSelected GPU profileSame selected GPU profile
UsageWorker waits between batchesStart via API; protect the task
One batch/day, with an assumed 30 minutes of combined startup/idle overhead per batch, capped at 24h. Same selected compute rate and retained storage in A and B. Trigger jobs from your own scheduler via the API; this is not a managed scheduling or automatic replica-scaling feature.
Your workload decides the right setup.
Benchmark a representative job. Include startup, idle thresholds, storage and your latency target. Then choose the smallest allocation that gets the work done.