Skip to main content
Deployment cost is based on the hardware selected and the execution time. Every time code runs or a machine is specified to stay running, compute is billed. GPU, CPU, and Memory usage are charged per second; persistent storage is charged per GB per month. View compute pricing on the pricing page. Listed rates apply to the default interruptible compute tier. Apps configured with the protected tier are billed at 2x these rates across GPU, CPU, and memory. See Compute Tier. Deploying a model incurs two billable processes:
  1. Build process — sets up the app environment: a Python environment with the specified parameters, required apt packages, Conda and Python packages, and any model files. A build is only charged when the environment needs rebuilding, i.e., a build or deploy command runs with changed requirements, parameters, or code. Each build step is cached, so subsequent builds cost substantially less than the first.
  2. App runtime — the time code runs from start to finish on each request. Three cost components apply:
  • Cold-start: The time to spin up server(s), load the environment, connect storage, etc. Cerebrium continuously optimizes cold-start latency. Cold-start time is not billed.
  • Model initialization: Code outside the request function that only runs on cold start (e.g., loading a model into GPU RAM, importing packages). This time is billed.
  • Function runtime: Code inside the request function, executed on every request
Example cost calculation A model deployment requires:
  • 24 GB VRam (A10): $0.000306 per second
  • 2 CPU cores: 2 * $0.00000655 per second
  • 20GB Memory: 20 * $0.00000222 per second
Assume the app works on the first deployment, incurring a single 2-minute build. The app has 10 cold starts per day with an average initialization of 2 seconds and an average runtime (predict) of 2 seconds. The expected monthly volume is 100,000 inferences.