Replicate
Replicate is a platform for running open source and custom machine learning models through a hosted API, without managing GPU infrastructure yourself. Most models are billed by how long they run on a given hardware tier, from CPU up to multi-GPU H100 and B200 configurations, while some popular models are billed per image or per token instead. There's no monthly subscription; you pay only for compute used.
Facts checked on against the sources listed at the end of this page.
- Pricing
- Pay as you goPaid from about $0.000025 a second for small CPU instances, varying by model and hardware
- Open source
- NoProprietary (the platform); hosted models vary by their own licenses
- GitHub stars
- 597
- npm downloads a week
- 705kup 15% on the week before
- Latest release
- v1.4.010 months ago
A good fit for
- Running open source AI models without managing GPU infrastructure
- Prototyping with many different models before committing to one
- Workloads with spiky or unpredictable usage, since there's no monthly minimum
- Deploying a custom or fine-tuned model behind a simple API
Not a good fit for
- Very high, constant-volume inference where owning or reserving dedicated GPU capacity elsewhere would be cheaper
- Teams that want a flat, predictable monthly bill rather than usage-based pricing
- Applications needing the absolute lowest latency, where a self-managed, always-warm deployment may outperform a shared platform
Replicate pricing
There's no subscription; you pay per second of compute or per output, depending on the model. CPU instances start at roughly $0.000025 a second, GPU instances range from about $0.000225 a second for a T4 up to higher rates for H100 and B200 hardware, and volume discounts are available for high spend.
Prices in USD, checked on 27 September 2026. Current prices on replicate.com.
Main features
- Hosted model API
- Call thousands of public open source models through a single REST API without hosting them yourself.
- Custom model deployment
- Package and deploy your own model or fine-tune using Cog, Replicate's open source tool for containerizing ML models.
- Per-second billing
- Most models are billed by the time they take to run, priced per second by hardware tier.
- Per-output pricing
- Some popular models charge a flat rate per image or per thousand tokens instead of compute time.
- GPU hardware range
- From CPU instances up to multi-GPU Nvidia H100 and B200 configurations.
- Webhooks
- Get notified when a long-running prediction finishes instead of polling for status.
About Replicate
Replicate lets developers run machine learning models, most of them open source, through a hosted API without setting up GPU servers. Thousands of public models are available to call directly, covering image generation, video, audio and language tasks, alongside the ability to deploy your own custom or fine-tuned model.
Billing is usage-based. Most models charge by the second the model runs on a given hardware tier, from small CPU instances to single and multi-GPU configurations up to 8x Nvidia B200. Some widely used models, such as certain image and language models, are billed per output (per image, or per thousand tokens) instead of raw compute time.
Private, dedicated deployments bill for setup and idle time as well as active compute, unless using fast-booting fine-tunes that only charge for active use. Replicate also offers volume discounts and account management for high-spend customers.
- Made by
- Replicate
- Free plan
- No ongoing free plan; billing is pay-as-you-go with no monthly minimum
- Available on
- REST API, Python, Node.js, Webhooks
Usage and activity
Weekly npm downloads of replicate, last 12 months.
Recent releases
Replicate alternatives
Replicate compared
Questions about Replicate
Does Replicate have a free plan?
No ongoing free tier. Replicate is pay-as-you-go, billed by compute time or by output, with no monthly subscription.
How is Replicate priced?
Most models are billed per second of compute on the hardware tier they run on. Some popular models are billed per image or per thousand output tokens instead.
Can I deploy my own model on Replicate?
Yes, using Cog, Replicate's open source tool for packaging models into containers that can run as a hosted API.
What hardware does Replicate offer?
CPU instances and GPU tiers ranging from Nvidia T4 up to multi-GPU H100 and B200 configurations.
Sources and links
Work on Replicate? Claim this listing. Something wrong or out of date? Suggest a correction. Stats come from the public GitHub and npm APIs and update every day.
