Maintained buyer guide

Free AI Inference: Credits, Limits and Catches

Free inference is not one market. Some providers give repeatable daily capacity, some expose a prototype sandbox, and others rotate zero-price models until demand changes. This guide compares what you can actually build before paying.

5 current providersClaims checked 2026-09-05
Hetzner

free while 'experimental'; Hetzner promises advance email notice before any status change

Verified
Useful monthly capacity
≈15B input + 150M output tokens per 30 days, if continuously available
Models / fit
Qwen/Qwen3.6-35B-A3B-FP8
Card
unknown
After free use
Requests are rate-limited. Hetzner does not promise continued availability.
Official source · Sep 5, 2026
Cloudflare Workers AI

Not stated (ongoing free allocation)

Verified
Useful monthly capacity
300,000 Neurons per 30 days
Models / fit
Cloudflare Workers AI model catalog
Card
no for Workers Free
After free use
Free-plan operations fail until reset. Workers Paid charges $0.011 per 1,000 Neurons above the free allocation.
Official source · Sep 4, 2026
Nous Research

Not stated (rotating catalog; availability can change)

Verified
Useful monthly capacity
Six hosted models at $0 prompt and $0 completion token pricing (rotating; down from 7). Current zero-priced IDs: inclusionai/ling-3.0-flash-fin:free, meituan/longcat-2.0:free, poolside/laguna-s-2.1:free, poolside/laguna-xs-2.1:free, stepfun/step-3.7-flash:free, upstage/solar-pro4:free.
Models / fit
Card
unknown
After free use
The provider's current standard pricing or limits apply.
Official source · Sep 4, 2026
Nous Research

not stated; catalog rotates while in free promotion

Verified
Useful monthly capacity
Seven zero-priced hosted models on Nous Portal: inclusionai/ling-3.0-flash-fin:free, inclusionai/ling-3.0-flash-sante:free, meituan/longcat-2.0:free, poolside/laguna-s-2.1:free, poolside/laguna-xs-2.1:free, stepfun/step-3.7-flash:free, upstage/solar-pro4:free.
Models / fit
Card
unknown
After free use
The provider's current standard pricing or limits apply.
Official source · Sep 5, 2026
OpenRouter

Not stated

Verified
Useful monthly capacity
Free router (`openrouter/free`) auto-selects from 20+ zero-cost models; no API key charges, no credit card; typical free-model limits around 20 requests/minute and 200 requests/day (per OpenRouter docs/community listings).
Models / fit
Card
unknown
After free use
The provider's current standard pricing or limits apply.
Official source · Sep 4, 2026

Start with workload shape, not the biggest number

A huge daily token ceiling is useful only if the endpoint supports your model, input type, latency target, and reliability needs. Hetzner's experiment is unusually generous but explicitly unstable. Cloudflare's allowance is durable but measured in Neurons, so model choice changes how far it goes. NVIDIA Build is a prototype environment rather than a promised monthly bucket.

For a small evaluation harness, a rotating free-model router may be enough. For a customer-facing feature, treat every free path as disposable capacity and keep a paid fallback.

Permanent free tier vs promotional access

Repeatable free tier

A documented allowance resets on a known cadence. Cloudflare's daily Neuron allocation is the clearest example in this snapshot.

Experiment or rotating catalog

The provider can remove a model, throttle access, or end the experiment without a published date. Hetzner and rotating :free catalogs belong here.

The privacy catch is part of the price

A zero-dollar request can still carry a data trade-off. Google marks free-tier usage differently from paid-tier usage on its pricing table. Router services may select providers with different retention rules. Never send secrets, production customer data, or regulated records until you have checked the current provider and model terms.

  • • Prefer synthetic evaluation data for first tests.
  • • Pin a provider when router privacy policies differ.
  • • Re-check terms before moving from prototype to production.

A practical selection order

  1. 1. Define the minimum model and modality.

    Text-only extraction, vision, coding agents, and embeddings need different catalogs.

  2. 2. Convert limits into your own requests.

    Estimate input, output, retries, and concurrency—not just headline tokens.

  3. 3. Check card, region, and data terms.

    These can disqualify an otherwise attractive tier.

  4. 4. Price the fallback before launch.

    Know what happens after the allowance and how quickly you can switch.

See only the offers that are current now

The guide explains the trade-offs. The live Deals page handles expiry, verification windows, and claim links.

Open verified inference deals