Skip to content
Free LLM Today

Free OpenAI-compatible LLM APIs

Free tiers you can call with the OpenAI SDK by changing the base URL and the model name. Each base URL below was read in the provider’s documentation on the date shown.

Confirmed in the documentation (6)

Free LLM APIs with a confirmed OpenAI-compatible endpoint
ProviderOpenAI SDKFree allowanceRequests / minRequests / day
GroqFree tier
Yes

No logprobs, logit_bias or n > 1.

30 req/min · 1,000 req/day

Chat models; 8K tokens/min and 200K tokens/day.

30

openai/gpt-oss-120b, openai/gpt-oss-20b, qwen/qwen3.8-27b.

1,000

Chat models. Prompt-guard models: 14,400/day.

Google Gemini APIFree tier
Not published

Limits are per project and shown only inside Google AI Studio.

Not published

Not published per model: “Rate limits … can be viewed in Google AI Studio.”

Not published

Not published per model: “Rate limits … can be viewed in Google AI Studio.”

OpenRouterFree tier
20 req/min · 50 req/day

1,000 req/day once the account has bought at least 10 credits.

20

All free model variants together.

50

1,000/day once the account has bought at least 10 credits.

Cloudflare Workers AIFree tier
Yes

Chat completions and embeddings; the base URL contains your account id.

10,000 Neurons/day

Resets at 00:00 UTC. Going above it needs Workers Paid ($0.011 per 1,000 Neurons).

Not published

The allowance is counted in Neurons, not in requests.

Not published

The allowance is counted in Neurons, not in requests.

Hugging Face Inference ProvidersFree tier
Yes

The docs use the OpenAI client against the router URL.

$0.10 of credits/month

Listed as “subject to change”. Past it, usage needs purchased credits.

Not published

The allowance is a monthly credit, not a request count.

Not published

The allowance is a monthly credit, not a request count.

CerebrasTrial only
$5 trial credit, 30 days

Granted “after adding a verified payment method”.

5

Free Trial: gpt-oss-120b and qwen-3.8-27b.

Not published

The Free Trial table lists no requests-per-day figure.

55 of 72 cells above were read on the provider’s own page in the last 30 days (oldest check: ). Each check mark opens that page. The other 17 say “not verified” instead of repeating a number from someone else’s article.

Not confirmed (1)

We did not read a compatibility page for these providers within the last 30 days.

Base URLs and a model to try

OpenAI-compatible base URLs per provider
Providerbase_urlModel id to tryKey variableRead on
Groqhttps://api.groq.com/openai/v1openai/gpt-oss-120bGROQ_API_KEY
Google Gemini APIhttps://generativelanguage.googleapis.com/v1beta/openai/gemini-3.8-flashGEMINI_API_KEY
OpenRouterhttps://openrouter.ai/api/v1openrouter/freeOPENROUTER_API_KEY
Cloudflare Workers AIhttps://api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1@cf/meta/llama-3.1-8b-instructCLOUDFLARE_API_TOKEN
Hugging Face Inference Providershttps://router.huggingface.co/v1deepseek-ai/DeepSeek-V3-0324HF_TOKEN
Cerebrashttps://api.cerebras.ai/v1qwen-3.8-27bCEREBRAS_API_KEY

The same request on each of them

Language
base_url
https://api.groq.com/openai/v1
model
openai/gpt-oss-120b
key from
$GROQ_API_KEY
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.groq.com/openai/v1",
    api_key=os.environ["GROQ_API_KEY"],
)

reply = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "Reply with the word: ok"}],
)
print(reply.choices[0].message.content)
Endpoint read on · console.groq.com

The key is read from an environment variable on your machine. This site has no field for keys and never asks for one.

What to check before you rely on it

Compatible is not identical

Every provider implements the chat completions route; fewer implement everything around it. Groq documents that logprobs, logit_bias and n above 1 are not supported. Cloudflare Workers AI offers chat completions and embeddings on its compatible endpoint. If your app uses tool calling, JSON mode or streaming, test those on the model you picked, because support is per model.

Put three values in configuration

The base URL, the model id and the key are what change between providers. With those three outside the code, moving from one free tier to another, or to a paid tier when the free one ends, is a configuration change.

Cloudflare’s URL contains your account

The Workers AI base URL includes the account id, so it is different for every user. Replace {account_id} with the id shown in your dashboard, and authenticate with an API token.

Model ids are not portable

The same open model has a different id on each provider: openai/gpt-oss-120b on Groq and gpt-oss-120b on Cerebras. Keep the id next to the base URL.

Questions and answers

What does “OpenAI-compatible” mean?

The provider exposes the same HTTP routes and JSON shapes as the OpenAI API, mainly /chat/completions. You keep the OpenAI SDK and change two things: the base URL and the model name.

Which free LLM APIs are OpenAI-compatible?

Confirmed in their documentation on 3 Oct 2026: Groq, Google Gemini API, OpenRouter, Cloudflare Workers AI, Hugging Face Inference Providers. Cerebras is also compatible, but its free offer is a 30-day trial.

Is compatibility complete?

No. Groq documents that logprobs, logit_bias and n greater than 1 are not supported. Cloudflare Workers AI supports chat completions and embeddings on its compatible endpoint. Test the parameters your app depends on.

Can I switch providers without changing code?

Mostly: put the base URL, the model name and the key in configuration. Rate-limit errors and model ids differ, so keep the retry logic and the model name configurable too.