Every model you can use for free.
Pick a model and see exactly which providers hand it to you at zero cost.
- Models
- 34
- Open weights
- 25
- Vendors
- 23
Qwen 3Alibaba’s family, from 0.6B you can run on a laptop to MoE giants. Hosted APIs name the exact weight; this page is the family.
Llama 3.3 70BMeta’s classic open model. Nearly every free provider here serves it, which makes it the safest default.
GPT-OSS 120BOpenAI’s open-weights release. Runs on your own hardware, or free on Groq. Cerebras still serves it on a $5 signup credit.
DeepSeek V4The current DeepSeek family: V4 Flash and V4 Pro, 1M of context, thinking as a switch rather than a second model.
Nemotron 3 UltraNVIDIA’s 550B Ultra, tuned for agents and function calling. Free on NIM, OpenRouter (`:free`) and Requesty.
DeepSeek V4 FlashThe current DeepSeek workhorse: 1M of context, thinking as a switch. Official id deepseek-v4-flash; AMD serves DeepSeek-V4-Flash-0731.GLM 5.2Z.ai’s current flagship. Excellent at coding and tool calling. Free on OpenRouter (`z-ai/glm-5.2:free`); on Z.ai itself GLM-5 bills per token.
Qwen CoderThe coding branch of Qwen. Small enough to run locally, good enough to power editor autocomplete.
Qwen 3.8 27BAlibaba’s current 27B: thinking or instruct in one id, images in. Groq serves it at 131K; Cloudflare’s free window reaches 262K.Mixtral 8x7BThe mixture-of-experts model that proved MoE works in the open. Retired from Mistral’s own API; still a great local pick.
Laguna XSPoolside’s small coding model. OpenRouter’s free id is poolside/laguna-xs-2.1; Requesty serves poolside/laguna-xs.2 at 33K.
Gemini 3.7 FlashGoogle’s current Flash workhorse — coding, agents and multimodal input, 1M of context, free to try in AI Studio.
Kimi K3Moonshot’s current open MoE. Strong at agentic tool use and coding. Free on NVIDIA NIM; Groq retired the old K2 ids.
DeepSeek R1The 2025 reasoner that made chain-of-thought open. Gone from DeepSeek’s own API; still free on a few hosts and as local weights.
Qwen3.8-Flash-NextThe Flash-Next 3.8 checkpoint. AMD Radeon Cloud serves this exact id on the $10/day Token Factory quota.Gemma 4 31BGoogle’s open 31B. Cerebras serves gemma-4-31b on a $5 signup credit; Requesty lists google/gemma-4-31b-it at $0.
Nemotron 3 SuperThe 120B Super. Requesty’s docs example and a $0 id on that gateway; also `:free` on OpenRouter.Gemini 3.5 Flash LiteThe cheapest current Gemini. Built for classification, extraction and anything you run thousands of times a day.
GPT-OSS 20BThe smaller open-weights GPT-OSS. Faster than 120B on Groq, same 131K window and 1,000 requests a day.
GPT-OSS Safeguard 20BThe safety classifier in the GPT-OSS line. Groq serves it on the same free chat quota as the other 20B.
DeepSeek V3.1The previous DeepSeek generation. Retired from the official API; SambaNova still hosts V3.1 on the free tier.GLM 4.7 FlashThe permanent free line on Z.ai. 200K of context. glm-4.7, glm-5 and glm-5.2 on that host bill per token.
Qwen 3.6 27BThe previous 27B Qwen still on Groq’s free list. Same 131K window and 30 requests a minute as 3.8.
Qwen 3.7 PlusThe current plus id on Alibaba Model Studio (Singapore / intl). 1M of context, 1M free tokens for 90 days. qwen3-plus is not a current id.
Qwen 3 30B-A3BThe 30B MoE (3B active). Cloudflare serves @cf/qwen/qwen3-30b-a3b-fp8 on the free Workers AI pool.
Qwen 3.5 397BThe 397B-A17B Qwen 3.5. Ollama Cloud’s current Qwen id is qwen3.5:397b.
Qwen 3 8BThe 8B Qwen 3 instruct. SiliconFlow’s free-tier example id; small enough to run locally too.Mistral SmallMistral’s current small line. Free mode on La Plateforme uses this id; Labs models (labs- prefix) are also free of charge.
Command ACohere’s enterprise model, built around RAG and citations. The trial key never expires.
Jamba 1.5 LargeA hybrid Mamba-Transformer. Unusually cheap on long documents thanks to its architecture.Groq CompoundNot one model but a system: Groq routes your request across models and built-in tools like web search.
Groq Compound MiniThe lighter Compound system on Groq. Same tool-routing idea, smaller default stack.
Mistral Medium 3.5Mistral’s balanced model: European hosting and 256K of context. Paid on La Plateforme; not the free-tier chip.
Phi-4Microsoft’s small model that punches far above its size on reasoning benchmarks.