October 10, 2026
A deep Dive into the traits of Seven Open-Weight Models
The short answer
By Zhang Tengfei
12 min read
The short answer
For coding and agent development, look at Qwen 3.8-Max and GLM-5.3-Flash. For long context and enterprise workloads, use DeepSeek V4.1 Flash. If you want to self-host something small and use it commercially without license headaches, pick Gemma 4. For multimodal and multilingual work, look at MiniMax M3 and Mistral Large 4.
API prices across the seven run from about $0.15 per million input tokens to about $15 per million output tokens. Compare input, output and cache-hit prices separately, because the cache-hit price is often where you save the most.
Seven models at a glance
If you're skimming: Qwen and GLM for coding, DeepSeek for long context, Gemma 4 for lightweight commercial use. Every model below gets the same three spec columns (scale, architecture, usage), followed by what the model is for.
Reading the specs
- Look at active parameters, not total parameters. The total tells you how much disk the weights need; the active count tells you how much compute each generated token costs. Kimi K3 has 2.8 trillion parameters and GLM-5.3-Flash has 320B, nearly a 9× difference, yet GLM is cheap to run because it activates only 18B per step.
- Weight memory ≈ total parameters × bytes per parameter. That's about 2 bytes at FP16 and about 0.5 bytes with INT4 quantization. Inference memory depends mostly on active parameters and the KV cache, and has little to do with the total.
- A 1M-token window costs nothing until you fill it. The KV cache grows linearly with the tokens you actually put in, and that's what you pay for.
DeepSeek V4.1 Flash: a 552B MoE with the KV cache cut to 1/4 of the previous generation
Long context · enterprise workloads · MIT license
Of the flagships here, this one needs the least memory for long-context and enterprise work.
DeepSeek built it for complex reasoning, code generation and agent tasks, and says it beats the company's own flagship, DeepSeek-V4-Pro, on performance, cost, speed and total runtime.
- The model is asymmetric: 8B parameters are active on the input side and 16B on the output side, so reading long documents takes less compute while answer quality holds up.
- At 1M context it needs a quarter of the previous generation's GPU memory and an eighth of its storage, which puts long-context work within reach of much smaller budgets.
- It takes text and images directly, with a 1M-token context and up to 384K output tokens.
- Since 2026–09–14, requests to
deepseek-v4-proare billed at Flash rates and old model names keep working, so existing DeepSeek API users don't have to change any code.
In practice, 552B is the size of the weights (about 1.1 TB on disk) and 8B / 16B is the compute that runs each step. Because so few parameters wake up, each token is cheap and throughput is high, but self-hosting means finding a lot of storage first. MoE, the asymmetric layout and KV cache compression together let the same memory hold a longer context, and the 1M window is big enough for whole-repository retrieval and long-running agents. The MIT license lets you use, modify and redistribute it commercially. The API model name is deepseek-flash, old names are routed automatically so upgrading costs nothing, and the Hugging Face repository includes a technical report.
Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. DeepSeek is an AI research company focused on building world-leading general artificial intelligence. We develop and…
Kimi K3: the first open-source flagship at the 3-trillion-parameter scale
Long context · long-horizon coding · custom license
Reach for it when you need the highest capability and cost isn't a concern.
It's Moonshot AI's most capable model so far and the first open-source model at the 3-trillion scale. It targets long-horizon coding, knowledge work and deep reasoning, and Moonshot puts its overall scaling efficiency at about 2.5× that of K2.
It's Moonshot AI's most capable model so far and the first open-source model at the 3-trillion scale. It targets long-horizon coding, knowledge work and deep reasoning, and Moonshot puts its overall scaling efficiency at about 2.5× that of K2.
_- KDA linear attention plus attention residuals make very long contexts faster to compute and lighter on memory.
- Only 16 of its 896 experts wake up for each token. That sparsity is what keeps a 2.8T model practical to run.
- Moonshot says it can keep going on long engineering tasks with very little human supervision: working through large codebases, coordinating terminal tools, and using screenshots and visual feedback to refine games, front ends and CAD models.
- Thinking mode is always on.
reasoning_effortsupports low / high / max (default max) to control reasoning depth and cost._
With 104B parameters active per step, it carries the heaviest compute of the seven, and both quality and latency sit at the top. The weights are about 5.6 TB, so self-hosting starts at multiple nodes (see section 4). Native vision accepts images and video, the 1M context covers long codebase tasks, and uploaded video files are referenced with ms://.
Research, deployment, fine-tuning and internal use are free. A model-serving business with more than $20M in revenue over 12 consecutive months needs a separate agreement, and very large products must give attribution. The API model name is kimi-k3. The flagship has to be unlocked with a top-up (minimum ¥10), and new-user vouchers don't apply to it.
Kimi K3 Tech Blog: Open Frontier Intelligence Kimi K3 is the world's first open 3T-class model - frontier performance across coding, knowledge work, and reasoning…
MiniMax M3: the first flagship to open-source three frontier capabilities at once
Multimodal · 1M context · community license (conditional commercial use)
Pick it if you need coding, long context and multimodality in one open-weight model.
MiniMax describes it as the first open-weight model with frontier coding, million-token context and native multimodality. It's aimed at AI coding assistants and long-running automation; until now, only a few closed models offered all three.
- Compared with M2 at 1M context, prefill is 9× faster, decode is 15× faster and per-token compute drops to 1/20, which makes long documents and long videos much cheaper to process.
- Text, images and video were mixed in from the first training step, so multimodality is part of the base model.
- It reproduced an ICLR outstanding paper on its own in 12 hours (18 commits, 23 experiment figures). In another run it optimized a CUDA kernel over 147 iterations and raised hardware utilization from 7.6% to 71.3%, a 9.4× speedup, with no human intervention.
- The
thinkingparameter accepts enabled / adaptive / disabled; adaptive decides for itself when to reason. API caching is automatic and needs no setup.
At 23B active out of 428B, it sits between GLM and the big flagships, and long agent runs stay affordable. The weights are about 856 GB, upper-middle tier for self-hosting (see section 4). MSA is a sparse attention operator designed for million-token contexts. The API guarantees at least 512K of context, and long-video understanding and long-horizon agents benefit most.
Non-commercial use is free. Commercial use requires the credit "Built with MiniMax M3", and annual revenue above $20M requires written authorization. It has the widest framework support of the seven: SGLang, vLLM, Transformers, KTransformers, unsloth and ATOM. The technical report is arXiv:2606.13392.
MiniMax M3 - Coding & Agentic Frontier, 1M Context, Multimodal MiniMax M3 reaches frontier-level performance on coding and agentic tasks, with a 1M context window powered by the MSA…
Gemma 4: an Apache 2.0 family that runs on everything from a Raspberry Pi to a workstation
Lightweight self-hosting · on-device · Apache 2.0
Pick it to run a model locally on a phone or a Raspberry Pi, or to use one commercially with no strings attached.
This is Google DeepMind's most capable Gemma family so far. Google says every size performs at the frontier for its class, from phones and edge devices (E2B, E4B) to consumer GPUs and workstations (26B, 31B), across reasoning, agent workflows, coding and multimodal understanding.
- Multi-token prediction (MTP) speeds up decoding by up to 2.2× on mobile GPUs and up to 1.5× on CPUs, with hardware such as the M4 MacBook seeing a clear gain. Google reports no quality loss.
- The E2B model in on-device format is 2.58 GB. It runs at about 8 tokens per second on a Raspberry Pi 5 and decodes at 56 tokens per second on the iPhone 17 Pro's GPU. The LiteRT-LM toolchain converts fine-tuned models to that format.
- E2B and E4B are for phones and edge devices, 12B sits in the middle, the 26B A4B MoE favors efficiency and the 31B Dense favors quality. No other family among the seven spans devices to workstations.
- Yale's Gemma-based Cell2Sentence-Scale 27B found a potential new pathway for cancer treatment, DolphinGemma is used to study dolphin communication, Living Models uses Gemma 4 to decode plant DNA, and Syngenta uses it to identify plants more accurately.
The 31B Dense weights are about 61 GB at FP16, while E2B in on-device format is 2.58 GB, two orders of magnitude apart within one family. That gives Gemma 4 the lowest self-hosting bar of the seven (see section 4) and makes offline use practical. Its 256K context (128K on small sizes) is short next to the others, but MTP makes generation fast. It handles reasoning, coding and agent tasks, and the quick on-device decoding makes it the strongest fit for offline local apps.
Apache 2.0 allows commercial use, modification and redistribution with no extra conditions. That makes it the easiest of the seven to customize and redistribute, and it's a common base for governments and companies building sovereign AI systems. You can try it free in AI Studio; Vertex hosted pricing covers only 26B A4B
Gemma 4 | Google AI Edge | Google for Developers Gemma 4 models are designed to deliver frontier-level performance at each size, targeting deployment scenarios from…
Mistral Large 4: the strongest open-weight flagship for cybersecurity
Cybersecurity · European sovereignty · weights pending
Pick it for offensive and defensive security work where a closed model's refusals would get in the way.
It's Mistral's largest and most capable flagship so far, nicknamed "Le Chonk". Mistral says it competes with the strongest open-source models in the world, clearly beats any US or European open-weight model, and leads open models on enterprise workloads such as cybersecurity, finance and law.
- It ranks in the global top five on the AA cyber index and first among open-weight models from outside China. It scored 82% on a test of reproducing and patching vulnerabilities, the highest of any model, and solved 93% of Cybench challenges. Claude Opus 5.5 and GPT-6 Astra scored close to zero on the same test because they refuse to carry out the tasks, and defenders need a system that won't.
- On Dense 200 it scored 42%, ahead of GPT-6 Astra at 41%. Mistral demonstrated it inspecting gigapixel satellite imagery for disaster relief, zooming in to check engineering drawings and pulling evidence out of PDFs.
- In vals.ai's evaluation it beat GPT-6 Astra on both legal and financial tasks. It outperformed every open-source model on Harvey's legal agent benchmark, and its 59.9% on AutomationBench is ahead of Kimi K3 and DeepSeek V4 Pro.
- Mistral trained it from scratch on 3,800 Grace Blackwell GPUs in its own European data centers. European deployments run end to end under EU law, independent of other cloud providers, and the training data covers more than 160 languages, including every official EU language.
At 1.05T parameters with 52B active per step, it's Mistral's largest model. For now it's only available through the preview API; once the weights are out, self-hosting will fall in the giant tier (see section 4). Mistral calls its image understanding "a step change", and perception-heavy agent work (satellite imagery, engineering drawings, evidence retrieval from PDFs) is where it does best.
The weights are due at the end of the month and the license hasn't been announced. You can try it in Mistral Studio, but check the current status of the weights and license before using it commercially. Mistral says it will publish more about the architecture and post-training methods when the weights come out.
Introducing Mistral Large 4 | Mistral The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and…
Price comparison
These numbers come from the Vals Index, which measures how often each model completes agent tasks in finance, coding, law and other fields, along with what each task costs. The cost differences run to an order of magnitude and beyond.
On the same set of real tasks, the most expensive closed model costs $28.92 per task and the cheapest open model $0.30. The price gap is 95×, and the completion rates are only about 11 percentage points apart. Data checked on 2026–10–09 against the official Vals Index page.
Closed vs. Open-Source LLMs: Current State, Capability Gap, and How to Choose Using Arena and Vals Index as a common yardstick, this post measures how far open-source LLMs trail closed models on…
AI Benchmarks for Real-World Tasks | Vals AI Independent AI evaluations and benchmarks on real-world tasks in finance, software, cybersecurity, healthcare, and…
Pick by use case
The picks below start from the job you need done. Each of the four common scenarios gets a top pick and the reasoning behind it.
Scenario 1: coding and agent development
Writing code, running tool calls and building workflows. Top picks: Qwen 3.8-Max and GLM-5.3-Flash.
By their makers' accounts, both lead in agentic coding. GLM-5.3-Flash suits high-frequency calls in particular: with a $0.15 input price and an MIT license, an agent loop can burn through tokens without the bill hurting. Switch to Qwen 3.8-Max when coding quality matters most.
Scenario 2: long context and enterprise workloads
Long documents, knowledge bases and enterprise automation. Top pick: DeepSeek V4.1 Flash.
KV cache compression brings the memory and storage cost of long context down to a quarter of the previous generation, and the 1M context fits long documents, knowledge bases and long-running agents. If you need stronger long-horizon engineering, Kimi K3 is the step up, but read its commercial license terms first.
Scenario 3: lightweight self-hosting and commercial compliance
Running locally, using the model commercially and keeping GPU spend down. Top picks: Gemma 4 and Kimi K3.
Gemma 4's Apache 2.0 license allows commercial use with no extra conditions, and the 31B Dense model runs on consumer hardware once quantized. Kimi K3 is the largest model here and can do the most, but its license adds conditions you'll need to check before commercial use. Teams on a tight budget can also start with free tiers.
Scenario 4: multimodal and multilingual
Understanding images, working across languages and serving several markets. Top picks: MiniMax M3 and Mistral Large 4.
MiniMax M3 is natively multimodal and already well proven in coding and automation. Mistral Large 4 supports more than 160 languages, which matters for products serving several markets. It's still in preview, though, and its pricing and weight status may change, so confirm the latest official details before you commit to a launch.
4. Self-hosting cost estimates
You can also deploy these models yourself. The figures in the table are rough orders of magnitude, meant only to tell you whether self-hosting is within reach.
These are estimates, not official hardware recommendations. Real requirements depend on the quantization scheme, batching and inference framework.
The memory column covers two separate costs: weight memory and inference memory. Weight memory is roughly total parameters × bytes per parameter. FP16 uses 2 bytes per parameter and INT4 uses 0.5, which is where the lightweight tier's 60 to 80 GB (24 to 40 GB quantized) comes from; Gemma 4 31B is about 61 GB at FP16 and about 15 GB at INT4. Inference memory depends mainly on active parameters and the KV cache. GLM-5.3-Flash has 640 GB of weights, but it wakes only 18B per step, so running it doesn't take 640 GB of GPU memory. For the giant tier, "TB range" refers to the weights themselves: about 4.8 TB for Qwen 3.8-Max and 5.6 TB for Kimi K3. Storing them alone takes several machines, which is why people mostly use this tier through APIs.
In the GPU column, "80 GB" means enterprise cards such as the H100 or A100, each several times the price of a consumer card. "1× 80 GB" means a single card is enough, "4–8× 80 GB" means four to eight cards in one machine, and "16+× 80 GB" means a cluster of several machines. A consumer card with around 24 GB is enough only for a quantized lightweight model or a small on-device model such as Gemma E2B or E4B.
Some practical advice:
- Don't rush to buy GPUs. Get your product working on an API, measure real usage, then work out whether self-hosting would pay for itself.
- If you want to try self-hosting, start with Gemma 4. One card, even a consumer one, is enough to begin.
- The mid-size tier (models like GLM) only makes sense for teams with high, steady usage. Even quantized it needs four to eight 80 GB cards, so compare a month of GPU rental against your API bill first.
- For the giant tier, start on the API.
- If your team won't keep the hardware busy over the long run, the API will cost less
Wrap-up
A few patterns hold across all seven models. Total parameters keep growing while active parameters stay small, and the active count is what drives cost. The price gap between open and closed models is measured in multiples: $0.30 against $28.92 for the same task, or 95×. Licenses fall into three groups: MIT, Apache 2.0, and licenses with commercial conditions. Choosing a model comes down to your use case, how you'll deploy it and which license terms you can accept. The table maps each model to the teams it suits.
The table draws on the earlier sections: official release information (section 1), Vals Index completion rates and costs (section 2) and the self-hosting estimates (section 4). Data checked on 2026–10–09.