For most enterprises in 2026, the open-weight shortlist is Qwen 3 (Apache 2.0, 119 languages, dense models to 32B and MoE to 235B), Llama 3.3 70B or Llama 4 for the broadest tooling, DeepSeek V3 or R1 (MIT) for frontier-class reasoning on an 8-GPU H200 or B200 node, Mistral Small 3.x for cost-efficient assistants with strong European-language coverage, Gemma 3 27B for single-GPU deployments, and Phi-4 for small on-premise or edge use. Choose by running your own task evaluations first, then filter on license terms, language coverage, hardware fit and ecosystem support rather than public leaderboard rank. Figures below are approximate as of 2026; check the current model card and license text before you commit.

The open-weight LLM landscape in 2026

"Open-weight" means the vendor publishes the trained parameters so you can run them on your own hardware. It does not always mean "open source": Apache 2.0 and MIT models fit the Open Source Initiative definition in spirit, while Llama and Gemma ship under vendor terms that allow commercial use with conditions. The distinction matters for redistribution, derivative naming, usage restrictions and the risk that terms change between versions.

The table covers the six families that dominate enterprise deployments as of 2026. For mixture-of-experts (MoE) models the active parameter count is shown because it drives compute cost, while total parameters drive memory. Context windows are model-card maximums, which are rarely the practical serving limit.

FamilyMain variants (as of 2026)LicenseLanguagesContext (approx.)Best fitTypical hardware
Llama 3.1 / 3.3 (Meta)3.1: 8B, 70B, 405B; 3.3: 70BLlama Community License8 official128KGeneral assistant, RAG, fine-tuning base, widest tooling70B: 2× H100 FP16 or 1× H200 / RTX PRO 6000 FP8; 405B: 8× H100 FP8
Llama 4 (Meta)Scout ~109B total / 17B active; Maverick ~400B / 17B active (MoE)Llama 4 Community License12 officialScout 10M advertised, Maverick 1M; practical limits far lowerLong context, image plus textScout: 1× H100 INT4 or 2× H100 FP8; Maverick: 8× H100 FP8 or 8× H200
Qwen 2.5 (Alibaba)0.5B to 72B dense; Qwen2.5-Coder 0.5B to 32B; VLApache 2.0 (3B and 72B: Qwen licenses)29+128K (32K native, YaRN)Code (Coder-32B), multilingual, fine-tuning7B to 14B: one 24 to 48 GB GPU; 32B: 1× H100; 72B: 2× H100 FP16 or 1× H200 FP8
Qwen 3 (Alibaba)Dense 0.6B to 32B; MoE 30B-A3B, 235B-A22B; 2507 Instruct/Thinking; Qwen3-Coder; Qwen3-Next 80B-A3BApache 2.011932K native, 128K YaRN; 2507 and Coder up to 256KMultilingual, reasoning (thinking mode), agents, code32B: 1× H100; 30B-A3B: one 48 to 80 GB GPU; 235B-A22B: 8× H100 BF16 or 4× H100 / 2× H200 FP8
DeepSeek V3 / R1V3, V3-0324, V3.1, V3.2 (671B total / 37B active); R1, R1-0528; R1-Distill 1.5B to 70BMIT (distills inherit base terms)English, Chinese strongest128KFrontier-class reasoning, agents, code, math8× H200 or 8× B200 at native FP8; 8× H100 only with INT4 or two nodes; distills on one GPU
MistralSmall 3.x 24B; Ministral 3 (3B/8B/14B); Mixtral 8x7B, 8x22B; Magistral Small 24B; Devstral Small; Large 2 123B; Large 3 675B-A41BApache 2.0 for current models; Large 2 and Codestral under research / non-production licensesDozens; strongest in European languages32K (Mixtral 8x7B) to 128K (Small 3.x), ~256K (Large 3)European vendor, cost-efficient assistants, function calling, edgeSmall 3.x: 1× H100 FP16 or one 24 GB GPU INT4; 8x22B: 4× H100; Large 3: 8× H200 or B200
Gemma 3 (Google)1B, 4B, 12B, 27B; QAT INT4 checkpoints; Gemma 3n E2B/E4BGemma Terms of Use140+ pretrained, 35+ tuned128K (1B: 32K)Single-GPU assistants, vision plus text, multilingual, edge27B: 1× H100 BF16 or one 24 GB GPU QAT INT4; 4B, 12B on workstation GPUs
Phi-4 (Microsoft)Phi-4 14B; Phi-4-mini 3.8B; Phi-4-multimodal 5.6B; Phi-4-reasoning 14B; mini-reasoningMITEnglish first; limited multilingual16K (Phi-4), 32K (reasoning), 128K (mini)Small on-premise, edge, STEM reasoning at small size14B: one 24 to 48 GB GPU (INT4 fits 16 GB); mini variants on laptops, CPUs, NPUs

Also track OpenAI gpt-oss (20B and 120B, Apache 2.0), Kimi K2 (modified MIT), GLM-4.5 (MIT) and NVIDIA Nemotron (NVIDIA Open Model License). Every family in this table is production-usable; license, language, hardware and ecosystem decide the choice, not headline benchmark rank.

License review: what each family lets you do

Read the license before you run the first benchmark, because a license problem is the only issue engineering cannot fix. The table is a summary, not legal advice; check the current license text on the vendor's site, since terms change between versions (DeepSeek moved V3 to MIT in 2025, and Mistral released Large 3 under Apache 2.0 after Large 2 was research-only).

FamilyLicenseCommercial useDerivatives and redistributionClauses to check
Llama 3.x / 4Llama Community LicenseYesYes, with "Built with Llama" attribution and "Llama" in derivative namesSeparate license above 700M monthly active users; acceptable-use policy; Llama 4 grants no multimodal-model rights to EU-domiciled entities
Qwen 3Apache 2.0YesYesKeep notices; patent grant. Qwen 2.5 3B (research license) and 72B (Qwen License, 100M MAU clause) differ
DeepSeek V3 / R1MITYesYesMinimal; R1-Distill models also carry Qwen or Llama base terms
Mistral (current)Apache 2.0YesYesLarge 2 (Mistral Research License) and Codestral (Non-Production License) exclude production use without a contract
Gemma 3Gemma Terms of UseYesYes; terms and Prohibited Use Policy must flow downGoogle may update the Prohibited Use Policy; not OSI-approved
Phi-4MITYesYesMinimal

Three checks catch most problems: confirm who the licensee is (an EU-headquartered company evaluating Llama 4 must read the multimodal clause with counsel), confirm what you will ship (an internal assistant carries fewer obligations than a product that redistributes fine-tuned weights), and record the license version and model commit hash in your architecture decision record. Apache 2.0 and MIT models (Qwen 3, DeepSeek, current Mistral, Phi-4) minimize legal review; Llama and Gemma are workable but need a documented reading of their conditions.

Hardware fit: what it takes to serve each family

Weights set the floor: about 2 bytes per parameter at FP16, 1 at FP8 and 0.5 at INT4, so a 70B model is roughly 140 GB, 70 GB or 38 GB of weights. Add 20 to 50 percent for the KV cache at realistic concurrency and context, more for 64K-plus contexts. For MoE models, memory follows total parameters (all experts stay resident) while per-token compute follows active parameters, which is why Qwen3-235B-A22B produces tokens faster than Llama 3.1 405B on the same node.

Map that to the NVIDIA parts enterprises buy as of 2026: H100 80 GB HBM3, H200 141 GB HBM3e, B200 with about 180 GB HBM3e, and RTX PRO 6000 with 96 GB GDDR7. One H200 or RTX PRO 6000 serves a 70B model at FP8 with headroom; one H100 serves 24B to 32B dense models at FP16 or 70B at INT4. The 671B-class models (DeepSeek V3/R1, Mistral Large 3) need an 8-GPU H200 or B200 node at native FP8; an 8× H100 node has 640 GB, not enough for FP8 weights plus cache, so it needs INT4 or a second node. Per-model GPU counts and tensor-parallel layouts are in How many GPUs do you need for 70B, 405B and DeepSeek R1?.

Small models change the economics: Phi-4 14B, Gemma 3 12B, Qwen3 8B and Mistral Small 24B at INT4 all fit in 24 GB, which means MIG slices on an H100, one RTX PRO card or existing workstation GPUs. Match the model class to the smallest GPU tier that leaves KV-cache headroom at peak concurrency; the serving stack, compared in vLLM vs TensorRT-LLM vs Ollama vs SGLang, then decides throughput.

How to choose: a five-step selection method

Public leaderboards help build a shortlist and are unreliable for the final decision: they use prompts, languages and tasks that are not yours, and training-data contamination inflates some scores. The method below is what Nanobase AI runs in model-selection engagements; it takes two to four weeks for a typical enterprise.

  1. Build a task benchmark from your own data. Collect 200 to 500 real examples per task (support tickets, contracts, code reviews, SQL questions) with reference answers. Score candidates with the same prompts, serving stack and quantization you will deploy. Track accuracy, latency at target concurrency and cost per million tokens.
  2. Run the license review. Map each shortlisted model to your use: internal tool, customer-facing product, fine-tuned derivative or redistributed weights. Note entity domicile, MAU thresholds, attribution and acceptable-use obligations, and get written sign-off from legal before the pilot.
  3. Test the languages you actually serve. A "supports 119 languages" claim says nothing about quality in Turkish legal text or German technical manuals. Measure tokenizer efficiency too: a language that needs two to three times the tokens per word of English costs proportionally more and consumes context faster.
  4. Confirm hardware fit. Compute weights plus KV cache at peak concurrency and context, and confirm the planned quantization does not degrade your benchmark beyond an agreed tolerance (1 to 2 points is typical). Prefer models with vendor-published FP8 or INT4 checkpoints.
  5. Check ecosystem support. Verify day-one support in vLLM, TensorRT-LLM or NVIDIA NIM, a working chat and tool-calling template, guardrail integrations and fine-tuning recipes. A model the ecosystem supports poorly costs more in engineering than it saves in quality.

Close with a pilot behind a model-agnostic API layer so you can swap models without changing applications. Benchmark first, because it eliminates most candidates cheaply, then apply license, language, hardware and ecosystem as filters on the survivors.

Recommendations by scenario

These shortlists assume self-hosting on NVIDIA GPUs and are inputs to step one of the method, not final answers.

General enterprise assistant (chat, summarization, RAG)

Start with Qwen3-32B or Llama 3.3 70B; both have mature tool-calling templates, reliable RAG behavior and wide fine-tuning tooling. If one node must serve many concurrent users cheaply, Qwen3-30B-A3B or Mistral Small 3.x give near-32B quality at a fraction of the compute. Gemma 3 27B is the best single-H100 option when image input matters. Whether you need fine-tuning at all is covered in RAG vs fine-tuning.

Multilingual workloads

Qwen 3 (119 languages) and Gemma 3 (140-plus in pretraining) have the widest declared coverage; Mistral Small 3.x is strongest for French, German, Spanish, Italian and other European languages; Llama 4 covers 12 languages officially. DeepSeek is strongest in English and Chinese and Phi-4 is English-first, so both need careful benchmarking before multilingual use. Test the specific language rather than trusting the count.

Reasoning and agents

DeepSeek R1 and V3.1/V3.2 (MIT) are the reference open-weight reasoning models as of 2026 and justify an 8× H200 node when agent quality drives revenue. Qwen3-235B-A22B in thinking mode is the most accessible alternative and fits on 4× H100 at FP8; Qwen3-32B thinking mode and Magistral Small 24B cover the single-GPU tier. For agents, check tool-call reliability and structured-output support in your serving stack, not only reasoning quality.

Small on-premise and edge deployments

For a single 24 to 48 GB GPU or a MIG slice, pick from Phi-4 14B (MIT), Gemma 3 12B, Qwen3 8B or 14B, and Mistral Small 3.x at INT4. Below that, Phi-4-mini 3.8B, Gemma 3 4B and 3n, Qwen3 4B and Ministral 3 run on laptops, CPUs and NPUs. Small models gain the most from task fine-tuning because they have less general capacity to spare.

Code generation and review

Qwen3-Coder (30B-A3B for one GPU, 480B-A35B for a node) and Qwen2.5-Coder-32B are the safest Apache 2.0 choices. DeepSeek V3.x is the strongest for agentic coding when the hardware budget allows. Devstral Small (Apache 2.0) targets agent-style repository tasks, while Codestral needs a commercial agreement for production use. Qwen 3 appears on every shortlist because of its license, size range and language coverage, which is why it is the default candidate in most 2026 selections.

Model update cadence and version pinning

Open-weight families now ship on a software-like cadence. As of 2026, Meta releases a major Llama version roughly every 6 to 12 months with point releases between; Alibaba ships Qwen updates every few weeks (dated snapshots plus Coder, VL and Next variants); DeepSeek has shipped a V3 or R1 revision every 2 to 4 months; Mistral, Google and Microsoft each publish several updates per year. A new snapshot under the same family name can change tokenizer behavior, chat template, tool-calling format and output style.

Treat the model as a versioned dependency: pin the exact Hugging Face revision (commit hash) in your deployment manifest, keep the previous version deployable for rollback, and re-run the step-one benchmark on every candidate update before promotion. Review the shortlist quarterly.

vllm serve Qwen/Qwen3-32B \
  --revision <commit-sha-from-model-card> \
  --served-model-name assistant-v3 \
  --max-model-len 32768

Pinning by commit hash and re-benchmarking on every update turns a fast-moving model ecosystem from a risk into an advantage.

Frequently asked questions

Which open-weight LLM is best for enterprise use in 2026?

There is no single best model. As of 2026 Qwen 3 is the most common default because it is Apache 2.0, covers 119 languages and spans 0.6B to 235B parameters; Llama 3.3 70B has the widest tooling; DeepSeek V3/R1 leads open-weight reasoning; Mistral Small, Gemma 3 and Phi-4 win at the single-GPU tier. Benchmark on your own tasks, then filter by license, language, hardware and ecosystem.

What is the difference between open-weight and open-source LLMs?

An open-weight model publishes its trained parameters for download and self-hosting. An open-source model, under the OSI definition, also grants unrestricted use, modification and redistribution and ideally documents its training data. Apache 2.0 and MIT models such as Qwen 3, DeepSeek and Phi-4 are open source in the license sense; Llama and Gemma are open-weight under vendor terms that allow commercial use with conditions.

Can I use Llama 4 commercially in the European Union?

The Llama 4 Community License states that rights to the multimodal models are not granted to individuals domiciled in, or companies with a principal place of business in, the European Union. Llama 4 Scout and Maverick are natively multimodal, so EU-headquartered organizations should review this clause with counsel and check the current license text. Llama 3.x has no such restriction, and Qwen 3, Mistral and DeepSeek are unaffected.

Is DeepSeek safe to deploy on-premise?

Self-hosting DeepSeek V3 or R1 keeps all data inside your infrastructure, and the MIT license carries no usage-reporting or data-sharing obligation. The security work is the same as for any model: verify checkpoint integrity from the official repository, run your own safety and bias evaluations, and apply guardrails and logging in the serving layer. The real constraint is hardware, since the 671B models need an 8× H200 or B200 node.

How much GPU memory does a 70B model need?

About 140 GB of weights at FP16, 70 GB at FP8 and roughly 38 GB at INT4, plus 20 to 50 percent KV-cache headroom at production concurrency. In practice that means two H100 80 GB GPUs at FP16, or one H200 141 GB or one RTX PRO 6000 96 GB at FP8. Contexts above 32K and high concurrency push the KV cache higher, so size for peak rather than average load.

Should I choose a dense model or a mixture-of-experts model?

MoE models such as Qwen3-235B-A22B and DeepSeek V3 give more quality per unit of compute because only a fraction of the parameters is active per token, but every expert must sit in GPU memory. Dense models such as Llama 3.3 70B and Qwen3-32B are simpler to fine-tune and quantize. Choose MoE when memory is available and throughput matters; choose dense for simplicity.

How often should we re-evaluate our chosen model?

Run a full re-evaluation quarterly, and a targeted one whenever the vendor ships a new snapshot of your deployed model or a competitor releases a version your benchmark shows is materially better. Pin the deployed model by commit hash, keep the previous version ready for rollback, and never promote an update that has not passed the same task benchmark used for the original selection.

How Nanobase AI can help

Nanobase AI helps enterprises select, license-review and deploy open-weight LLMs end to end: building task benchmarks from your own data, running Llama, Qwen, DeepSeek, Mistral, Gemma and Phi candidates side by side on the same serving stack, sizing H100, H200, B200 or RTX PRO infrastructure, and putting the winner into production with vLLM, TensorRT-LLM or NVIDIA NIM, RAG pipelines, fine-tuning and MCP integrations into SAP, Salesforce and Microsoft 365. As a Silicon Valley company and an NVIDIA Inception Program member, we keep the evaluation harness in place after go-live so every model update is re-tested before promotion. See our solutions or book a live demo to run the benchmark on your own data.

Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.