Open-weight means a model's trained parameters are published for anyone to download and run, while open-source in the strict software sense also requires the training data, training code and full methodology to be public and freely licensed. Llama 4, Qwen 3, DeepSeek V3 and Gemma 3 are open-weight: you get the weights and a usage license, but Meta, Alibaba, DeepSeek and Google do not publish the full pretraining datasets or the exact recipe used to build them. A genuinely open-source model, by the stricter definition used by groups like the Open Source Initiative, would let anyone reproduce it from scratch, which almost no frontier-scale model satisfies today because of data licensing and cost. This distinction matters for procurement and compliance teams because open claims in marketing do not automatically mean auditable provenance or reproducibility. It also affects risk assessment, since open-weight models can still carry undisclosed biases or filtered training choices that cannot be verified directly. Typically, treat open-weight as the accurate technical term and confirm license terms and model card disclosures separately. Nanobase AI helps enterprise teams read past the marketing label and evaluate what a given open-weight license actually permits.

Openness is a spectrum, not a label

Vendors use the word open loosely, and that looseness creates real procurement risk when a compliance team assumes open means fully auditable. In practice, models fall along a spectrum of disclosure that runs from fully reproducible research releases to weights-only releases with usage restrictions to fully closed APIs with no downloadable artifact at all. Understanding where a specific model sits, rather than trusting the marketing label, is the first step in any procurement review.

TierWeights publicTraining data disclosedCode/recipe publicLicense restrictionsExample
Fully openYesYesYesMinimalSome academic releases, e.g. OLMo
Open-weight, permissiveYesNoPartialNone on usage scaleQwen 3, Mistral Small (Apache 2.0)
Open-weight, restrictedYesNoPartialScale or output-use limitsLlama 4 (Community License), Gemma 3
Weights-available, gatedRequires approvalNoNoAccess and usage gatedSome research-only releases
ClosedNoNoNoAPI terms onlyFrontier proprietary APIs

Where a model sits on this spectrum, not whether it is called open, determines what you can actually audit and what a legal team needs to review.

Why the distinction affects compliance work

The EU AI Act, in force since 1 August 2024 with general-purpose AI obligations applying from 2 August 2025, includes specific carve-outs and lighter documentation duties for models released under free and open-source licenses with published weights, parameters and architecture. A model that is open-weight but attaches a restrictive commercial license, or that does not disclose its training data summary, may not qualify for those lighter obligations even though it is colloquially called open. Compliance teams should verify the exact license text and published documentation against the regulation's definitions rather than assuming an open-weight model automatically clears every bar. This is a rapidly evolving area, so check the current regulatory guidance and license text before relying on a specific exemption.

Regulatory treatment depends on the precise license and disclosure, not on marketing language, so verify both separately for any compliance filing.

What you can and cannot verify without full openness

Without published training data and code, an enterprise adopting an open-weight model cannot independently confirm what the model was trained on, whether copyrighted or personal data was included, or exactly how safety alignment was performed. This is not unique to open-weight models, since proprietary APIs disclose even less, but it means due diligence has to rely on the model card, third-party red-teaming reports and your own testing rather than full reproducibility. Enterprises in regulated industries should treat this gap as a known limitation to document, not a reason to avoid open-weight models altogether, since the alternative typically offers even less transparency.

Partial disclosure is the norm even for open-weight models, so build due diligence around testing and model-card review rather than expecting full reproducibility.

A practical checklist for procurement review

  1. Confirm the exact license name and version, not just "open-weight" as a description.
  2. Check whether training data provenance is disclosed at even a summary level in the model card.
  3. Verify if usage restrictions exist above a scale threshold, such as Llama's monthly active user clause.
  4. Determine whether the model qualifies for any relevant regulatory carve-out, and get that confirmed against current legal text.
  5. Document what cannot be verified (training data, exact alignment process) as a known and accepted risk.

Run this checklist for every open-weight candidate before procurement sign-off, since two models both called open-weight can land in very different places on the spectrum.

Frequently asked questions

Are Llama, Qwen and DeepSeek open-source by the strict definition?

No, by the Open Source Initiative's traditional software definition, none of them are, since none publish full training data and exact reproduction code. They are accurately described as open-weight: the parameters are downloadable and usable under a stated license, but the model cannot be fully reproduced from scratch by an outside party.

Does open-weight mean the model has no safety filtering?

No. Open-weight only describes how the parameters are distributed. Publishers like Meta, Alibaba and Google apply their own safety alignment during training regardless of whether the resulting weights are later published openly.

Can an open-weight model still have hidden restrictions?

Yes. Always read the full license text, including any acceptable use policy referenced separately from the main license, since restrictions on specific use cases or redistribution can exist even under a broadly permissive license.

How Nanobase AI helps

Nanobase AI, an enterprise AI engineering company and NVIDIA Inception Program member, reviews exactly this kind of licensing and disclosure gap before recommending a model for regulated deployments. We map each candidate against your compliance requirements and document what can and cannot be verified, alongside our EU AI Act and GDPR compliance checklist. See how this fits into our solutions.

Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.