SaaS Tech Watch
saas

Open-weight models after DeepSeek: what actually changed for enterprise AI stacks

Every figure states its provenance: measured (we ran it) · reported (vendor says) · derived (we calculated).

In January 2025, DeepSeek released R1, an open-weight reasoning model that reset industry cost and performance assumptions. It was not the first strong open-weight model. It was the first to make frontier-class reasoning available as downloadable weights under a permissive licence at a moment when API models were the default. Ten months on, the enterprise AI stack has absorbed the change. What it looks like in production is more specific than either the hype or the dismissal suggested.

This piece covers what is actually running in enterprise stacks now, the governance conversations that have formed, and when open weights beat API models in practice. We are not benchmarking models — benchmarks are volatile and vendor-contested. Deployment patterns are not.

Three deployment patterns, not two

Public discussion frames the choice as “host a model or call an API.” In production, there are three patterns, and the third is the one that has grown most this year:

PatternWhat it meansWho runs the infraLatency profile
Managed open-weightsProvider hosts the open model behind their API (e.g. hosted Llama, Mistral, Qwen variants)ProviderComparable to frontier API
Self-run, vendor hardwareWeights run on rented GPUs in a cloud or GPU-cloud accountTeam manages deploymentTunable; egress-dependent
Self-run, owned hardwareWeights run on owned or colocated GPUsTeam’s infraPredictable under owned capacity

Managed open-weights has quietly become the default first step. It preserves the API contract engineering teams know, while changing the properties underneath: the model is reproducible, fine-tunable, and cannot be silently swapped. That last property — the model is a named artefact, not a mutable endpoint — is why regulated teams are interested.

The governance questions that now get asked

When a team proposes an open-weight model, internal review reliably surfaces five questions:

  1. Licence. “Open weight” is not one licence. DeepSeek R1 ships under MIT; Llama under a community licence with commercial terms; Mistral varies by model. Legal review of the specific licence is step zero.
  2. Provenance. Who trained it, on what, and do we have recourse? For models trained in other jurisdictions the answer is usually “none.” Some organisations accept this; some do not.
  3. Evals. How does it perform on our task, our data, our failure modes? Public benchmarks are a starting signal at best. Governance teams increasingly require an internal eval suite before production.
  4. Update policy. An open-weight model does not silently update. That aids reproducibility and makes safety patches our problem.
  5. Exit path. If the model falls behind, how hard is migration? Open weights genuinely win here — switching is a deployment decision, not a data egress project.

Where open weights win in practice

The honest list of cases where self-run open weights are the better economic and architectural answer:

  • Sustained, high-volume inference where unit economics are dominated by GPU utilisation. A workload serving tens of millions of requests per day is often cheaper on owned or committed GPU capacity than per-token billing.
  • Strict data-residency or data-egress constraints. If prompts may not leave a jurisdiction or network zone, self-hosting resolves the problem architecturally rather than contractually.
  • Fine-tuning as a product requirement. Where a model must absorb domain-specific behaviour beyond what prompts can express, an open-weight base is far more flexible than an API model.
  • Reproducibility requirements. Audit and regulatory work demanding “the exact model artefact that produced this output” cannot be satisfied by providers who silently update models.

This list is shorter than conference talks suggest. Most workloads fall outside it.

Where API models still win

The unglamorous counter-list:

  • Spiky, unpredictable usage. Owning capacity that sits idle between bursts is expensive. Serverless inference from an API provider absorbs the variance.
  • Small teams without platform headcount. A well-run self-host deployment is real ops work: model registry, GPU scheduling, observability, security patching. If the team cannot staff it, API wins.
  • Frontier capability needs. The strongest end of reasoning, multimodal, and tool-use capability is still disproportionately proprietary. Teams needing that ceiling use API models until open weights catch up on that axis.
  • Fast-moving product work. Proprietary endpoints ship features (tool use, structured output, image input) at a cadence open releases follow months later.

The pragmatic answer for many buyers is a per-feature decision, not a per-company one.

Hybrid is the actual shape

Most enterprise stacks we see this year are hybrid by workload, not ideology. Embeddings and classification run open on owned capacity; long-horizon agentic tasks run on frontier APIs; domain generation runs on fine-tuned open weights. Model routing has become infrastructure, and the routing layer is now a product category.

The governance posture that makes hybrid workable is explicit workload classification: what data classes may touch external APIs, what latency budgets apply, what reproducibility requirements exist. Without it, sensitive workloads drift to external endpoints because it is convenient.

What we would tell a team planning 2026

  • Assume hybrid; design routing and classification layers deliberately, not by accident.
  • Pick at least one open-weight workload for owned or committed capacity to build operational competence. When the next frontier open model lands, teams that have run one before move faster.
  • Write down the per-workload governance answers once, in reviewable form. Individual teams repeating the analysis is the real cost of not documenting.

The bottom line

DeepSeek R1 did not make API models obsolete, and it did not make open weights a panacea. It made open weights a legitimate design choice for a specific set of workloads, at a moment when the default assumption was API-first. Ten months later, the mature position is neither “we host everything” nor “we rent everything” — it is “we choose per workload, and we know why.” That is a more complicated place to be, and a more honest one.

Related reading

from the desk ▸