BlogAll opinions are my own
Back to all posts

Local Models Are Catching Up Faster Than Most People Realize

For a long time, the assumption around advanced AI was simple: if you wanted real capability, you had to rent it from a major lab.

That assumption is starting to crack.

With models like Qwen 3.5/3.6 27B and Qwen 3.6 35B-A3, it is now possible to run surprisingly strong models on local hardware, including something like a used RTX 3090, which can often be found for around 800€. That matters not just because it is cheaper. It matters because it changes who controls the system, how stable that system is over time, and how independent users can become from providers whose incentives will eventually be shaped by margins, pricing pressure, and platform control.

The most interesting part is not that local models have caught the absolute frontier. They have not. The best hosted models are still ahead, especially on the hardest reasoning tasks, multimodal performance, and the most polished end-to-end agentic workflows.

But the gap feels much smaller than many people assume.

In benchmark terms, models like these can feel roughly a few months behind the frontier, maybe around four months depending on what exactly you compare. And that is already a remarkable shift. A few months behind is not the same as an era behind. It means leading intelligence is diffusing into the hands of everyday users at a speed that would have seemed absurd not long ago.

The hardware story has changed

A used 3090 for around 800€ is especially interesting because it sits in a sweet spot. It is still powerful enough to run serious local models, but cheap enough to be a plausible one-time purchase for individuals, hackers, researchers, and small teams.

That becomes even clearer when you compare it to the top end of the market. An RTX 6000 Pro Blackwell can cost close to 10,000€. Of course that card is in a different league. It is a workstation product for professional budgets. But that is exactly the point.

The important comparison is not whether a 3090 beats a flagship enterprise GPU. It is whether 800€ now buys enough intelligence to matter.

Increasingly, it does.

For the price of one top-end workstation card, you can build an entire local inference machine around a used 3090 and still spend dramatically less. That changes the accessibility of serious AI. It moves capable local inference out of the lab and into the range of normal builders.

Behind the frontier, but close enough to matter

It is worth being honest here. Local models are still behind the best frontier systems. That gap is real.

The frontier still leads in:

  • hardest reasoning problems

  • multimodal depth

  • tool reliability

  • polished long-horizon agent behavior

  • overall consistency at the top end
  • But the important thing is how narrow the gap is becoming.

    If a local model is only months behind the frontier on many benchmark-style comparisons, that is a huge deal. It means you are no longer choosing between "state of the art" and "toy." You are increasingly choosing between "best possible" and "more than good enough, but fully under your control."

    For many practical tasks, coding help, writing, summarization, private document analysis, structured extraction, note systems, internal agents, the local option has crossed an important threshold. It is not merely impressive for local. It is operationally useful.

    And once a model becomes useful enough, the surrounding advantages start to matter more.

    The license matters just as much as the weights

    When people talk about local AI, they often focus on size, VRAM, speed, and benchmark performance. All of that matters, but the license matters just as much.

    A model is not strategically useful just because you can download it. It becomes strategically useful when the license actually gives you room to build with confidence, whether as an individual, a researcher, or a company.

    That is the deeper value of strong open-weight or broadly usable local models. They do not just offer access. They offer stability.

    If your workflows depend entirely on closed APIs from labs, you are downstream of their business model. Prices can rise. Terms can change. Access can narrow. Features can be withdrawn. Entire categories of use can become less attractive to the provider once profit pressure increases.

    That is not a moral failure. It is just how centralized infrastructure behaves.

    A local model changes that relationship. Once the weights are available and the license permits your use case, you are no longer waiting for a company to decide whether your workflow remains worth serving.

    That gives you bargaining power. It gives you fallback options. And it makes the overall ecosystem healthier.

    Independence is not a slogan, it is a feature

    A strong local model gives you more than privacy.

    It gives you:

  • control over latency

  • control over deployment

  • control over data locality

  • insulation from pricing changes

  • insulation from provider churn

  • a more durable foundation for workflows that need to keep working
  • And in some cases, it gives you something even more important: the ability to run off grid.

    That point is easy to miss, but it matters. A local model running on owned hardware can operate without a live dependency on a lab, a cloud provider, a billing account, or even an internet connection. For remote environments, unstable infrastructure, field deployments, privacy-critical workflows, or resilience-minded systems, that is not just a convenience. It is a fundamentally different reliability model.

    Frontier APIs are powerful, but they are rented intelligence. Local models can become owned infrastructure.

    That difference matters more than many benchmark charts can capture.

    A more stable AI future

    The real significance of local models like Qwen 3.5/3.6 27B and Qwen 3.6 35B-A3 is not that they replace the frontier outright. Frontier APIs will continue to matter, and for the highest-end tasks they will continue to lead.

    The significance is that they create an alternative center of gravity.

    That is good for users, because it reduces total dependence.
    It is good for builders, because it creates fallback infrastructure.
    And it is good for the market, because it forces hosted providers to compete on real value rather than pure lock-in.

    If this trend continues, the future of AI looks healthier than the most centralized version people sometimes imagine. Not one monolithic dependency, but a layered stack:

  • frontier APIs when you need the absolute best

  • local models for private, stable, controllable daily work

  • hybrid systems that use both intentionally
  • That is a much stronger position for users to be in.

    Conclusion

    The most exciting thing about local models right now is not that they are perfect. It is that they are getting close enough, fast enough, and cheaply enough to change the balance of power.

    A used 3090 for 800€ will not beat a 10,000€ RTX 6000 Pro Blackwell. And local Qwen models will not beat every frontier model on every benchmark.

    But that is not the real story.

    The real story is that serious intelligence is becoming available on hardware normal people can actually buy, under conditions they can actually control, for workflows that can keep running even if the market shifts around them.

    That is not just a technical milestone.

    It is a stability milestone.