Bumblebee Technology → Notes

Running AI on your own hardware: Ollama, Qwen, and when it's worth it

Every so often a client asks whether they can run this stuff themselves, on their own machine, with nothing going out to the internet. The answer is yes, it is genuinely good now, and for a specific set of situations it is the right call. Here is the honest version.

Two things made this practical. The first is Ollama, which turned running a model locally from a Python-and-CUDA adventure into roughly the experience of installing any other application. The second is that open-weight models got dramatically better: something in the 8-to-30-billion-parameter range now does work that needed a far larger model two years ago.

The models we reach for most are the Qwen family, which are open-weight, strong at both general work and code, and available in a range of sizes that map neatly onto hardware you might actually own.

What hardware you need

The constraint is video memory (VRAM) on the graphics card, not general RAM or CPU. Rough guide, using the standard four-bit quantisation that most people run — a compression that trades a little quality for a lot of memory:

You can also run on CPU with no GPU at all — it works, it's just slow enough that nobody enjoys it. Treat that as a way to try the idea, not a way to deploy it.

In practice, for an organization, this means one modest workstation or a small server with a consumer graphics card in it. Not a data centre. The current capable Qwen releases offer both a dense model around 27B and a mixture-of-experts variant around 35B with a very large context window, which covers most of what a small organization would want to do locally.

The three situations where local wins

1. The data genuinely cannot leave

Not "we'd prefer it didn't" — cannot. Health records under provincial rules, certain funder agreements, legal work with confidentiality undertakings, defence-adjacent manufacturing. If a contract or a statute says the data stays on your premises, local inference is the only honest way to use this technology on it, and it works.

2. High volume, simple task

Classifying every inbound email, extracting three fields from ten thousand documents, tagging a decade of records. Per-token cloud pricing is very cheap until you multiply it by a large number. A local model doing an unglamorous job at volume has effectively zero marginal cost once the hardware is bought, and the tasks that suit high volume are exactly the ones small models handle well.

3. It has to work when the internet doesn't

A plant floor, a remote site, a vehicle, a building whose connection is one backhoe away from an outage. If the workflow has to keep running regardless, local is not a preference but a requirement.

What you give up

This is the part vendors of local AI tend to skip.

Capability. A frontier hosted model is meaningfully better than anything you'll run on one graphics card — at hard reasoning, at long multi-step work, at not confidently making things up. For the easy 80% of tasks the gap is small. For the hard 20% it is the whole ballgame.

Someone else's problem becomes yours. Patching, model updates, the card that dies, capacity when six people want it at once, the backup story. Cloud pricing includes an operations team. Local doesn't.

Ecosystem. The polished integrations — the Excel add-in from last month's post, the mature tooling, the documentation — are built for hosted models. Locally you assemble more yourself.

A hidden cost that catches people out: the moment a local model is genuinely useful, demand grows, and one card serving one person at a time becomes a queue. Plan for that before you promise it to a department.

The pattern that actually works

The best deployments we've built aren't purely local or purely cloud. They're split by sensitivity:

The local model handles the bulk work on sensitive material — reading the case files, extracting the fields, classifying the intake. A hosted model handles the hard reasoning on material that is safe to send out, or on data the local step has already stripped of identifiers. The routing rule is written down and enforced in code, not left to whoever is using it.

That gets you most of the privacy benefit and most of the capability. It is also more work to build than either extreme, which is the honest trade.

If you want to try it this week

Install Ollama on a machine with a decent graphics card. Pull a Qwen model sized to your VRAM per the guide above. Ask it to do something real from your actual work — not a riddle, an actual task you did last week — and form your own view. It takes an afternoon and it is the fastest way to calibrate your expectations, in both directions.

One caution on model choice: the very newest flagship releases from the same families are sometimes API-only with no published weights. If a model name is generating excitement, check that the weights are actually available before you plan around running it yourself.

Deciding between local and cloud?

It usually comes down to what your obligations actually say, which is a shorter conversation than people expect. Tell us what data is involved and we'll tell you which way we'd go, and why.

Request help

Next month: where your data actually goes — consumer plans, business plans, APIs, and what PIPEDA requires of you.