I want an AI that keeps working
Why dependable access, data privacy and control make locally hosted, open-weight AI worth considering for Canadian businesses.
How much of your AI infrastructure are you comfortable renting from someone else?
I see a bright future for open-weight AI for small and medium-sized enterprises (SMEs). Locally deployed open-weight models can offer client data privacy as well as mitigate sovereignty threats of cloud model providers from both the US and China. This is not for every SME. There is incredible utility in a per-seat subscription that is good enough for most tasks. You can’t beat the convenience to bundle backups, connectors, logs included in that subscription.
But I have been thinking about the sovereignty angle since 2025 when it became very real for Canadians to listen to week after week of threats on our sovereignty from the United States. This has me reassessing professional and personal use of brands like Microsoft, Apple, Google, OpenAI, Anthropic and many others. A salient example this year was when Mythos/Fable was slated for general release this summer, only for the USG to order suspension of access to Fable 5 + Mythos 5 by any foreign national. Anthropic pulled access for everyone to these models as a response to the order. This had me reconsidering Anthropic for business use in a Canadian context. I want dependability in a vendor for a business relationship. Especially if that vendor’s pitch is top of the line intelligence models; they should be accessible for day to day operations. In the same situation I would assume OpenAI would also fold to the USG but I can only guess this on vibes.
The Canadian business operator perspective has changed. For anyone running a business where access to frontier models is important; the question is now “what is the expected availability of this subscription as a daily driver?” I am not going to be happy if I am in the middle of a system conversion and I get shut off access to the best model a provider has to offer. I also care about where sensitive client information is being stored. If it is stored internationally what happens to that data, how secure is it and what happens if it becomes a bargaining chip between countries in the future?
There are two choices here: whether you can obtain the model’s weights, the learned parameters needed to run it, and where the model runs. Open-weight models let you download those weights, subject to their licence. Local hosting means running inference on hardware you control; cloud hosting means running it on a provider’s infrastructure. For illustration:
| Local / on premises | Cloud / provider hosted | |
|---|---|---|
| Open weights | You run downloaded weights on your own hardware. You manage the equipment and can keep a working copy of the model. | A provider runs an open-weight model for you. You depend on that service for access, but can arrange another deployment of the model. |
| Closed weights | Where a vendor offers an on-premises deployment, inference runs in your environment under its software and licensing restrictions. | You access the model through the provider’s app or API. The provider controls access to both the model and the service. |
If you are interested in control, a combination to familiarize with is open weights, hosted locally. A working local model keeps working when a hosted service withdraws access. Local models don’t require internet to keep answering once running, only to serve the model to your team. The responsibility you inherit with local is the hardware, maintenance and security requirements. What is tricky is AI companies often force your hand on this choice, so if you aren’t paying attention you may pick the wrong combination for your business simply because it was what is offered. This is risk mitigation and future strategy as we grow used to AI tools. Don’t ignore it. Access to weights affects your ability to keep or move a model. Hosting affects where prompts are processed, which infrastructure you depend on, and whose electricity and cooling you use.
AI is a political topic, much more so now than a year ago. There are concerns about the use of water to cool inference and training equipment. This suggests ChatGPT needs 500 mL for a conversation of 10-50 questions and answers, depending on where and when it runs. This is also a bit dated and is often shared without the nuance of the training footprint not being repeated for each individual query. A large part of AI energy usage is not serving the AI to a user but in the training of the model (calculations) to build the model. This is a sunk environmental cost with these tools. Local or in house/on premises hardware can run AI. This bypasses sending sensitive data outside of your organization. It also lets you see exactly how much energy and compute requirements go into serving AI to your selected users. This doesn’t absolve AI of using water or energy but it does let you see the energy use and thereby control the environmental footprint.
AI’s impact on the physical world is also visible with Datacentres which are in my opinion ugly buildings that sometimes bypass municipal and small town governance. The media landscape around AI is also steeped in negativity that is both overselling the capabilities and job displacement potential and also calling it a bubble. As a heavy user of the agentic coding tools this is mostly noise to me, simply because I find the tools are generally useful. Even if advancement stopped today the tools would be extremely useful for years to come.
The performance of local AI depends on how fancy your hardware is (a laptop is not going to give you the same speed as a Claude subscription), but for businesses the amount of money you need to spend on hardware to get “good enough” inference may surprise you. The energy footprint is also comparable with other networking infrastructure like servers, cooling, batteries and electricity you use in operations and again; keeps data local. For an SME with 50-200 occasional users, an illustrative hardware budget is roughly US$4,700-$18,800 for one to four DGX Sparks (for simplicity, this isn’t the only local hardware option), assuming only a few people per machine need responses at once.1 By comparison, ChatGPT Business and Claude for Teams both cost US$20 per person per month with annual billing: US$12,000 a year for 50 seats or US$48,000 for 200. Hosted subscriptions of course provide the models, apps and managed infrastructure included with that price tag.
Open-weight AI thankfully, is not far behind the frontier. Prior to the release of Astra aka GPT-6 (a substantial leap in model capabilities) open-weight models I would estimate were on par with Anthropic and OpenAI top models from four months ago (May 2026). Those were already very powerful models. What could be done with frontier models four months ago is probably sufficient for most of your business use-cases in your day-to-day operations. I know we always want the best and brightest but you should be wary of any claims that you need the best models right now. Most knowledge work tasks are not that demanding of cutting edge intelligence. There is also a right tool for the job effect. Models all excel and lag relative to the median model at certain tasks. A lightweight model may be perfectly suited for classifying emails or reading a database of records. These don’t have to be one-size-fits-all tools, there are performance gains in matching tools to tasks.
The Canadian AI coverage I encounter often feels dominated by risks. Most articles I read focus on the externalities and division AI is supposed to bring. Common points are AI will flood the web with misinformation, natural resources and water especially is being diverted to train and serve AI, energy prices are rising due to demand brought by datacentres. The tone can vary quite a bit by topics related to AI. Public opinion also varies sharply between countries. The 2025 and 2026 Stanford AI Index reports show the same broad pattern in Ipsos polling from 2024 and 2025: respondents in China and Indonesia were much more likely to see AI’s benefits outweighing its drawbacks than respondents in Canada, the United States, Japan and Great Britain. Canada remained at 40% and Japan at 48%, while the United States rose from 39% to 42% and Great Britain fell from 46% to 43%.2
Source: Ipsos, 2024–2025, via Stanford AI Index 2025, fig. 8.1.2 and Stanford AI Index 2026, fig. 9.1.2. Survey years shown; six selected countries. Ipsos supplied China’s 2025 data separately to Stanford.
What’s shocking in the survey data is the gulf in sentiment between China and North America. The share of Chinese respondents who say AI provides more benefits than drawbacks was 87% according to Ipsos in 2025. Contrast the business approach of American AI companies vs Chinese AI companies. Deepseek, Moonshot, Z.ai have been pricing their models much more as an abundant commodity for high repetition and throughput tasks. The American frontrunners Anthropic and OpenAI have been focusing mainly on best-in-class intelligence at a premium (10-25x the price of popular Chinese models on a given week). Top competing Chinese AI models are much more commonly released with open-weights meaning the model can be downloaded, optionally tweaked and run on local hardware - for free. I wonder if the polling divide has to do with this openness of the product.
I have been a bit of a fanboy for OpenAI for the past few years but I started to question that once I tried Kimi K2.5 and have been testing various open weight models since. GLM-5.2 wowed me, and Deepseek V4 Flash is really impressive for routine work that needs quick answers. I now reach for these models to get a difference of opinion or a different perspective from GPT or Claude. Beyond the personality and taste differences of models, when you are building AI capabilities into a program you have to pay the market rate (API pricing) for use. For low level tasks that require some fuzzy reasoning these open weight models are excellent. I should note American models can be open-weight and hosted locally too. Google Gemma for example is a great open weight model I’ve leaned on. It’s the same open-weight, locally hosted square from above despite being from Google. Gemma has allowed me to better appreciate the power of this technology, for example when I’m stuck without internet on a rainy cottage weekend and have a slow but dependable Gemma instance for settling trivia debates.
As a business owner you should be considering who controls the AI served to your team. You need to consider local, open-weight models if you want to insulate your business from geopolitical spats and data privacy concerns. There is upfront work in picking the right hardware investment for your team’s needs but over time this can have significant cost and qualitative advantages over a per-seat subscription to OpenAI or Anthropic’s plans. For me though it is predictability. If I plan to use something daily for work, I want it to be dependable and not a future bargaining chip.
A compelling vision for how AI works for your organization and employees starts with fitting your business context with a deployment that mitigates risks unique to your situation. But control over the model isn’t the same as control over how it’s going to be used every day. The combo of local and open-weights gives you control of dependability and ownership of the model. It doesn’t answer the question your employees are still asking. Who can see my prompts, why are we changing the way we did things, and what happens to my job if this can do some of my tasks?
If the solution to those problems was as easy as a $20 per person per month subscription we’d have solved the AI rollout problem by now.
Footnotes
-
Illustrative purchase budgets September 2026. NVIDIA lists a DGX Spark with 128 GB memory and 4 TB storage at US$4,699
- 50 occasional users: one Spark, US$4,699, assuming up to four requests running at once and a queue for bursts.
- 200 occasional users: four Sparks running separate copies of the model, US$18,796, assuming up to sixteen requests running at once, distributed evenly across the machines.
For a usage reference, NVIDIA’s March 2026 benchmark, table 2 ran Qwen3 Coder Next in FP8 using vLLM on one Spark: four simultaneous requests, each with 32,000 input tokens and 1,000 output tokens, took about 91 seconds end to end, with a median 15 seconds to the first token. These scenarios extrapolate from there. Before making an investment like this you’d need testing for the team’s actual tasks to understand needs. Prices are hardware only, in US dollars, before tax, shipping, setup, networking, electricity, maintenance and staff time. ↩
-
The chart shows selected country results from Ipsos polling, not regional averages. Online samples have country-specific representativeness limits. See the chart source notes for data provenance and reproduction details. ↩