We are starting to talk about artificial intelligence like a race. And the winner of that race will own the future. Which laboratory has the smartest model? Who leads on coding? How much better is the newest release at mathematics? How long before the next company overtakes it? Is this benchmark played out? Should I dismiss the current model as “benchmaxxed”?
I follow that race too. The frontier matters. Especially for my world. A model that can solve a previously intractable problem expands what we can imagine doing with computers. But there is another race, and its results may reach many more people, and change more companies: the race to make useful intelligence cheap enough, small enough, and practical enough to run wherever we need it.
That race has several engines. Models are getting better. Quantization stores them more efficiently. Distillation teaches smaller models some of what larger models can do. Hardware delivers more for the money, while better software makes more of the hardware already installed. I think this hints at intelligence being everywhere.
And I mean everywhere. On a laptop. On a desktop beside a researcher. Inside a small business. On the devices and software that surround us. Intelligence chips become like the 386 chip. Everywhere. I still remember being shocked one day to see a 386 chip inside a fridge I was fixing. I’m old enough to remember the humble 286 and the majesty of the turbo button. Yes you don’t need a 386 chip to control a fridge, but the chip is cheap enough and the development is faster, and one chip replaces a lot of control circuits leading to lower cost. This was driven by the cost to make the hardware falling. But what if the software costs less and gains in capability?
These changes mean that the computer you own can gain capabilities without being replaced. A new download can give the same machine a substantially better assistant. Will we get excited for iOS releases again?
The frontier shows us what is possible. What people can run at home tells us what is becoming ordinary.
That is what I mean by watching the tail: following useful capability as it reaches progressively smaller budgets and more widely available machines. I’m interested in a very simple question: what can a person do with the memory, money, electricity, and patience they actually have? When is a cheap and local model good enough for all of someone’s daily needs?
Yesterday’s frontier is ringing your doorbell.
There is already evidence that this migration happens on a surprisingly short timescale.
Epoch AI’s 2025 analysis estimates that leading open models small enough for a single high-end consumer GPU trail frontier models by roughly six to twelve months on the benchmarks it examined. Its estimates range from 6.3 months on the Artificial Analysis Intelligence Index to 12.4 months on LM Arena. These are fitted historical gaps, but you can see the trend. Epoch AI’s consumer-hardware analysis.
There are assumptions in this work. Epoch estimates what fits using four-bit weights, an 8,000-token context, and allowances for working memory. I think this is a good reflection of what resources many people can actually access. It combines benchmark results from several sources with estimates of which models fit on consumer hardware. And we know benchmark optimization can make the apparent gap smaller than the gap in everyday usefulness.
Still, this is a remarkable shift in perspective. A capability that attracts enormous attention when it first appears can become something an individual downloads within roughly a year. This might have been a better way to describe why we should be concerned and have a discussion about the rate of development. These local models can have a lot of guardrails obliterated. What would it mean when everyone has access to x level capability with the potential for reduced alignment?
Another historical comparison points in the same direction. Stanford’s 2025 AI Index reports that the smallest model exceeding 60% on MMLU shrank from the 540-billion-parameter PaLM in 2022 to the 3.8-billion-parameter Phi-3-mini in 2024. Now that’s a particular benchmark threshold, not showing the models have the same capabilities across all areas, but the reduction is striking. Stanford AI Index.
We should follow both curves: how far the best model advances, and how quickly a useful level of performance becomes accessible.
Quantization changes which machine can run the model
A language model contains billions of numerical parameters, usually called weights. Storing those numbers takes memory. Moving them through the computer takes bandwidth. Quantization represents those numbers with fewer bits. At the simplest level, moving from sixteen bits to four cuts the raw weight storage to a quarter. Real formats have additional overhead and there is a reason we need higher precision in some parts of a model.
Same structure, fewer shades.
The difficult engineering question is how to make that reduction without damaging the behavior we care about. Methods such as activation-aware weight quantization, or AWQ, use information about how the model behaves to protect sensitive parts of the computation. Compression works better when it accounts for which numerical errors matter. AWQ research paper.
We can see this in work from Jemin Lee and colleagues. Their evaluation reports Llama 3.1 70B at 140 GB of model storage and an OpenLLM Leaderboard v2 average of 40.64 in sixteen-bit precision. A four-bit AWQ version uses 35 GB and scores 40.39: one quarter of the reported storage, with a decline of 0.25 points. Lee et al., Table 2.
Less storage, similar measured performance. The chart selects the best reported result at each precision tier for each Llama 3.1 model. The horizontal axis is logarithmic: every step doubles storage. These are benchmark scores and reported model storage, rather than percentages of intelligence or total operating memory. Source: Lee et al.
This practical change is a hardware threshold. A model that previously needed much more memory can become a candidate for a smaller system. That changes who can experiment with it and which projects can afford it. Solar and batteries with local inference is as close to free tokens as you can get after depreciating the initial investment.
Other models are being released closer to the capabilities of home computers. Google’s April 2025 Gemma 3 QAT release reported that quantization-aware training reduced the memory needed for the 27B model’s weights from 54 GB to 14.1 GB, and the 12B model’s weights from 24 GB to 6.6 GB. Google explicitly described running the larger version on a 24 GB RTX 3090. Conversation memory still has to fit alongside the weights. Google’s Gemma 3 QAT announcement.
Quantization-aware training lets the model adapt to reduced numerical precision during training. Making the model easier to deploy is an explicit goal in its development. This can include additional quantization-aware training after the original model has been trained.
This is why the tail deserves attention. An improvement in how numbers are stored can change the accessible audience for a model as much as an improvement in its benchmark score.
Distillation gives the smaller model a better teacher
Distillation provides another route down the hardware ladder.
Instead of representing the same model with fewer bits, it trains a smaller model to reproduce useful behavior learned from a larger teacher. The teacher can provide predictions or examples of successful responses. The student learns from those examples without needing the teacher’s full architecture. The general approach is old, predating today’s language models. Hinton, Vinyals, and Dean on knowledge distillation.
DeepSeek made this especially visible with R1. Alongside its January 2025 release, it published six distilled models ranging from 1.5 billion to 70 billion parameters. Smaller descendants were part of the release itself. DeepSeek’s R1 announcement.
The team’s published results put its 32B Qwen-based distilled model at 72.6 on AIME 2024 and 57.2 on LiveCodeBench, compared with 63.6 and 53.8 for o1-mini in the same comparison. Now these are developer-reported results for specific mathematics and coding evaluations, so it’s not showing it’s better across all types of work. But they do show that a much smaller downloadable model can acquire substantial reasoning capability through training on a larger model’s outputs. DeepSeek’s evaluation tables.
Clearly distillation and quantization can be combined. Simply train a more capable small model; then reduce the memory required to run it. A 32-billion-parameter model represented at four bits would need about 16 GB for the raw weights alone, before format overhead and working memory. And hopefully it’s as good on the tasks you care about.
The larger implication is that improvements at the frontier can help create better teachers. Some of their behavior can then become training material for models designed for smaller machines. This is why we have the debate about distillation of frontier lab model outputs by other companies.
So you can see how the frontier can help improve the tail even while continuing to move ahead.
GLM-5.3-Flash shows how far the upper end has moved
GLM-5.3-Flash belongs at the expensive end of this story, but it’s a good example.
Z.AI reports approximately 320 billion total parameters, with 18 billion activated per token. Its mixture-of-experts design selects parts of the model for each step. And it has a combination of sparse and linear attention helping reduce the cost of handling long context. These are clever architectural approaches to reducing computation and memory pressure. Z.AI’s model documentation.
The distinction between total and active parameters is essential. I often see folks thinking this has reduced the size of the model. Think of a library: only a few books may be open on the desk, but the collection still needs shelves. Sparse computation does not make the need for the rest of the weights disappear. We still need to store them. Interestingly MOE and KV Cache to Cache communication might make for a new type of swarm in a model. I’ll post some results about this later.
That is where compression enters. The weights are available in a broad range of compressed versions. Official model card, Unsloth’s quantized releases.
Selected file sizes from Unsloth’s repository
The smallest listed version uses about 85.5% less storage than the BF16 reference. That brings a very large model into the range someone might investigate for a powerful personal workstation.
But a 93 GB file remains enormous for an ordinary laptop. A 120 GB file also leaves little nominal headroom on a 128 GB machine. Loading the weights is only the start: we need to add the operating system, conversation cache, and other tools, and the final system must be fast enough to use. No one would use a token/minute for example. But 13 tokens/second? With a bit of the aforementioned patience it might be very useful.
You should not think about how much intelligence is preserved from looking at the compressed size. One way to evaluate these quantized models is to compare their next-token choices with those of the reference model. This is a measure of fidelity; it won’t tell you whether a coding agent will finish a difficult job correctly. A long-context study found that four-bit degradation varied sharply by model, method, and task, with severe losses in some settings. Mekala et al..
GLM-5.3-Flash is important because a model of this scale is available for people to adapt to machines they control. Now it won’t be around forever and will be displaced, but I’m confident the process and market forces that we have discussed mean that this local capability will increase over time.
Better hardware meets better software
We need to be careful when we talk about hardware. The trend I want to discuss is how much useful computation a budget buys. The cost of the machine does not need to reduce to enable that, and with the projected RAM shortages, RAM prices may keep rising into 2027.
If you look at Epoch AI’s August 2026 analysis you see the performance per dollar of AI chips purchased from 2023 through 2025 improved by an average of 49% a year. (this is theoretical chip performance per dollar) So evidence of improving hardware economics, which is relevant for home-computer prices or local-model speed, because this used to be a leading indicator for improvements in the capability and economics of consumer grade hardware. Although perhaps the data centre demand is so high this might be more lagged than usual. Epoch AI’s hardware analysis.
Consumer hardware is responding to the need to hold larger models. AMD’s Ryzen AI Max+ 395 systems have configurations with up to 128 GB of unified memory, accessible to the CPU and GPU. Apple’s March 2025 Mac Studio announcement offered up to 512 GB in its M3 Ultra configurations, and other systems are appearing all the time. The demand is there, and interestingly talk of ‘pacing the frontier’ has made the more 2nd amendment/fiercely independent friends I know start to seriously build their own AI rigs. “You’ll take my model out of my cold dead hands”. Got to have AI in the prepper shelter. These folks are very into alignment removal on principle.
Memory capacity alone does not determine performance. Memory bandwidth, software support, and the model’s architecture affect whether the machine feels responsive. But more usable memory expands the set of models that can be considered at all.
Meanwhile, inference software makes existing machines more useful and allows non-experts to run this. One example is the llama.cpp project, which supports compressed models across CPUs and several GPU backends, including combinations of CPU and GPU execution. It is easy to get a working local application. Then, use another AI model or harness to optimize serving, or get the local model to help optimize its own serving! Reminds me of ye old days tuning BLAS and playing with ATLAS. llama.cpp project.
This is what makes the trajectory so interesting. A person can benefit from a better model, a better compression method, or better execution software without buying another computer. Someone buying a machine later can benefit from those improvements and better hardware together.
Progress reaches the home through several routes at once.
The release is becoming the start of a rapid adaptation process
Open weights allow other people to work on deployment immediately. They can create compressed versions, improve hardware support, fix conversion problems, and package the model for local applications.
GLM-5.3-Flash’s release-period histories show original-model updates and multiple quantized uploads arriving within days of one another. The histories also record subsequent fixes ;). Z.AI’s release history, Unsloth’s conversion history.
That gives us two different clocks to watch. One measures how long a capability takes to move from a frontier system into a model small enough for consumer hardware. The other measures how long a newly released model takes to acquire useful local versions. The first can be measured in months; the second can sometimes be measured in days, or disappear when smaller versions ship at launch.
The evidence supports fast diffusion. We can’t say that every capability reaches every hardware tier faster with each successive generation, but a stable gap would still mean continual improvement at home as the frontier advances.
What I would like to see beside every major model announcement is a simple record: what works on 16 GB, 32 GB, 64 GB, and 128 GB; how quickly it runs; how reliably it completes real tasks; and when those usable versions became available.
That would make the tail immediately visible.
Intelligence becomes part of the surroundings
The consequence I expect is a world in which useful machine intelligence is increasingly available close to the work.
A writer could keep an assistant beside years of notes. A small business could search and organize its records locally. A researcher could connect a model to papers, experimental logs, and analysis tools. A household could use software that helps make sense of personal documents without sending every question and intermediate result to a remote service. Maybe there will be an LLM in my fridge now, and LLM+chip is the new ‘transistor’ (which is a bit of a lame phrase, but you get what I mean).
The economics change when someone can keep a capable model running on equipment they control. Experimentation no longer necessarily carries a separate provider charge for each exchange. The costs become hardware, electricity, upkeep, and the time spent making the system useful. For occasional work, a cloud service may remain cheaper. For repeated local work where the model has sufficient ability, ownership becomes an increasingly credible option.
Local operation also allows more control over where data goes and which model version stays in use. That may really benefit security in control systems.
I expect many people and organizations to combine the two approaches. A local model will handle work it does reliably; a stronger remote model will be available for tasks that need it. Good routing will require tests and checks, because a model’s own confidence is an unreliable guide. However, there are some new approaches such as Jev and new local models like Laya.
As the local system improves, more work can stay close to the user. The practical measure of progress becomes the share of useful tasks it completes successfully, within an acceptable time and cost. That changes what deserves investment and depreciation times, and we might see local systems trading off against cloud spend.
This is why I keep coming back to the tail. The benchmark winner is a snapshot of the best result available under a particular setup. The spread of capability across ordinary machines tells us something about how widely people may be able to use it.
Intelligence will be everywhere, in the practical sense that more of our tools will be able to interpret, explain, draft, search, and act. Its quality will vary. Its reach will expand.
The next frontier breakthrough is worth watching, but so are the less visible moments when a person downloads a better model, opens it on the computer they already own, and discovers that something useful has become possible.
The tail might even begin to wag the dog in the future.
Sources checked September 20 2026.






