LLMs are here to stay
The press, the machine, the railway, the telegraph are premises whose thousand-year conclusion no one has yet dared to draw. - Nietzsche
The recent Qwen3.8-27B release really made me realize how capable locally runnable LLM models have gotten, even though this particular model tends to overthink a lot. Additionally, the success of the improved DeepSeek V4 Flash and Pro versions shows that you do not need multi-trillion parameter models to do useful things. Apparently, the immense quality difference between the first and newer versions of DeepSeek V4 Flash can largely be attributed to better post-training, meaning that the base model itself has stayed the same. Part of this improvement in smaller models comes from the fact that distillation from larger models works surprisingly well. You can now run models locally matching the frontier of ~7-10 months ago with around 10k euros for DeepSeek V4 Flash and Qwen with a relatively specced-out MacBook.
Even if every single one of these AI research companies like Google, OpenAI, and Anthropic were to disappear overnight, these local models would still exist and would be capable of accelerating work. While their output in software development can be underwhelming, they can still be immensely useful for speeding up certain aspects of the software development workflow. I focus on software development partly because software engineers tend to work closely with frontier models and are therefore more exposed to their current capabilities than people in many other fields. The potential in many different fields still seems to be under-realized, and only time will tell if this technology will provide the promised capabilities in tasks with unclear constraints which are hard to verify.
While more and more intelligence can be squeezed into smaller models, this technology will become increasingly commoditized and be able to run on lower-end consumer hardware while giving increasingly better results. Even if the upper limits of intelligence that these models can achieve plateau, I believe that the cost of running these models and the cost of intelligence will ultimately continue to decline. Maybe in a few years you will have the intelligence equivalent of today's multi-trillion-parameter models running on a regular phone, who knows.
This is why the future feels so uncertain. There is this revolutionary technology that somehow produced intelligence from relatively simple principles, even though the high-level architecture of these systems can be understood in a few hours. In addition to all of this, the intelligence of this technology is increasing at a rapid pace. Humanity has unleashed this technology that cannot be removed from existence even if we wanted to. In the past, older technologies enabled further innovations and changed the values and expectations of humans. We can only imagine what kind of technological advancements increasing efficiency of work by even a factor of 1.2x could allow. While I don't necessarily believe that LLMs can achieve AGI, the efficiency improvements in different industries might eventually lead there. We have this premise that we don't know the conclusion of.