I don’t doubt that locally hostable models will be important for agentic tasks but I suspect they are still going to be bigger than most can comfortably host for time being. I still think there will be a place for the super large models for more complex reasoning although how much will be due to the intrinsic knowledge in the weights and how much due to the plumbing around them remains too be seen.
And if course LLM’s are not going to be the end point of the search for AGI. Whatever their architecture they will still need copious amounts of compute.














I wouldn’t say it’s totally legacy. (v)ram bandwidth does matter and while the Apple M-series chips do well with their unified memory architecture don’t forget it’s fixed because it’s part of the CPU chip.