That’s one benchmark that they focused on, but having double checked, you’re right, I was thinking of this one where it lags behind gpt 5.6 https://deepswe.datacurve.ai/
I’m just waiting for the closed-source AI industry to have their Black Swan moment.
Something’s going to come out of left field from the open AI community and all that investment in proprietary models and mega-compute is going to be rendered useless.
You’re typing your comments using a blob of goo with about a hundred million neurons in it that cycles under a hundred hertz and draws less than 20 watts.
I don’t think that we’ll be running packs of goo in our PCs any time soon. But I do think some entirely different way of looking at the problem will emerge that will reduce computational requirements by many orders of magnitude. And it won’t involve gigantic statistical engines trying to find the best average response to a question.
It could be TurboQuant or something like it. The biggest detraction to local LLM models is being able to close the gulf between obscenely-expensive 512GB NPUs, to house the 230GB uncompressed models (+ context), and more common 24GB GPUs. Quantized 15-18GB models are already working pretty well, but context size is still a bit of a problem.
Of course, the whole industry need to ramp up memory production and wrestle duopolies from the few that can make the raw silicon. It was pretty fucking pathetic that parts of the PC industry decided to leave these silicon processing weaknesses in various places. Large corps could have easily jumped into the industry and made bank in the long-term, but that would require not funneling into short-term quarterly profit bullshit.
It does.
Except cost. Except open source. Except the entire fucking global economic trillion dollar US AI model.
Those are very big exceptions.
(Also, I’m mostly stealing from Yσɠƚԋσʂ’s post.)
That’s one benchmark that they focused on, but having double checked, you’re right, I was thinking of this one where it lags behind gpt 5.6 https://deepswe.datacurve.ai/
I’m just waiting for the closed-source AI industry to have their Black Swan moment.
Something’s going to come out of left field from the open AI community and all that investment in proprietary models and mega-compute is going to be rendered useless.
You still need the mega-compute. Even for inference, 1.5 terabytes of RAM in modern servers isn’t cheap.
You can run Claude Sonnet equivalent quantised models locally on much less RAM.
Local LLMs are close to being viable. We’re almost at the point where they fit a Macbook.
You’re not thinking black-swan enough.
You’re typing your comments using a blob of goo with about a hundred million neurons in it that cycles under a hundred hertz and draws less than 20 watts.
I don’t think that we’ll be running packs of goo in our PCs any time soon. But I do think some entirely different way of looking at the problem will emerge that will reduce computational requirements by many orders of magnitude. And it won’t involve gigantic statistical engines trying to find the best average response to a question.
It could be TurboQuant or something like it. The biggest detraction to local LLM models is being able to close the gulf between obscenely-expensive 512GB NPUs, to house the 230GB uncompressed models (+ context), and more common 24GB GPUs. Quantized 15-18GB models are already working pretty well, but context size is still a bit of a problem.
Of course, the whole industry need to ramp up memory production and wrestle duopolies from the few that can make the raw silicon. It was pretty fucking pathetic that parts of the PC industry decided to leave these silicon processing weaknesses in various places. Large corps could have easily jumped into the industry and made bank in the long-term, but that would require not funneling into short-term quarterly profit bullshit.