

Honestly, I think the most reasonable approach is just to see what other people’s experience is like and which models are well regarded, then try them out and see which one is the best fit for what you’re doing. You might not even need the top performing one necessarily, and speed or lower resource usage might be a bigger factor.


lol might have to smuggle it directly from China 🤣


Yeah, the new policy is they’re not going to seek reunification, but I think that mostly implies that they see no reason to take over the south by force. If there was a collapse in the south, integrating it would mean removing US presence from their border. And that would be very valuable.


My prediction is that social collapse in the south will eventually lead to reunification.


I expect this is gonna be far worse because the economy was way more diversified back in enron days.


Sure, a benchmark doesn’t capture all the subtleties and different use cases, but it does give a general idea of the capabilities of a model. Obviously, you have to run the model and see if it does what you need. But the chart isn’t really about the nuance, it’s showing how drastically the efficiency of the models has improved in just a year. The fact that we can even reasonably compare a model you can run on a desktop to one that needed a data center just a year ago is phenomenal.


ah gotcha, and I’ve made the list myself apparently


I’m fairly optimistic that people will figure out how to optimize the models a lot further going forward. One obvious path is to try and separate the reasoning network from the trivia that gets baked into the model, and some work is being done in this area. If you could have a context free reasoning engine and then feed the facts it needs to know on the fly based on the context you’re running it in, then you could likely have a much smaller model that’s very capable.


Not sure what Moore’s law has to do with anything here to be honest. The models you can run locally on a consumer desktop can do real work, and their resource usage is no different from any other software like games that you’d run.


You need a GPU with around 16gb vram at a minimum to run qunatized version.


that’s the other huge advantage of open models you can run locally


The difference is that you can run Qwen completely local though.


yeah, it’s not a completely insane amount of data, and a db like postgres can do fast text search on that too with fuzzy matching


Ok, but that’s a completely nonsensical statement. If you ever used Qwen in an agentic loop, you’d know that it delivers working code, and it takes about same resources as playing a modern game, and I don’t see anybody whinging that game are too inefficient for what they deliver.


I mean baking knowledge into a model isn’t really all that useful to begin with. Just download wikipedia locally and have it access it through tool use, it’s way more efficient and more accurate. And yeah, I find Q6 tends to be the sweet spot where it’s close enough to full 16 bit in performance, but doesn’t chew up too much memory.


some turbolib made a Lemmy client that had a hardcoded blacklist of users and instances they considered to be communists


Ok that’s fair, what Iran is doing is very similar to Russia’s approach of launching combination of drones and missiles to overwhelm Ukrainian defences.
China’s CXMT just broke the global memory chip duopoly held by Samsung, SK Hynix and Micron. Their IPO caused a major global semiconductor stock rout and Korean markets halted trading because it’s basically propped up by these companies. And on top of that, there are now increasing worries about the whole AI bubble built on circular investments. If data center build out stops then the demand for chips and memory is going to collapse too.