☆ Yσɠƚԋσʂ ☆

  • 13K Posts
  • 13.4K Comments
Joined 7 years ago
cake
Cake day: January 18th, 2020

help-circle















  • Looking at 2 out of 8 links provided is the definition of cherry picking. Nowhere did I say that you have to ignore anything. What I gave you is a broad example of many different studies. You zeroed in on two examples, ignored the rest, and then started making wild claims about the quality of research. If you had any genuine interest in the subject, you would’ve actually gone through all the links provided and then examined the available evidence before making further comments. But obviously you just wanted to confirm your existing biases and that’s what you proceeded to do.



















  • The process takes time because even when you’re distilling answers, you still need to actually do reinforcement training on the model. And given that Fable and GPT 5.6 just came out there simply hasn’t been much time to do that. However, models like Kimi also do better than Fable or GPT on a lot of tasks, which means it’s not just distillation but also difference in architecture. You can watch this talk from Kimi founder to see how Kimi was actually trained and why it performs well.

    It’s also absolutely hilarious that people think only Chinese companies use distillation, as if Anthropic or OpenAI are above that or something. Not to mention that they basically ignored copyrights on all the data the siphoned and are now crying that people aren’t respecting their terms of use.