• eleitl@lemmy.zip
    link
    fedilink
    arrow-up
    6
    ·
    17 hours ago

    You still need the mega-compute. Even for inference, 1.5 terabytes of RAM in modern servers isn’t cheap.

    • ℍ𝕂-𝟞𝟝@sopuli.xyz
      link
      fedilink
      English
      arrow-up
      1
      ·
      2 hours ago

      You can run Claude Sonnet equivalent quantised models locally on much less RAM.

      Local LLMs are close to being viable. We’re almost at the point where they fit a Macbook.

    • Dave.@aussie.zone
      link
      fedilink
      arrow-up
      1
      arrow-down
      1
      ·
      12 hours ago

      You’re not thinking black-swan enough.

      You’re typing your comments using a blob of goo with about a hundred million neurons in it that cycles under a hundred hertz and draws less than 20 watts.

      I don’t think that we’ll be running packs of goo in our PCs any time soon. But I do think some entirely different way of looking at the problem will emerge that will reduce computational requirements by many orders of magnitude. And it won’t involve gigantic statistical engines trying to find the best average response to a question.