• stuner@lemmy.world
    link
    fedilink
    arrow-up
    3
    ·
    5 hours ago

    I run it using LM Studio, which defaults to Q4 quantization, I think. I was able to put about 10 layers on the GPU with 64k token context. That put me at about 9.1 GB VRAM usage, leaving some room for Video playback xD