cross-posted from: https://programming.dev/post/51407459

Check what can you use and at what rate of token per seconds would it be… It has examples of many models and quantization levels. Huge resource!

  • comrade | 🇵🇸@lemmy.ml
    link
    fedilink
    arrow-up
    1
    ·
    edit-2
    2 days ago

    It recommends Qwen3.5 35B for an Apple A18 8GB… 🙃 also it doesn’t seem to understand how unified memory works