mbuchel-hn 19 hours ago

does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?

  • rhgraysonii 13 hours ago

    Yes that is exactly what this does.

    • kennywinker 10 hours ago

      Could you explain what happens when you try to shoehorn a 2.4T parameter model into a 24gb m4 mac?

jaylane 16 hours ago

tried it out but based on the model sizing result i got i got an insufficient memory error when the server started running

  • rhgraysonii 13 hours ago

    If you could post an issue if you still have the error around that would be awesome.