For learning it's all you need. Pick a good small model, I would recommend Qwen3.5 there are models ranging from 0.8b params (will run on a smart phone) all the way up to 397b params .. I run the 122b param model daily with 128GB vram and get 20+tk/sec. regardless of the model, you can learn all about how to host and harness a model with any size
For learning it's all you need. Pick a good small model, I would recommend Qwen3.5 there are models ranging from 0.8b params (will run on a smart phone) all the way up to 397b params .. I run the 122b param model daily with 128GB vram and get 20+tk/sec. regardless of the model, you can learn all about how to host and harness a model with any size
With this amount of VRAM, do you happen to be using a Strix Halo or a Mac?