1 comments

  • floathub an hour ago

    The thesis being that if you can pack multiple models onto a single GPU, you can actually increase processing capacity with a fixed amount of hardware.