There was a fairly convincing discussion on HN about a month or so ago that argued that this is enabled by the large difference between the monthly subscription costs and the API costs to access the very same models. That the primary purpose is not distillation but pricing arbitrage. Data collection for distillation is just a bonus side effect.
Bringing the prices of monthly subscriptions much closer to API rates would kill the arbitrage business and make distillation a much more expensive (and harder to hide) activity. Raise one, lower the other - whatever.
The "distillation" argument always seemed pretty weak to me. If one argues that training an AI on copyrighted content is merely analogous to reading a book (and therefore is fully transformative), then it stands to reason that training an AI on other AI outputs would also be fully transformative.
It seems to me you can't have it both ways; either training is a violation of copyright or it isn't, and there's no consistent argument where "distillation" is a violation but other things aren't.
There was a fairly convincing discussion on HN about a month or so ago that argued that this is enabled by the large difference between the monthly subscription costs and the API costs to access the very same models. That the primary purpose is not distillation but pricing arbitrage. Data collection for distillation is just a bonus side effect.
Bringing the prices of monthly subscriptions much closer to API rates would kill the arbitrage business and make distillation a much more expensive (and harder to hide) activity. Raise one, lower the other - whatever.
The "distillation" argument always seemed pretty weak to me. If one argues that training an AI on copyrighted content is merely analogous to reading a book (and therefore is fully transformative), then it stands to reason that training an AI on other AI outputs would also be fully transformative.
It seems to me you can't have it both ways; either training is a violation of copyright or it isn't, and there's no consistent argument where "distillation" is a violation but other things aren't.
This argument only makes sense if you equate a printed book with a chatbot. They aren't the same thing, and the chatbot's outputs aren't copyrighted.
How does distillation actually work? If I wanted to distill a model and I have enough money and compute, what’s the process?
So China can't steal the data that US companies stole fair and square?
https://www.techopedia.com/trump-administration-openai-train...