Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I believe the argument from GitHub is that using public code as training data for machine learning falls under fair use [1]. If something is found to be fair use then (as far as I understand) copyright does not apply at all, and hence terms in the license do not make a difference. It is not clear whether this argument will stand up if tested in court (fair use is an affirmative defense, so something cannot actually definitively be said to be fair use until someone is accused of copyright violation and a judge rules that it actually is fair use).

That said, at the very least it seems like it would be rude to include code in the training data if the developer has expressly said they don't want that.

[1] From their FAQ: "Training machine learning models on publicly available data is considered fair use across the machine learning community."



I hope it'll get tested in court soon and ruled against GitHub/Microsoft - because otherwise this will mean that Copilot and other GPT-3-like models become perfect copyright laundering machines.


Testing the fair use argument for training won't necessarily answer the question about copyright laundering.

You could easily make the case that training is fair use, but that doesn't have to imply the model's output is non-infringing.

For example, it seems reasonable to train a model by feeding copyrighted texts and images, and that model could be useful for analyzing the content, finding facts, or detecting features. But we're in murky waters when the model also starts outputting the original content (be it verbatim or "derived").

Not all that different from human learning: you can study and learn from publicly available books but that doesn't grant you the right to recite their contents and claim it as your own, original work.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: