2 comments

  • nycdatasci an hour ago

    Post from Mark Chen for those not on X:

      Two things to distinguish:
      Did any human or agent look at user data as part of the Navier Stokes effort? No.
      Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.
  • mmooss an hour ago

    > de-identified

    The issue always is, was the de-identification effective?

    With a birtdate, gender, and zip code, ~85% of Americans can be uniquely identified.[0] Much data contains much more unique information than that; I imagine most data about you has identifiable fingerprints - where you go, what you bought at the grocery store, your medical conditions, movies you watch, music you listen to, entertainment choices, hobbies, etc.

    An LLM is the perfect tool to identify someone based on that data.