A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]
damaru2 · 421 points · 136 comments · ontem · Open original
Comments
5 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
PBpbasistaontem
Tangential:
I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.
This partial prompt data might potentially be used to "pre-warm" some kind of cache.
But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.
KDkdaniel_03ontem
It's the same lesson as the Navier-Stokes credit fight earlier this month. Buckmaster and Alpoge had their unpublished drafts in private Codex sessions and OpenAI says nobody saw them but admits de-identified product data may have improved its models. There it's training data, here it's ad trackers. Either way, prompts and results that should stay private don't.
Thats why even though open models aren't perfect it has to win. You can skip the app and run the model yourself.
POpostalcoderontem
My least favorite trend I’ve noticed with so many AI chat services is they seem to equate a UUID in the url with privacy.
Perplexity does this. Visiting a past perplexity search url exposes your full conversation.
DEdelis-thumbs-7eontem
In an old Simpsons episode Lisa gets to visit the Teachers room, where all the staff are making fun of the children. Groundskeeper Willie is pantomiming Milhouse “Oh I am Milhouse, I tell all my secrets to Willie since I have no friends!” and the teachers laugh. Later something embarrassing happens to Milhouse and he immediately runs away crying “I have to tell this to Willie!”.
We have all become Milhouse now.
J4j4k0bfrontem
This is a bit surprising to me, considering how much AI companies love to hoard data. Especially since some of these ad companies are direct competitors!
My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.
Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).
Comments
5 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
Tangential: I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it. This partial prompt data might potentially be used to "pre-warm" some kind of cache. But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.
It's the same lesson as the Navier-Stokes credit fight earlier this month. Buckmaster and Alpoge had their unpublished drafts in private Codex sessions and OpenAI says nobody saw them but admits de-identified product data may have improved its models. There it's training data, here it's ad trackers. Either way, prompts and results that should stay private don't. Thats why even though open models aren't perfect it has to win. You can skip the app and run the model yourself.
My least favorite trend I’ve noticed with so many AI chat services is they seem to equate a UUID in the url with privacy. Perplexity does this. Visiting a past perplexity search url exposes your full conversation.
In an old Simpsons episode Lisa gets to visit the Teachers room, where all the staff are making fun of the children. Groundskeeper Willie is pantomiming Milhouse “Oh I am Milhouse, I tell all my secrets to Willie since I have no friends!” and the teachers laugh. Later something embarrassing happens to Milhouse and he immediately runs away crying “I have to tell this to Willie!”. We have all become Milhouse now.
This is a bit surprising to me, considering how much AI companies love to hoard data. Especially since some of these ad companies are direct competitors! My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest. Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).