250 documents ingested from somewhere is enough to become part of the knowledge of a model of arbitrarily large size.
I would expect that a good idea that fits in a framework that is already being ingested would be more easily taken up than some random thing unassociated with anything else. Could that go down to a single transcript? If the model is consciously focusing on everything X related, quite possibly.
Thats not what they are asking. This paper is discussing documents in the training dataset poisoning the LLM for malicious behavior. This person are asking if anyone has deliberately put something in a private chat (presumably with retrain on my data turned off), to see if they can get it to leak across sessions from distinct users. I am positive this happens but I have not seen the proof. I also want to know the answer to this question.
* with train on my data turned ON, yes. Though OFF would of course be even more notable!
Thank you – the non-adversarial reproduction paper ( https://arxiv.org/abs/2411.10242 ) nails it – from chat, to training corpus, to subsequent model. Though in my hasty read, it is not entirely clear whether the snippets it finds are nonces, i.e. present exactly once in the internet.
I presented the research that I knew was somewhat relevant. Then made it clear that that wasn't what was being asked, and why my expectation is what it is.
The point of mathematics is not to prove results. It is to build conceptual thinking about mathematics. Important problems are important because in order to solve them we have to build concepts tying different things together.
We're not searching for answers. We're searching for insights. Trying to understand the problem causes us to draw the connections and find those insights.
AI gives us answers. But it doesn't help us build those insights. AI has a complete mastery of existing human insights. But doesn't build new ones from its own experience. In a real way, it does not find the opportunity to really learn.
So it tackles problems and either solves them or not. If solved, we now have an answer. If not, it's too hard for humans.
That's the viewpoint of everyone sensible outside of Alzheimer's research.
Those in it are still throwing billions per year at the idea.
Meanwhile, back in reality, no amyloid-beta drug has had any clinical effect in humans, other than reducing the plaques. But both the shingles and RSV vaccines are proven to reduce Alzheimer's risk.
Which did not stop the FDA from approving a useless anti-amyloid drug, leading to the resignation of several experts, one of whom called it "probably the worst drug approval decision in recent U.S. history" in his resignation letter. [0]
Bad example. You can do all of this with constructivism. Any constructable Cauchy sequence converges to a constructable member of the space.
What you get for the formalism around computable numbers is this. Every mathematical object in the theory is something that can be, at least in principle, actually written down. When we say that it exists, this existence is of the most tangible form that any mathematical thing could have.
Having constructible Cauchy sequences doesn't guarantee that we can construct unbounded operators. I'm no expert, but the little searching I've done suggests this is an open research question.
I don't see the benefit of being able to write something down "in principle." A number can only ever be computed to a finite number of digits in practice. If we're talking about finite approximations, then the standard approach using numerical solutions to the Schrödinger equation handles this just fine, no alternative mathematics needed. If we're talking about theories, then we should choose whatever abstraction is most convenient for expressing the theory.
Personally, I don't believe numbers "exist." The physical universe exists, and numbers are abstractions that we invent to describe it. In that sense, uncomputable numbers are just as "real" as computable ones.
But there are numbers in constructivism for which it is unknown whether they are zero. Some of which must remain unknown, if mathematics is consistent. This is a rather important and weird edge case.
There is no human only proof of this.
reply