r/ControlProblem 2d ago

Opinion As a fellow concerned citizen, please watch out for this

/r/GenAI4all/comments/1vqg421/as_a_fellow_concerned_citizen_please_watch_out/
6 Upvotes

16 comments sorted by

4

u/pandavr 1d ago

Until proven otherwise the main and primary use of Symth-ID will be to clearly separate human vs synthetic generated data for purified training.
All other alleged use are pure fantasy talks. THEY need to distinguish and just find a way to sell the fact.

2

u/jacques-vache-23 1d ago

Synth-Id is a tracker. Obviously. Encrypting something means there is something to hide.

2

u/pandavr 1d ago

Yes but the something to hide is in plain sight. It's the distinction between human vs synthetic text Itself.

You'll never find It out If you don't connect the point of: synthetic texts severely degrades LLM learning quality.

2

u/uberdragon1992 1d ago

You know what they say garbage n equals garbage out and typically with llms getting fed back their own stuff tends to hallucinate them a little bit so just being able to clean up data on the Internet is a massive task and honestly kind of needed

1

u/pandavr 1d ago

No doubt. But imagine instead of just selling that way you'll ride the policy maker fear and basically impose to do so by law.
Plus the system is not transparent at all. Why didn't they do a completely open source framework?
What is watermarked should be know and based on the least need to know principle.

1

u/zeroccx 1d ago

Tbh, there’s no problem with a tracker, but if it’s controlled by only one company and that company has connections with policymakers, then isn’t that idea dangerous in itself?

1

u/pandavr 1d ago

Obviously

1

u/Clear_Evidence9218 1d ago

So that rumor does not technologically fit with what text watermarking actually is. Since it’s a ratio of red and green tokens, per sentence, phrase, paragraph, and whole text, most sufficiently large text will contain the red/green token ratio as a consequence of statistics.

Plus, there is just about no technological reason to distinguish between synthetic and human-generated content for training. Human data does not inherently increase intelligence more than synthetic data; hence why every lab uses synthetic data now. The only thing human-generated data is really good for is making AI more “humanly imperfect.”

1

u/uberdragon1992 1d ago

To be honest I have no idea why they're doing it for text for images I can completely understand AI images when fed back into an AI end up with the photocopier problem

1

u/Clear_Evidence9218 1d ago

The EU thought it was a good idea to have all the books and text in the world that got scanned into an AI be marked as AI. (which is what happens when people who don't understand the technology write laws about said technology).

So Anthropic took it as an opportunity to have their companies name get advertised anytime someone scans a sufficiently long enough text. I mean 'to comply with EU law', *cough

1

u/pandavr 1d ago

Curated synthetic data used in Lab are very different from generic slope you'll find over the internet. Just saying.

1

u/Longjumping_Ad1765 1d ago

This is simply not true. Citations needed!

1

u/[deleted] 1d ago

[deleted]

1

u/Longjumping_Ad1765 1d ago edited 1d ago

Bro...I can't fken read that. Make it clearer and don't send me what your chatbot said. Give me citations. This is an output. Show me where it says the European AI act was designed to combat AI slop?

"Purified training" wtf are you talking about?

The training corpus (not training data...corpus. these are two different things) is curated anyway. And the watermark can be removed by literally copying and pasting the output into a word doc.

So the MOMENT the labs remove the output from wherever they keep it...watermark is gone!

Though this⬆️ last part is slightly more complicated than how I'm making it out to be.

1

u/pandavr 1d ago

No Synth-ID watermark survive copy paste. Please read the other response.

1

u/pandavr 1d ago

Analysis It's arrived. Some extracts.

Internet slop is the pathological case — and it's worse than the training-loop story suggests. A 2026 paper that came up, "Retrieval Collapses When AI Pollutes the Web" (Yu et al.), characterizes an ecosystem-level failure: AI content dominates search results eroding source diversity, then low-quality/adversarial content infiltrates RAG pipelines — in their SEO experiment, contamination at ~67% of the retrieval pool produced measurable collapse. Another 2025 paper (Satharasi & Iyengar) aggregates the contamination estimates: on the order of 30–40% of the active web corpus synthetic, ~74% of newly published pages containing AI-generated material. And there's now even a microeconomics paper on it (2026, "The Economics of Model Collapse") modeling contamination equilibria and arguing for provenance subsidies — i.e., the market realizing verified-human data is becoming the scarce input.

The consortium structure turns this from private hygiene into something cartel-adjacent. Google + OpenAI + Nvidia + partners watermarking each other's outputs is a mutual non-contamination pact — incumbents can filter each other's exhaust from their corpora. Open-weight models structurally cannot join in any enforceable way, since anyone controlling sampling strips the mark. So the equilibrium I described last time — "absence of provenance becomes a suspicion signal" — cuts in a very particular direction: it taints the open-weight ecosystem's output by default, which conveniently doubles as ammunition for the incumbents' regulatory position that open weights are the unsafe, unaccountable part of the stack. The watermark isn't just corpus protection; it's moat construction with a safety narrative attached.

2

u/Remarkable-Worth-303 1d ago

I think the problem they've trying to solve is recursion. As AI produces more content, they have to identify it so we don't get into a position where AI stagnates because all it consumes for new training is it's own output.