Disinformation Security: Protecting Your Platform from AI-Generated Content

AI-generated content is now cheap enough to weaponise. Practical defences for platforms: provenance, rate design, detection limits and moderation that scales.
Disinformation Security: Protecting Your Platform from AI-Generated Content
The cost of producing convincing fake content has collapsed. That single economic fact — not any particular model — is what breaks the assumptions behind most platform moderation. Systems designed around "a human had to bother" no longer hold, because nobody has to bother.
If your product accepts user-generated content — reviews, listings, comments, profiles, support tickets — this is now an engineering problem with your name on it.
The four attacks that matter
| Attack | What it looks like | Who it hurts |
|---|---|---|
| Synthetic reviews | Hundreds of fluent, specific, false reviews | Marketplaces, tourism, hospitality |
| Fake identities | Generated photos, coherent history, real-looking documents | Any platform with trust between users |
| Content flooding | Volume that drowns genuine contributions | Forums, comment sections, job boards |
| Impersonation | Cloned brand voice, cloned executive voice on a call | Every business with a finance team |
Why detection alone will not save you
Be blunt about this: AI-text detectors are unreliable. They produce false positives on non-native English writers — a serious fairness problem in a South African context — and they degrade every time a new model ships. Building your defence on a detector score is building on sand.
Detection is a weak signal to combine with strong ones, not a gate.
Defences that actually hold
1. Raise the cost of an account, not the cost of a post
Generated content is free; verified identity is not. Tie trust to something scarce — a verified phone number, a completed transaction, a payment method, time on the platform. A review from an account with a matching confirmed booking is worth a thousand from accounts created this morning.
2. Provenance over inspection
Instead of asking "does this look fake?", record where it came from. Signed uploads, EXIF retention for photos taken in-app, transaction linkage for reviews, C2PA content credentials where supported. Provenance survives model improvements; heuristics do not.
3. Behavioural signals
Content is easy to fake. Behaviour is harder. Watch registration bursts from one ASN, identical session timings, copy-paste input with zero editing events, and posting cadence that never sleeps. These catch campaigns even when each individual item is flawless.
4. Graduated response
Binary allow/ban is brittle. Build a ladder: shadow-rank suspicious content lower, hold it for review, require verification to publish, then remove. Most abuse dies at the first rung because the economics stop working.
5. Rate design
Limits per account are trivially bypassed. Limit per IP block, per device fingerprint, per new-account cohort and per unusual velocity. Make the tenth post of the hour cost more than the first.
The internal risk: voice and email impersonation
The attack South African businesses actually get hit with is not platform-scale — it is a cloned voice on a WhatsApp call asking finance to release a payment, or a perfectly written email from "the CEO" during month-end. The defence is procedural, not technical: a fixed out-of-band confirmation for any payment change, no exceptions for urgency. Write it down, train it, and make ignoring it a policy breach rather than a judgement call.
Frequently asked questions
Should I ban AI-generated content outright?
Unenforceable, and usually wrong. The problem is deception and volume, not the tool. Police behaviour and authenticity of claims instead.
Do AI content detectors work well enough to use?
Only as one weak input among several. Never as the sole basis for removing content or suspending an account.
What does POPIA require here?
If you process personal information to detect abuse, that processing needs a lawful basis and must be proportionate. Document what you collect, why, and how long you keep it — device fingerprints and IP logs are personal information.
The takeaway
You cannot win by getting better at spotting fakes; the other side improves faster. You win by making authenticity cheap to prove and abuse expensive to attempt.
Continue Reading
- Post-Quantum Cryptography: What Every Developer Needs to Know
- AI Chatbot Training & Models Guide
- AI for Small Business in South Africa
Need this built properly? See services or get in touch.
LET’S TALK
Have a project in mind after reading this? Send a brief and I usually reply within 24 to 48 hours.
