TENDAI . G
Back to journal
security10 Jun 2025

Disinformation Security: Protecting Your Platform from AI-Generated Content

Tendai Gumunyu11 min read
Disinformation Security: Protecting Your Platform from AI-Generated Content

AI-generated content is now cheap enough to weaponise. Practical defences for platforms: provenance, rate design, detection limits and moderation that scales.

Disinformation Security: Protecting Your Platform from AI-Generated Content

The cost of producing convincing fake content has collapsed. That single economic fact — not any particular model — is what breaks the assumptions behind most platform moderation. Systems designed around "a human had to bother" no longer hold, because nobody has to bother.

If your product accepts user-generated content — reviews, listings, comments, profiles, support tickets — this is now an engineering problem with your name on it.

The four attacks that matter

AttackWhat it looks likeWho it hurts
Synthetic reviewsHundreds of fluent, specific, false reviewsMarketplaces, tourism, hospitality
Fake identitiesGenerated photos, coherent history, real-looking documentsAny platform with trust between users
Content floodingVolume that drowns genuine contributionsForums, comment sections, job boards
ImpersonationCloned brand voice, cloned executive voice on a callEvery business with a finance team

Why detection alone will not save you

Be blunt about this: AI-text detectors are unreliable. They produce false positives on non-native English writers — a serious fairness problem in a South African context — and they degrade every time a new model ships. Building your defence on a detector score is building on sand.

Detection is a weak signal to combine with strong ones, not a gate.

Defences that actually hold

1. Raise the cost of an account, not the cost of a post

Generated content is free; verified identity is not. Tie trust to something scarce — a verified phone number, a completed transaction, a payment method, time on the platform. A review from an account with a matching confirmed booking is worth a thousand from accounts created this morning.

2. Provenance over inspection

Instead of asking "does this look fake?", record where it came from. Signed uploads, EXIF retention for photos taken in-app, transaction linkage for reviews, C2PA content credentials where supported. Provenance survives model improvements; heuristics do not.

3. Behavioural signals

Content is easy to fake. Behaviour is harder. Watch registration bursts from one ASN, identical session timings, copy-paste input with zero editing events, and posting cadence that never sleeps. These catch campaigns even when each individual item is flawless.

4. Graduated response

Binary allow/ban is brittle. Build a ladder: shadow-rank suspicious content lower, hold it for review, require verification to publish, then remove. Most abuse dies at the first rung because the economics stop working.

5. Rate design

Limits per account are trivially bypassed. Limit per IP block, per device fingerprint, per new-account cohort and per unusual velocity. Make the tenth post of the hour cost more than the first.

The internal risk: voice and email impersonation

The attack South African businesses actually get hit with is not platform-scale — it is a cloned voice on a WhatsApp call asking finance to release a payment, or a perfectly written email from "the CEO" during month-end. The defence is procedural, not technical: a fixed out-of-band confirmation for any payment change, no exceptions for urgency. Write it down, train it, and make ignoring it a policy breach rather than a judgement call.

Frequently asked questions

Should I ban AI-generated content outright?

Unenforceable, and usually wrong. The problem is deception and volume, not the tool. Police behaviour and authenticity of claims instead.

Do AI content detectors work well enough to use?

Only as one weak input among several. Never as the sole basis for removing content or suspending an account.

What does POPIA require here?

If you process personal information to detect abuse, that processing needs a lawful basis and must be proportionate. Document what you collect, why, and how long you keep it — device fingerprints and IP logs are personal information.

The takeaway

You cannot win by getting better at spotting fakes; the other side improves faster. You win by making authenticity cheap to prove and abuse expensive to attempt.

Continue Reading

Need this built properly? See services or get in touch.

Work with me

LET’S TALK

Have a project in mind after reading this? Send a brief and I usually reply within 24 to 48 hours.

Start a project →

More reading

All articles