Insights

The moderation cascade: why most content should never reach an LLM

Fast filters, classifiers, an LLM policy judge and humans, each doing what it is best at. How to order them, what each layer costs, and where teams go wrong.

Nemanja Jeremenkovic · · 5 min read · Last reviewed September 22, 2026

Large language models are good at moderation. They read context, follow written policy and explain their decisions. That makes it tempting to send every post, message and profile straight to one. For most platforms that is the wrong design. It is slow, it is expensive, and it hides the easy cases inside the hard ones.

A better design is a cascade: a series of layers where each one settles what it can and passes the rest down. The cheap layers see everything. The expensive layers see only what they must.

The keel: a seven layer moderation system1. Intake: Every post, message, image and profile enters one queue, with user and context attached. 2. Fast filters: Hash matching for known abuse, plus scam and spam rules. Most traffic is settled here in milliseconds. 3. Classifiers: Models tuned on your data for text, images and behavior. They score risk, they do not decide alone. 4. Policy judge: An LLM reads your written policy and decides the hard cases. Every decision comes with a reason. 5. Human review: Only truly ambiguous items reach your team, ranked by risk, with the context already gathered. 6. Action and reporting: Takedowns, appeals, NCMEC reports and 48 hour removal clocks, logged for auditors. 7. Evals and cost: Every decision measured for accuracy, speed and cost, so the system gets cheaper and better each month.
Stages 2, 3, 4, 5 of the keel.

The four layers

1. Fast filters

Fast filters are rules and lookups that run in milliseconds. They include:

  • Hash matching against known abusive images and video, using perceptual hashes so that resized or recompressed copies still match.
  • Blocklists and patterns for known scam domains, shortened links, phone numbers and payment handles.
  • Rate and behavior rules: a new account sending the same message to 40 people in two minutes does not need a model to be judged.

Fast filters should settle most of your traffic. On a typical platform the large majority of content is ordinary, and much of what is not is repetitive: the same spam, the same scam script, the same known image. A filter that catches it costs almost nothing per item.

2. Classifiers

Classifiers are models that score risk for a specific harm: nudity, harassment, spam, self harm, solicitation. They can be vendor APIs, open models or models you train on your own labels. They are fast, usually tens of milliseconds, and cheap per call.

The key rule: classifiers score, they do not decide alone. Each one gives a number. Your thresholds turn numbers into actions:

  • Above a high threshold, act automatically.
  • Below a low threshold, allow.
  • In between, pass the item down to the next layer.

The band in the middle is where the cascade earns its money. Set it too wide and you flood the expensive layers. Set it too narrow and you make confident mistakes.

3. The policy judge

The policy judge is an LLM that reads your written policy and the item, with its context, and returns a decision and a reason. It handles the cases a score cannot: sarcasm, reclaimed slurs, a scam that uses no known keywords, a message that is fine alone and harmful in the conversation it sits in.

The judge is where your policy becomes executable. Two things matter most:

  • Structured output. A verdict from a fixed set, a category and a short reason that cites the policy section. Anything that does not parse becomes a review item, not an allow.
  • Honest uncertainty. When the judge is not confident, it should say so and route the item to a person. That is also what a good human moderator does.

4. Human review

Humans see what is left: the ambiguous, the high severity and the appeals. A good cascade does not just reduce the queue, it improves it. Each item arrives with the classifier scores, the judge's reason and the user's history already gathered, ranked by risk. Reviewers spend their time deciding, not searching.

What each layer costs

The numbers below are illustrative orders of magnitude, not benchmarks. Measure your own.

| Layer | Typical latency | Relative cost per item | Share of traffic it should see | |---|---|---|---| | Fast filters | Under 20 ms | Near zero | All of it | | Classifiers | 20 to 200 ms | Low | Most of it | | Policy judge | 0.5 to 3 s | Medium | A small share | | Human review | Minutes to hours | High | A very small share |

The point of the table is the shape. Every step down costs more and takes longer, so the share of traffic that reaches it should fall faster than its cost rises.

Where teams go wrong

Sending everything to the LLM. It works in the demo and fails on the invoice. It also adds seconds of latency to content that could have been settled instantly.

Treating vendor scores as decisions. A vendor's default threshold was tuned on someone else's platform. Your harms, your users and your tolerance for mistakes are different. Tune thresholds on your own labeled data.

No path for uncertainty. If a layer has to choose between allow and remove, it will make confident mistakes. Every layer needs a third option: pass it down.

Measuring accuracy but not cost. A cascade that is accurate and expensive is not finished. Track cost per decision for each layer so you can see where the money goes.

Forgetting the audit trail. Regulators and app stores increasingly want to know why content came down, and how fast. Log which layer decided, what it saw and what it said. The policy judge's reason is a gift here: store it.

How to start

You do not need to rebuild everything. Most teams already have one or two of these layers. Start by measuring:

  1. What share of your traffic does each existing layer settle?
  2. What does each layer cost per item?
  3. Where do human reviewers spend their time, and how much of it is on items a rule could have settled?

The answers usually point to one or two changes that pay for themselves quickly: a hash matching layer in front of an image classifier, a threshold band that routes to a judge instead of a person, or a report queue ranked by risk instead of time.

Relevant for communities and social apps, marketplaces, dating apps.