Almost every moderation bot starts the same way: the admin writes a list of banned words. It holds up for about a week. Then someone posts q u i c k c a s h with spaces between the letters, and the list is blind again.
Two problems with a word list
It is trivial to slip past. Letters get spaced out, an "o" becomes a zero, Latin and Cyrillic characters get mixed inside one word, invisible characters are dropped between letters. A human still reads the message perfectly; the filter sees a word it has never heard of.
It cannot read intent. "Crypto" in an investing chat is an ordinary word. The same word in "dm me, I'll explain crypto, $500 a day guaranteed" is an advert. A list treats both the same, deletes both, and the admin gets a thread of annoyed members.
What ModerBot does instead
Checks run in two stages, and the order matters.
First, fast local rules. Links, invites, channel forwards, flooding, invisible characters and text-direction control symbols. This takes milliseconds and never reaches a model. Deobfuscation happens here too: mixed alphabets and swapped characters are normalised before anything is matched, so "qu1ck ca$h" and "quick cash" are the same string by the time a rule looks at them.
Then the model, but only where it is needed. If the local rules settled nothing, the message is judged on meaning: what is being offered, to whom, and in what conversation. The model sees the surrounding messages, not one isolated line — which is why a member's question stays a question while the same words in a pitch become an advert.
Categories, not one big list
A decision is never just "banned or allowed". It belongs to a category: scams and phishing, advertising, illegal substances, insults, 18+ content and others. Each category has its own action, from "log it and do nothing" to a ban. That is how an advert can be removed quietly while a phishing link gets the sender banned.
Critical categories — scams, phishing, doxxing, threats — are always on and always strict: they harm members directly. The debatable ones, like profanity, links or bot mentions, are off by default, because what counts as noise in one chat is normal conversation in another.
Images are checked like text
Spam moved into pictures years ago: a screenshot defeats any text filter. The bot reads the text on an image and runs it through the same rules as a normal message. The image itself is not stored — the log keeps the verdict, not the file.
The real difference: you can see why
A word list can only tell you that something matched. The decision log shows the category, the reason and the action for every message. That changes the admin's job: instead of "the bot banned someone again, go figure out why", there is a line you can read, and a single setting to adjust if you disagree.
And when a decision is wrong, you change the sensitivity of that one category — not of moderation as a whole.