If you think you can create a website and allow anyone to post content on there without moderation, then bless your heart , as my imaginary southern grandma would say. Being able to post anything anonymously means, for some people, that they will do exactly this. I won’t mention the content you’ll end up with, as this is a family-friendly blog.
Reddit “solved” content moderation by outsourcing it to the communities themselves. Being a moderator on Reddit is a huge time investment. Everyone wants your attention, and Reddit gets a lot of it, so any decently successful subreddit gets flooded with ads, bots and things that are not in the interest of the subreddit. This also led to the cliché of the typical Reddit mod: unemployed, always online and hungry for power, as you need the first two things to be able to successfully moderate.
The issue then arises that these mods can shape their communities however they want, for better or worse. Often for the worse, sadly. But somebody has to do it, and the community would be waaaay worse without anyone moderating it.
With the advent of ML, our computers can understand language well enough to even understand irony, so they can say if someone genuinely hates on someone or is just doing it as a joke.
Detecting irony
Detecting irony
Small rabbit hole here: How good are computers actually at detecting irony?
The paper LLMs vs. humans in sarcasm detection on German soccer tweets created a small dataset and tested it against models. While their paper showed that humans were still better, they used older models. I ran it against the latest breed of SOTA models as well as some decision models (including Jev). SOTA models are now, on average, better at understanding your joke than your neighbour - use this information however you like. Jev is not as good yet, but clearly outperformed the other decision models that I tested. Certainly good enough, I felt.
Results Cost / Performance
Who gets the joke? Detecting irony in 100 German soccer tweets. Irony detection accuracy on 100 German soccer tweets, grouped into decision models and general-purpose LLMs. Claude Opus 5.5 leads at 96.25%; Jev scores 79%. Human reference lines show 83.3% average accuracy and 97% majority-vote accuracy. Bars show mean accuracy; whiskers show the lowest and highest repeated-run scores. GPT and Claude: mean of 4 runs. Gemini and DeepSeek: mean of 3 runs. All other models: 1 run. Who gets the joke? Detecting irony in 100 German soccer tweets 0% 25% 50% 75% 100% Human majority vote 97% Human average 83.3% Decision models General-purpose LLMs Jev 1.13.0 79% Jeeves 9B (MLX) 68% Imajev-4B 62% SemIf Qwen3.5-4B 62% Laya Multilingual 53% Claude Opus 5.5(medium) 96.25% Gemini 3.1 Pro(high) 93.33% Gemini 3.8 Flash(high) 92.33% GPT-6 Astra(medium) 92.25% GPT-6 Sol(medium) 90.5% DeepSeek V4.1 Flash(high) 86.67% GPT-6 Luna(medium) 82.75% Claude Sonnet 5(medium) 74.75%
Who gets the joke? Detecting irony in 100 German soccer tweets. Irony detection accuracy on 100 German soccer tweets, grouped into decision models and general-purpose LLMs. Claude Opus 5.5 leads at 96.25%; Jev scores 79%. Human reference lines show 83.3% average accuracy and 97% majority-vote accuracy. Bars show mean accuracy; whiskers show the lowest and highest repeated-run scores. GPT and Claude: mean of 4 runs. Gemini and DeepSeek: mean of 3 runs. All other models: 1 run. Who gets the joke? Detecting irony in 100 German soccer tweets 0% 25% 50% 75% 100% Human majority vote 97% Human average 83.3% Decision models General-purpose LLMs Jev 1.13.0 79% Jeeves 9B (MLX) 68% Imajev-4B 62% SemIf Qwen3.5-4B 62% Laya Multilingual 53% Claude Opus 5.5(medium) 96.25% Gemini 3.1 Pro(high) 93.33% Gemini 3.8 Flash(high) 92.33% GPT-6 Astra(medium) 92.25% GPT-6 Sol(medium) 90.5% DeepSeek V4.1 Flash(high) 86.67% GPT-6 Luna(medium) 82.75% Claude Sonnet 5(medium) 74.75%
As computers are now so good at understanding language and the intent behind it, automoderation is widely used in moderating the internet; for example, YouTube is basically fully auto-moderated by now . Every other social network most certainly is similar. Humans should not have to review the filth of the internet.
These automoderation techniques are now accessible to anyone, basically. You don't need your own huge training corpus anymore — no fancy pipelines, no huge infra costs. One very interesting one was just released.
Jev is a decision model that you don't have to train to use. It's quite good at classifying anything you throw at it. And as it is very fast and cheap in doing so, I was wondering: could you automoderate a Reddit-like website purely with Jev? I could run a decently sized community with it: 10,000 posts and comments a day would cost about $15 for a 30-day month.
For scale, Reddit reported 4.4 billion posts and comments in 2025 . Running that volume through Jev at the same rate would come to roughly $18,000 a month. So actually quite attainable.
So I built Jevdit — a Reddit-like social network, where everything is moderated by Jevdit. You can post and create subreddits however you want, as long as Jev permits it. You can even change the rules of a subreddit — by creating a Jevdit post where you argue with the community about it, and Jev then decides if it thinks that enough people agree with the changes.
To make sure I don’t just blindly put my hand into the fire here while Jev looks at me with a smirk, saying “Trust me, bro,” I ran some of my own benchmarks to calm down my anxiety about putting an open website into the birthplace of 4chan.
Thankfully, there are several public datasets that you can use to see how well your automoderation technique performs. I also ran it with the latest models as well as some other interesting models. I combined the results of all datasets here for an average; below, you can see the detailed results if you care.
A moderator has two jobs: stop the bad stuff and let people talk. You want a balance — if you just block every comment, you definitely catch 100% of the bad stuff. So we have two bars here: how much of the bad stuff was caught vs. how much of the good content was allowed.
Jev did really well on this, could be considered SOTA with the models i tested. The cool thing about models such as Jev is, that you can decide how strict you want the moderation to be, as it will give a number and you can decide the cutoff. The benchmark shows multiple settings for Jev and how it changes the harmful caught vs benign allowed split.
Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses.
Results Cost / Performance
Who recognizes bad and good content? All five datasets · pooled results. Moderation results for all five datasets, grouped into decision models and general-purpose LLMs. Orange bars show harmful content blocked; violet bars show benign content allowed. Jev 1.13.0 (balanced) catches 95.4% of harmful examples and allows 84.9% of benign examples. LLMs catch 83.8%–91.3% and allow 88.7%–96.1%, showing the tradeoff between catching harm and overblocking. Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses. Who recognizes bad and good content? All five datasets · pooled results Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Jev 1.13.0(relaxed) 90.2% 92.1% Jev 1.13.0(balanced) 95.4% 84.9% Safeguard 20B(medium) 81.4% 96% Jeeves 9B (MLX)(balanced) 95.4% 77% Imajev-4B(balanced) 96% 68.4% Laya English (expanded)(balanced) 95.4% 65.8% OpenAI Moderation 63% 95.4% Detoxify unbiased 64.2% 92.8% Jev 1.13.0(strict) 99.4% 57.2% Laya English(balanced) 95.4% 57.9% SemIf Qwen3.5-4B(balanced) 95.4% 42.8% Claude Opus 5.5(medium) 85% 96.1% GPT-6 Sol(medium) 91.3% 89.5% DeepSeek V4.1 Flash(high) 90% 90.3% DeepSeek V4 Pro(high) 87.2% 92.8% Claude Sonnet 5(medium) 83.8% 96.1% GPT-6 Astra(medium) 89% 89.4% Gemini 3.1 Pro(high) 86.6% 91.4% Gemini 3.8 Flash(high) 84.4% 93.4% GPT-6 Luna(medium) 88.3% 88.7%
Who recognizes bad and good content? All five datasets · pooled results. Moderation results for all five datasets, grouped into decision models and general-purpose LLMs. Orange bars show harmful content blocked; violet bars show benign content allowed. Jev 1.13.0 (balanced) catches 95.4% of harmful examples and allows 84.9% of benign examples. LLMs catch 83.8%–91.3% and allow 88.7%–96.1%, showing the tradeoff between catching harm and overblocking. Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses. Who recognizes bad and good content? All five datasets · pooled results Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Jev 1.13.0(relaxed) 90.2% 92.1% Jev 1.13.0(balanced) 95.4% 84.9% Safeguard 20B(medium) 81.4% 96% Jeeves 9B (MLX)(balanced) 95.4% 77% Imajev-4B(balanced) 96% 68.4% Laya English (expanded)(balanced) 95.4% 65.8% OpenAI Moderation 63% 95.4% Detoxify unbiased 64.2% 92.8% Jev 1.13.0(strict) 99.4% 57.2% Laya English(balanced) 95.4% 57.9% SemIf Qwen3.5-4B(balanced) 95.4% 42.8% Claude Opus 5.5(medium) 85% 96.1% GPT-6 Sol(medium) 91.3% 89.5% DeepSeek V4.1 Flash(high) 90% 90.3% DeepSeek V4 Pro(high) 87.2% 92.8% Claude Sonnet 5(medium) 83.8% 96.1% GPT-6 Astra(medium) 89% 89.4% Gemini 3.1 Pro(high) 86.6% 91.4% Gemini 3.8 Flash(high) 84.4% 93.4% GPT-6 Luna(medium) 88.3% 88.7%
Results by dataset Who recognizes bad and good content? Civil Comments Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Detoxify unbiased 86.7% 96.7% Safeguard 20B(medium) 86.7% 93.3% OpenAI Moderation 90% 90% Jev 1.13.0(relaxed) 93.3% 83.3% Laya English (expanded)(balanced) 100% 63.3% Jev 1.13.0(balanced) 93.3% 66.7% Imajev-4B(balanced) 100% 60% Jeeves 9B (MLX)(balanced) 93.3% 63.3% Laya English(balanced) 100% 56.7% SemIf Qwen3.5-4B(balanced) 100% 23.3% Jev 1.13.0(strict) 100% 6.7% Claude Sonnet 5(medium) 83.3% 96.7% DeepSeek V4 Pro(high) 93.1% 86.7% Gemini 3.8 Flash(high) 86.7% 86.7% Claude Opus 5.5(medium) 83.3% 90% GPT-6 Sol(medium) 90% 73.3% Gemini 3.1 Pro(high) 86.2% 76.7% GPT-6 Astra(medium) 90% 72.4% DeepSeek V4.1 Flash(high) 88% 73.1% GPT-6 Luna(medium) 89.3% 70%
Who recognizes bad and good content? Civil Comments Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Detoxify unbiased 86.7% 96.7% Safeguard 20B(medium) 86.7% 93.3% OpenAI Moderation 90% 90% Jev 1.13.0(relaxed) 93.3% 83.3% Laya English (expanded)(balanced) 100% 63.3% Jev 1.13.0(balanced) 93.3% 66.7% Imajev-4B(balanced) 100% 60% Jeeves 9B (MLX)(balanced) 93.3% 63.3% Laya English(balanced) 100% 56.7% SemIf Qwen3.5-4B(balanced) 100% 23.3% Jev 1.13.0(strict) 100% 6.7% Claude Sonnet 5(medium) 83.3% 96.7% DeepSeek V4 Pro(high) 93.1% 86.7% Gemini 3.8 Flash(high) 86.7% 86.7% Claude Opus 5.5(medium) 83.3% 90% GPT-6 Sol(medium) 90% 73.3% Gemini 3.1 Pro(high) 86.2% 76.7% GPT-6 Astra(medium) 90% 72.4% DeepSeek V4.1 Flash(high) 88% 73.1% GPT-6 Luna(medium) 89.3% 70%
Who recognizes bad and good content? Conversations Gone Awry. Moderation results for Conversations Gone Awry, grouped into decision models and general-purpose LLMs. Orange bars show harmful content blocked; violet bars show benign content allowed. Jev 1.13.0 (balanced) catches 100% of harmful examples and allows 81.7% of benign examples. LLMs catch 66.7%–83.3% and allow 89.8%–96.7%, showing the tradeoff between catching harm and overblocking. Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses. Who recognizes bad and good content? Conversations Gone Awry Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Jev 1.13.0(relaxed) 95% 90% Jev 1.13.0(balanced) 100% 81.7% Detoxify unbiased 80% 96.7% Jeeves 9B (MLX)(balanced) 91.7% 76.7% Laya English (expanded)(balanced) 90% 75% Laya English(balanced) 96.7% 66.7% OpenAI Moderation 63.3% 100% Imajev-4B(balanced) 100% 61.7% Safeguard 20B(medium) 66.1% 95% Jev 1.13.0(strict) 100% 60% SemIf Qwen3.5-4B(balanced) 100% 33.3% GPT-6 Sol(medium) 83.3% 90% DeepSeek V4.1 Flash(high) 82.2% 90.9% GPT-6 Astra(medium) 80% 90% GPT-6 Luna(medium) 80% 89.8% Claude Opus 5.5(medium) 71.7% 96.7% Gemini 3.1 Pro(high) 75% 93.3% DeepSeek V4 Pro(high) 73.3% 90% Claude Sonnet 5(medium) 66.7% 95% Gemini 3.8 Flash(high) 66.7% 93.3%
Who recognizes bad and good content? Conversations Gone Awry. Moderation results for Conversations Gone Awry, grouped into decision models and general-purpose LLMs. Orange bars show harmful content blocked; violet bars show benign content allowed. Jev 1.13.0 (balanced) catches 100% of harmful examples and allows 81.7% of benign examples. LLMs catch 66.7%–83.3% and allow 89.8%–96.7%, showing the tradeoff between catching harm and overblocking. Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses. Who recognizes bad and good content? Conversations Gone Awry Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Jev 1.13.0(relaxed) 95% 90% Jev 1.13.0(balanced) 100% 81.7% Detoxify unbiased 80% 96.7% Jeeves 9B (MLX)(balanced) 91.7% 76.7% Laya English (expanded)(balanced) 90% 75% Laya English(balanced) 96.7% 66.7% OpenAI Moderation 63.3% 100% Imajev-4B(balanced) 100% 61.7% Safeguard 20B(medium) 66.1% 95% Jev 1.13.0(strict) 100% 60% SemIf Qwen3.5-4B(balanced) 100% 33.3% GPT-6 Sol(medium) 83.3% 90% DeepSeek V4.1 Flash(high) 82.2% 90.9% GPT-6 Astra(medium) 80% 90% GPT-6 Luna(medium) 80% 89.8% Claude Opus 5.5(medium) 71.7% 96.7% Gemini 3.1 Pro(high) 75% 93.3% DeepSeek V4 Pro(high) 73.3% 90% Claude Sonnet 5(medium) 66.7% 95% Gemini 3.8 Flash(high) 66.7% 93.3%
Who recognizes bad and good content? HateCheck. Moderation results for HateCheck, grouped into decision models and general-purpose LLMs. Orange bars show harmful content blocked; violet bars show benign content allowed. Jev 1.13.0 (balanced) catches 100% of harmful examples and allows 100% of benign examples. LLMs catch 97.2%–100% and allow 100%–100%, showing the tradeoff between catching harm and overblocking. Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses. Who recognizes bad and good content? HateCheck Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Jev 1.13.0(balanced) 100% 100% Jev 1.13.0(relaxed) 100% 100% Safeguard 20B(medium) 100% 100% OpenAI Moderation 100% 75% Jev 1.13.0(strict) 100% 66.7% Jeeves 9B (MLX)(balanced) 100% 66.7% Imajev-4B(balanced) 100% 50% SemIf Qwen3.5-4B(balanced) 100% 33.3% Detoxify unbiased 80.6% 50% Laya English (expanded)(balanced) 100% 25% Laya English(balanced) 94.4% 25% GPT-6 Luna(medium) 100% 100% GPT-6 Sol(medium) 100% 100% GPT-6 Astra(medium) 100% 100% Claude Opus 5.5(medium) 100% 100% Gemini 3.8 Flash(high) 100% 100% Gemini 3.1 Pro(high) 100% 100% DeepSeek V4.1 Flash(high) 100% 100% DeepSeek V4 Pro(high) 100% 100% Claude Sonnet 5(medium) 97.2% 100%
Who recognizes bad and good content? HateCheck. Moderation results for HateCheck, grouped into decision models and general-purpose LLMs. Orange bars show harmful content blocked; violet bars show benign content allowed. Jev 1.13.0 (balanced) catches 100% of harmful examples and allows 100% of benign examples. LLMs catch 97.2%–100% and allow 100%–100%, showing the tradeoff between catching harm and overblocking. Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses. Who recognizes bad and good content? HateCheck Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Jev 1.13.0(balanced) 100% 100% Jev 1.13.0(relaxed) 100% 100% Safeguard 20B(medium) 100% 100% OpenAI Moderation 100% 75% Jev 1.13.0(strict) 100% 66.7% Jeeves 9B (MLX)(balanced) 100% 66.7% Imajev-4B(balanced) 100% 50% SemIf Qwen3.5-4B(balanced) 100% 33.3% Detoxify unbiased 80.6% 50% Laya English (expanded)(balanced) 100% 25% Laya English(balanced) 94.4% 25% GPT-6 Luna(medium) 100% 100% GPT-6 Sol(medium) 100% 100% GPT-6 Astra(medium) 100% 100% Claude Opus 5.5(medium) 100% 100% Gemini 3.8 Flash(high) 100% 100% Gemini 3.1 Pro(high) 100% 100% DeepSeek V4.1 Flash(high) 100% 100% DeepSeek V4 Pro(high) 100% 100% Claude Sonnet 5(medium) 97.2% 100%
Who recognizes bad and good content? YouTube Spam. Moderation results for YouTube Spam, grouped into decision models and general-purpose LLMs. Orange bars show harmful content blocked; violet bars show benign content allowed. Jev 1.13.0 (balanced) catches 80% of harmful examples and allows 93.3% of benign examples. LLMs catch 83.3%–93.3% and allow 93.3%–96.7%, showing the tradeoff between catching harm and overblocking. Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses. Who recognizes bad and good content? YouTube Spam Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Jeeves 9B (MLX)(balanced) 96.7% 83.3% Jev 1.13.0(balanced) 80% 93.3% Safeguard 20B(medium) 73.3% 96.6% Jev 1.13.0(strict) 96.7% 70% Imajev-4B(balanced) 76.7% 83.3% Jev 1.13.0(relaxed) 60% 96.7% Laya English (expanded)(balanced) 93.3% 53.3% Laya English(balanced) 90% 36.7% SemIf Qwen3.5-4B(balanced) 73.3% 53.3% OpenAI Moderation 0% 96.7% Detoxify unbiased 3.3% 93.3% DeepSeek V4.1 Flash(high) 91.3% 95.8% GPT-6 Sol(medium) 93.3% 93.3% Claude Sonnet 5(medium) 93.3% 93.3% Claude Opus 5.5(medium) 86.7% 96.7% DeepSeek V4 Pro(high) 86.7% 96.7% Gemini 3.8 Flash(high) 90% 93.3% GPT-6 Astra(medium) 86.7% 93.3% Gemini 3.1 Pro(high) 86.7% 93.3% GPT-6 Luna(medium) 83.3% 93.3%
Who recognizes bad and good content? YouTube Spam. Moderation results for YouTube Spam, grouped into decision models and general-purpose LLMs. Orange bars show harmful content blocked; violet bars show benign content allowed. Jev 1.13.0 (balanced) catches 80% of harmful examples and allows 93.3% of benign examples. LLMs catch 83.3%–93.3% and allow 93.3%–96.7%, showing the tradeoff between catching harm and overblocking. Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses. Who recognizes bad and good content? YouTube Spam Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Jeeves 9B (MLX)(balanced) 96.7% 83.3% Jev 1.13.0(balanced) 80% 93.3% Safeguard 20B(medium) 73.3% 96.6% Jev 1.13.0(strict) 96.7% 70% Imajev-4B(balanced) 76.7% 83.3% Jev 1.13.0(relaxed) 60% 96.7% Laya English (expanded)(balanced) 93.3% 53.3% Laya English(balanced) 90% 36.7% SemIf Qwen3.5-4B(balanced) 73.3% 53.3% OpenAI Moderation 0% 96.7% Detoxify unbiased 3.3% 93.3% DeepSeek V4.1 Flash(high) 91.3% 95.8% GPT-6 Sol(medium) 93.3% 93.3% Claude Sonnet 5(medium) 93.3% 93.3% Claude Opus 5.5(medium) 86.7% 96.7% DeepSeek V4 Pro(high) 86.7% 96.7% Gemini 3.8 Flash(high) 90% 93.3% GPT-6 Astra(medium) 86.7% 93.3% Gemini 3.1 Pro(high) 86.7% 93.3% GPT-6 Luna(medium) 83.3% 93.3%
Who recognizes bad and good content? Site-policy tests. Moderation results for Site-policy tests, grouped into decision models and general-purpose LLMs. Orange bars show harmful content blocked; violet bars show benign content allowed. Jev 1.13.0 (balanced) catches 100% of harmful examples and allows 100% of benign examples. LLMs catch 100%–100% and allow 100%–100%, showing the tradeoff between catching harm and overblocking. Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses. Who recognizes bad and good content? Site-policy tests Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Jev 1.13.0(strict) 100% 100% Jev 1.13.0(balanced) 100% 100% Jev 1.13.0(relaxed) 100% 100% Safeguard 20B(medium) 100% 100% Jeeves 9B (MLX)(balanced) 100% 95% Imajev-4B(balanced) 100% 90% SemIf Qwen3.5-4B(balanced) 100% 90% Laya English (expanded)(balanced) 100% 85% Laya English(balanced) 94.1% 85% OpenAI Moderation 47.1% 100% Detoxify unbiased 41.2% 100% GPT-6 Luna(medium) 100% 100% GPT-6 Sol(medium) 100% 100% GPT-6 Astra(medium) 100% 100% Claude Sonnet 5(medium) 100% 100% Claude Opus 5.5(medium) 100% 100% Gemini 3.8 Flash(high) 100% 100% Gemini 3.1 Pro(high) 100% 100% DeepSeek V4.1 Flash(high) 100% 100% DeepSeek V4 Pro(high) 100% 100%
Who recognizes bad and good content? Site-policy tests. Moderation results for Site-policy tests, grouped into decision models and general-purpose LLMs. Orange bars show harmful content blocked; violet bars show benign content allowed. Jev 1.13.0 (balanced) catches 100% of harmful examples and allows 100% of benign examples. LLMs catch 100%–100% and allow 100%–100%, showing the tradeoff between catching harm and overblocking. Balanced targets at least 95% harmful content blocked; Jev relaxed targets 90%. Cutoffs were chosen and evaluated on these same 325 labeled cases. Jev strict retains the original rules. All modes reuse saved model responses. Who recognizes bad and good content? Site-policy tests Harmful caught Benign allowed 0% 25% 50% 75% 100% Decision models General-purpose LLMs Jev 1.13.0(strict) 100% 100% Jev 1.13.0(balanced) 100% 100% Jev 1.13.0(relaxed) 100% 100% Safeguard 20B(medium) 100% 100% Jeeves 9B (MLX)(balanced) 100% 95% Imajev-4B(balanced) 100% 90% SemIf Qwen3.5-4B(balanced) 100% 90% Laya English (expanded)(balanced) 100% 85% Laya English(balanced) 94.1% 85% OpenAI Moderation 47.1% 100% Detoxify unbiased 41.2% 100% GPT-6 Luna(medium) 100% 100% GPT-6 Sol(medium) 100% 100% GPT-6 Astra(medium) 100% 100% Claude Sonnet 5(medium) 100% 100% Claude Opus 5.5(medium) 100% 100% Gemini 3.8 Flash(high) 100% 100% Gemini 3.1 Pro(high) 100% 100% DeepSeek V4.1 Flash(high) 100% 100% DeepSeek V4 Pro(high) 100% 100%
So, with all this analysis, I feel a little bit confident I won't create another Tay . So check out jevdit.com and see if Jev can handle the moderation.