Jev recipe / content quality
AI Slop Detection with Jev: A Typed Quality Gate
Screen mass-produced filler content before it ships: one evaluate call returns an is_slop verdict and an ordered quality score, and calibrated confidence decides between auto-block, flag, and human review.
Build the slop quality gate in six steps
{
"is_slop": {
"type": "noul",
"instructions": "Does this text read as AI slop: mass-produced filler with template phrasing, formatting abuse, and low fact density?"
},
"quality": {
"type": "score",
"instructions": "Rate the editorial quality of this text",
"criteria": {
"1": "Pure slop: template phrases and filler, no original information",
"2": "Mostly filler: heavy repetition, vague sourcing, few concrete facts",
"3": "Mixed: readable but generic, thin sourcing",
"4": "Good: specific and sourced, reads like a human wrote it",
"5": "Excellent: original insight, dense facts, clear voice"
}
}
}What counts as AI slop: six tell families
Slop is a family resemblance, not a single signal. The checklist below groups the recurring tells into six families — use them as criteria vocabulary when you define what slop means for your product. The widely quoted "35 tells" figure comes from the public description of the madewithjev AI slop checker, not an official spec; treat the exact count as one tool's taxonomy, not ground truth.
Template phrasing & repetition
Repeated n-grams, stock openers, and a sentence rhythm with no variation.
- Stock openers like "In today's fast-paced world" or "In the ever-evolving landscape of"
- The same n-gram recurring across paragraphs
- Formulaic transitions: "Moreover," "Furthermore," "In conclusion"
- The prompt or title restated verbatim as the first paragraph
- Uniform sentence length with no rhythm variation
- Filler hedges: "It's important to note that," "It's worth mentioning"
List & bold abuse
Formatting doing work the prose should have done.
- Bullet lists where a single sentence would do
- Bolded phrases scattered mid-sentence for fake emphasis
- A heading for every two sentences of body text
- Numbered "steps" that are actually just topics
- Bold-on-first-use glossaries nobody asked for
Hollow summary phrases
Conclusions that compress nothing and add less.
- "In conclusion, X is a powerful tool that can help you..."
- Verb stacks: "unlock," "elevate," "empower," "streamline," "supercharge"
- "It's not just X — it's Y" constructions
- A TL;DR that restates instead of compressing
- Padding nouns: "the world of," "the realm of," "the landscape of"
Em-dash & emoji tells
Punctuation habits that survive across models and prompts.
- Em-dash as the default sentence joiner — several per paragraph
- Emoji used as bullet markers or section headers
- Curly and straight quotes mixed on one page
- One-line "punchy" paragraphs stacked back to back
Low fact density & vague sourcing
Nothing checkable: no names, numbers, or dates where they would naturally appear.
- "Studies show" with no study named
- "Many experts agree" without a single expert
- No figures, versions, or dates where they would be natural
- Abstract claims where a concrete example would fit
- Facts that merely echo the question back
Structural symmetry
Sections machined to equal size — the geometry of batch production.
- Every H2 section the same length
- Perfectly parallel headings ("5 Benefits... 5 Challenges...")
- An FAQ answering questions the article itself invented
- A "final thoughts" block that adds one last generic claim
AI slop detection & three-lane routing simulator
Pick a sample text and run the two-question check — confidence decides the quarantine, flag, or human lane (front-end simulation only, no API calls):
Press "Run Jev check" to see the is_slop verdict, quality score, and lane routing.
Slop detection benchmark: Jev typed gate vs generative LLM detector
input-only billing, zero output tokens
input plus generated verdict tokens
single forward pass, no generation
token-by-token verdict generation
Noul and Score are native outputs, nothing to parse
the policy is inspectable and diffable
one number gates the block / flag / review lanes
uncalibrated without extra logprob plumbing
Production code: the slop gate in TypeScript and Python
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
// Policy lives in application code — the model only reports facts.
const AUTO_BLOCK_CONFIDENCE = 0.85; // >= 0.85 -> quarantine
const FLAG_MIN_CONFIDENCE = 0.6; // 0.60-0.85 -> flag, below -> human
const QUESTIONS = {
is_slop: {
type: "noul",
instructions:
"Does this text read as AI slop: mass-produced filler with template phrasing, formatting abuse, and low fact density?",
},
quality: {
type: "score",
instructions: "Rate the editorial quality of this text",
criteria: {
"1": "Pure slop: template phrases and filler, no original information",
"2": "Mostly filler: heavy repetition, vague sourcing, few concrete facts",
"3": "Mixed: readable but generic, thin sourcing",
"4": "Good: specific and sourced, reads like a human wrote it",
"5": "Excellent: original insight, dense facts, clear voice",
},
},
} as const;
export type SlopLane =
| "auto_quarantine"
| "flag_review"
| "human_review"
| "publish";
export async function checkSlop(
text: string,
metadata: Record<string, unknown>,
) {
const response = await fetch(JEV_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
state: { text, ...metadata },
questions: QUESTIONS,
}),
});
if (!response.ok) throw new Error("Jev evaluate failed");
const data = await response.json();
const { is_slop, quality } = data;
const lane: SlopLane = !is_slop.answer
? "publish"
: is_slop.confidence >= AUTO_BLOCK_CONFIDENCE
? "auto_quarantine"
: is_slop.confidence >= FLAG_MIN_CONFIDENCE
? "flag_review"
: "human_review";
return {
lane,
is_slop: is_slop.answer,
slop_confidence: is_slop.confidence,
quality: quality.answer,
quality_confidence: quality.confidence,
};
}
// Typical path: ~100-500ms per text, input tokens only, 0% schema errorsimport requests
JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_BLOCK_CONFIDENCE = 0.85 # >= 0.85 -> quarantine
FLAG_MIN_CONFIDENCE = 0.60 # 0.60-0.85 -> flag, below -> human review
QUESTIONS = {
"is_slop": {
"type": "noul",
"instructions": "Does this text read as AI slop: mass-produced filler with template phrasing, formatting abuse, and low fact density?",
},
"quality": {
"type": "score",
"instructions": "Rate the editorial quality of this text",
"criteria": {
"1": "Pure slop: template phrases and filler, no original information",
"2": "Mostly filler: heavy repetition, vague sourcing, few concrete facts",
"3": "Mixed: readable but generic, thin sourcing",
"4": "Good: specific and sourced, reads like a human wrote it",
"5": "Excellent: original insight, dense facts, clear voice",
},
},
}
def check_slop(text: str, source: str) -> dict:
resp = requests.post(
JEV_ENDPOINT,
json={"state": {"text": text, "source": source}, "questions": QUESTIONS},
timeout=5,
)
resp.raise_for_status()
result = resp.json()
slop, quality = result["is_slop"], result["quality"]
if not slop["answer"]:
lane = "publish"
elif slop["confidence"] >= AUTO_BLOCK_CONFIDENCE:
lane = "auto_quarantine"
elif slop["confidence"] >= FLAG_MIN_CONFIDENCE:
lane = "flag_review"
else:
lane = "human_review"
return {
"source": source,
"lane": lane,
"slop_confidence": slop["confidence"],
"quality": quality["answer"],
}
def scan_backlog(pages: list[dict]) -> dict:
# ~100-500ms per page: a 10,000-page CMS audit finishes in about an
# hour single-threaded and costs pennies at input-only pricing.
lanes: dict[str, int] = {}
for page in pages:
verdict = check_slop(page["text"], page["url"])
lanes[verdict["lane"]] = lanes.get(verdict["lane"], 0) + 1
return lanesAI slop detection FAQ
Can Jev detect AI slop reliably?
Reliably enough to triage with, not to convict on. Slop is a family resemblance — template phrasing, formatting abuse, hollow summaries, low fact density — and a typed two-question pass (is_slop Noul plus a 1–5 quality Score) catches those tells in a single ~100–500ms forward pass with zero format errors. The reliability you can automate against comes from calibrated confidence: thresholds become accuracy contracts, and every sub-threshold or borderline call lands in a review lane instead of an automatic deletion.
What confidence threshold should I use?
Start with 0.85 for auto-quarantine, 0.60–0.85 for flag-and-review, and below 0.60 for human review — then derive your own from labeled samples. Plot predicted confidence against editor verdicts on a calibration curve and place each cutoff where a mistake is cheap enough to absorb. Run the policy in shadow mode (log-only) before enforcing it, widen the middle band for flip-prone text (hysteresis), and re-validate quarterly as your content mix drifts.
How is this different from generative LLM detectors or AI text classifiers?
A generative detector produces its verdict as text: 1,500–3,000ms of token-by-token generation, output tokens billed on every call, and 2–8% of responses arriving as malformed JSON that needs parse-and-retry loops. Its rubric lives inside a prompt, so the policy drifts from run to run. Jev inverts this: criteria are typed, versioned code, the answer arrives as calibrated probabilities over Noul/Score labels in one forward pass with 0% format errors, and billing is input-only — roughly $0.0008 per thousand texts.
Does it work on human text?
Yes — with one honest caveat: the gate measures textual qualities, not authorship. Humans wrote template-riddled, filler-heavy copy long before LLMs existed, so a confident is_slop=true on a human-written piece is a correct slop verdict, not a false one. The real false-positive risk runs the other way: terse, fact-dense human copy scores cleanly, and heavily edited AI-assisted drafts land near the boundary — which is exactly what the 0.60–0.85 flag lane and the human review lane are for.
Where does the "35 tells" number come from?
From the public description of the madewithjev AI slop checker — a free web tool that flags slop tells in a single pass (~243ms per check). It is one tool's taxonomy, not an official Jev specification. The six families on this page (template phrasing, list & bold abuse, hollow summaries, punctuation tells, low fact density, structural symmetry) organize the same territory into criteria vocabulary you can put in your own schema.
Can I scan an entire site or CMS backlog with this?
Yes — that is the batch pattern in the Python example. At ~100–500ms per page, a 10,000-page audit finishes in about an hour single-threaded, costs pennies at input-only pricing, and never needs a GPU. Map each lane to a CMS status (publish / review / quarantine), log answer + confidence + lane for every page so the audit trail survives policy changes, and re-scan on a schedule — the jev-radar CASEBOOK tracks roughly eleven open-source tools built on exactly this loop.