As of early 2026, mainstream LLM-generated text exhibits strong statistical patterns that can be effectively distinguished from human-written content using traditional machine learning models【1†L1-L3】. This isn’t marketing fluff; a straightforward scikit-learn SVM trained on TF‑IDF features hits ~85% single‑sentence accuracy, and the same approach survives common anti‑detection tricks【1†L13-L15】. For SaaS teams drowning in AI‑generated spam, fake reviews, or undisclosed AI‑assisted content, the implication is clear: you don’t need a giant LLM to catch an LLM.

The reality: how it actually works The author built a detector by scraping pre‑2022 human‑written fiction (≈10,000 documents) and using seven different LLMs—Gemini‑3‑Flash for summaries, then Gemini‑3‑Pro, Qwen‑Coder‑Plus, GLM‑5, GLM‑4.7, Kimi‑K2.5, Doubao‑Seed‑Code, and DeepSeek‑V3.2—to regenerate chapter‑level text, creating a balanced corpus【1†L30-L38】. After splitting into sentences and stripping non‑CJK/English characters, they fed the raw text into a TF‑IDF vectorizer (500 k features) and a LinearSVC classifier【1†L48-L52】. Individual binary models scored 84‑89% accuracy; majority voting (≥2 of 7 models flagging a sentence) gave the final AI probability【1†L55-L60】. The JavaScript port for the web demo keeps the same 500 k features, bloating to 107 MB raw but gzipping to ~38 MB, with only a ~1% accuracy drop【1†L70-L73】.

The pain point: who this hits For enterprise content moderation, the cost of running a 7‑B‑parameter LLM just to detect AI text is prohibitive—both in dollars and latency. The classical approach runs on CPU; the JS demo processes a few‑thousand‑character input in milliseconds and a million‑character document in ~10 seconds on a modest laptop【1†L68-L70】. False positives are vanishingly low: at a 70% AI‑score threshold the false‑positive rate on pre‑2022 human fiction is <0.01%【1†L80-L82】. This means SaaS platforms can auto‑flag likely AI‑generated submissions without drowning moderators in noise, saving both operational overhead and potential reputational damage from missed AI‑generated plagiarism or astroturfing.

Failure modes: where this breaks in the real world The detector assumes a statistical shift between LLM and human word choice; it degrades when LLMs are fine‑tuned to mimic a specific human style or when adversarial editing follows generation【3†L1-L4】. Simple evasion tactics—round‑trip translation or prompting the LLM to “reduce AI flavor”—only shave a few points off the score (e.g., 89.9% → 85.0% after Google Translate, → 79.3% after a complex rewrite prompt)【1†L88-L94】. The method also struggles with very short texts (<50 words) where TF‑IDF features become sparse, and it offers no explainability beyond feature weights, which may frustrate compliance teams needing granular audit trails【3†L5-L8】. Finally, the model is trained on a narrow corpus of web fiction; performance on highly technical or legal documents may vary, requiring domain‑specific retraining.

The blueprint: what the reader does about it

  1. Gather domain‑specific data: collect a few thousand human‑written samples from your target channel (e.g., support tickets, product reviews) and generate matching LLM outputs using the same models you suspect are in play【1†L30-L38】.
  2. Build a baseline: split text into sentences, TF‑IDF vectorize (max_features=500k), train a LinearSVC, and validate accuracy; aim for >80% on a hold‑out set【1†L48-L52】.
  3. Deploy lightweight: export the vectorizer and model to ONNX or, for maximum simplicity, reimplement TF‑IDF + linear inference in JavaScript as the author did—this keeps the detector serverless and edge‑friendly【1†L68-L73】.
  4. Set thresholds: use the AI‑score (percentage of sentences flagged by ≥2 models) with a 70% cutoff for near‑zero false positives; adjust to 60% if you need higher recall and can tolerate ~0.04% FP【1†L80-L82】.
  5. Monitor drift: weekly, re‑score a sample of known human content; if the false‑positive rate creeps above 0.1%, retrain with fresh human data and recent LLM outputs【3†L9-L12】. This approach gives you a cheap, interpretable, and surprisingly robust first line of defense against undisclosed AI‑generated text—without betting your budget on the next hyped “AI detector” startup.

Sources

  1. Detecting LLM-Generated Web Fiction with "Classical" Machine Learning (AIGC Text Detection)
  2. Detecting LLM-Generated Texts with “Classical” Machine Learning
  3. Detecting LLM-Generated Texts with Classical Machine Learning in 2026