Every article on MGRID aggregates multiple sources. Some of those sources may contain AI-generated text. We believe readers deserve to know. Our originality scoring system analyzes both our own writing and every source we cite, providing a transparent signal of content authenticity.

Two-Layer Detection

Our system uses a two-layer approach inspired by commercial AI detection research. Neither layer alone is sufficient. Together they provide a reliable signal.

Layer 1: Statistical Analysis

Before any AI model is involved, we compute seven forensic linguistic signals from the raw text:

  • Burstiness — measures variation in sentence length. Human writers naturally alternate between short and long sentences. AI text tends toward uniformity.
  • Type-Token Ratio — vocabulary diversity across a normalized window. AI models cluster within a narrow lexical range.
  • Lexical Density — ratio of content words to total words. AI text tends to be hyper-formal with no conversational filler.
  • Phrase Fingerprinting — scans for 50+ phrases characteristic of large language models: filler, hollow transitions, and marketing buzzwords.
  • Transition Density — frequency of hollow transition phrases per sentence.
  • Repetition Score — repeated phrasing patterns that indicate algorithmic text generation.
  • Structural Analysis — sentence count and average length distribution.

Layer 2: Ensemble Judge

The statistical signals are injected into a prompt sent to a large language model acting as a weighted aggregation function. It synthesizes the hard statistical evidence with its own reading of tone, structure, and domain specificity to produce a final score.

We use multiple models with automatic fallbacks to ensure consistent scoring regardless of any single provider’s availability.

Reading the Score

Originality is displayed as a number from 0 to 100. The spectrum below shows where scores fall:

0 – 29 Synthetic
30 – 69 Mixed
70 – 100 Original

70 – 100

Original

Strong signals of human authorship. Varied sentence rhythm, specific details, domain expertise, and no AI phrase patterns detected.

30 – 69

Mixed

Some AI patterns alongside human signals. Common in AI-assisted drafts with substantial editing, or in formal technical writing.

0 – 29

Synthetic

Strong AI signals throughout. Uniform sentence structure, AI phrase fingerprints, generic language lacking domain-specific detail.

Where You See It

  • Source panel — each cited source displays a badge with its originality score, showing which sources contain original reporting versus AI-generated content.
  • Feedback panel — the article’s own originality score appears alongside pipeline status and content quality metrics.

What the Score Doesn’t Tell You

  • Technical writing — formal regulatory or engineering text can score lower because it shares structural characteristics with AI output.
  • Government publications — policy documents often score in the mixed range due to their formal, templated structure.
  • Short texts — passages under 200 characters are skipped; insufficient signal for reliable analysis.
  • Adversarial evasion — paraphrasing attacks can reduce detection accuracy. This system is designed for editorial transparency, not forensic attribution.