Every article on MGRID aggregates multiple sources. Some of those sources may contain AI-generated text. We believe readers deserve to know. Our originality scoring system analyzes both our own writing and every source we cite, providing a transparent signal of content authenticity.
Two-Layer Detection
Our system uses a two-layer approach inspired by commercial AI detection research. Neither layer alone is sufficient. Together they provide a reliable signal.
Layer 1: Statistical Analysis
Before any AI model is involved, we compute seven forensic linguistic signals from the raw text:
- Burstiness — measures variation in sentence length. Human writers naturally alternate between short and long sentences. AI text tends toward uniformity.
- Type-Token Ratio — vocabulary diversity across a normalized window. AI models cluster within a narrow lexical range.
- Lexical Density — ratio of content words to total words. AI text tends to be hyper-formal with no conversational filler.
- Phrase Fingerprinting — scans for 50+ phrases characteristic of large language models: filler, hollow transitions, and marketing buzzwords.
- Transition Density — frequency of hollow transition phrases per sentence.
- Repetition Score — repeated phrasing patterns that indicate algorithmic text generation.
- Structural Analysis — sentence count and average length distribution.
Layer 2: Ensemble Judge
The statistical signals are injected into a prompt sent to a large language model acting as a weighted aggregation function. It synthesizes the hard statistical evidence with its own reading of tone, structure, and domain specificity to produce a final score.
We use multiple models with automatic fallbacks to ensure consistent scoring regardless of any single provider’s availability.
Reading the Score
Originality is displayed as a number from 0 to 100. The spectrum below shows where scores fall:
70 – 100
Original
Strong signals of human authorship. Varied sentence rhythm, specific details, domain expertise, and no AI phrase patterns detected.
30 – 69
Mixed
Some AI patterns alongside human signals. Common in AI-assisted drafts with substantial editing, or in formal technical writing.
0 – 29
Synthetic
Strong AI signals throughout. Uniform sentence structure, AI phrase fingerprints, generic language lacking domain-specific detail.
Where You See It
- Source panel — each cited source displays a badge with its originality score, showing which sources contain original reporting versus AI-generated content.
- Feedback panel — the article’s own originality score appears alongside pipeline status and content quality metrics.
What the Score Doesn’t Tell You
- Technical writing — formal regulatory or engineering text can score lower because it shares structural characteristics with AI output.
- Government publications — policy documents often score in the mixed range due to their formal, templated structure.
- Short texts — passages under 200 characters are skipped; insufficient signal for reliable analysis.
- Adversarial evasion — paraphrasing attacks can reduce detection accuracy. This system is designed for editorial transparency, not forensic attribution.