Comment Generator logoComment Generator
2026-08-098 min readFelix Melchnerby Felix Melchner

What Actually Makes a Comment Read as AI-Generated

Comment Generator header graphic breaking down the research on what makes text read as AI-generated

"This reads like AI" gets said constantly and defined rarely. We covered the mechanics of LinkedIn's new reporting button in LinkedIn's AI slop button: what it actually flags. This post is about the other half: what does the actual research say makes a piece of writing detectable as AI-generated in the first place, and is any of it something a reader can reliably catch?

The honest answer is more specific, and more useful, than "you can just tell." A handful of real studies have measured this, and they point at word choice and sentence variety more than anything else, while also showing that most of what people believe about detecting AI text, including their own ability to do it, doesn't hold up.

How much AI content is actually out there?

The most-cited number comes from Pangram Labs, an AI-detection company that scanned just over one million posts through an opt-in Chrome extension between April and July 2026, and found more than 40% of long-form LinkedIn posts came back as fully AI-generated, with LinkedIn accounting for nearly two-thirds of all the AI content flagged across every platform sampled. Worth stating plainly: the sample is self-selected (people who installed an AI-detection extension and chose to scan specific posts), and the company selling that detection published the numbers. That doesn't make the figures wrong, but it does mean the true base rate across all of LinkedIn is unknown and probably lower than a self-selected sample of the most suspicious-looking posts would suggest.

What's better sourced is what LinkedIn itself says it's doing about it: blocking hundreds of thousands of automated spam comment attempts a day, and rolling out a "seems like AI slop" reporting option in July 2026, alongside quietly retiring its own AI post-enhancement tool in favor of a conservative proofreader that fixes grammar without rewriting anyone's voice. That's a company acting on a problem it considers real, even without a fully audited number attached to it.

The vocabulary tell is real and it's measurable

The clearest, best-verified finding in this entire area comes from biomedical publishing rather than social media. A 2025 study in Science Advances analyzed over 15 million PubMed abstracts published between 2010 and 2024 and measured which words suddenly spiked in frequency after large language models became widely available, comparing actual 2024 usage against a projection based on the 2021 to 2022 trend. The excess was large and specific: "delves" appeared 28 times more often than the pre-AI trend predicted, "underscores" 13.8 times more often, "showcasing" 10.7 times more often. The same method estimated that at least 13.5% of all 2024 abstracts, roughly 200,000 papers, had been processed with an LLM somewhere in their drafting.

That's academic writing, not comment sections, but the mechanism is the same one that makes a comment sound synthetic: certain words don't occur naturally as often as language models reach for them. If you notice a comment leaning on "delve," "underscore," "showcase," or their close relatives, you're not imagining a pattern. It's been measured at scale.

Em dashes are a real, quantified signal, just not for short text

This one gets repeated as a joke and it's actually backed by data, with a caveat that matters. Research tracking medical preprints found em dash usage in discussion sections rose from 4.23% before ChatGPT to 11.58% after, and a separate study of U.S. congressional press releases found unspaced em dash usage roughly doubled during 2025 alone. Both are real, population-level shifts, and both are measured in long-form prose: preprints and press releases, not one-paragraph comments. An em dash in a three-sentence comment is a much weaker signal than the same punctuation in six hundred words of prose, simply because there's less text for the pattern to show up in. (For what it's worth, we don't use them in anything we publish either, on principle rather than to dodge a detector.)

The tell that actually matters for something as short as a comment

A 2023 study that looked specifically at AI-generated news comments, rather than essays or press releases, found that human comments consistently showed higher lexical diversity than the GPT-generated ones the researchers produced for comparison, and that a fine-tuned classifier could separate human from AI comments reliably on that basis alone. Lexical diversity, in plain terms, means not reaching for the same handful of words and sentence shapes every time. A generic AI comment often isn't wrong or even badly written. It's just narrower than a person's actual vocabulary tends to be when they're reacting to something specific.

There's a counterintuitive wrinkle worth knowing: a separate 2025 study found that newer, more capable models sometimes score higher on diversity metrics while still producing text that reads as less human overall, meaning the newest models aren't automatically solving this by getting more "diverse" on paper. What seems to matter more than raw vocabulary range is whether the comment contains something specific to the post it's replying to, which is a harder thing for a generic prompt to fake than word variety alone.

Can people actually tell? Not as well as they think

A 2025 study recruited 254 native speakers and had each of them judge 20 texts as human or AI-written. Without any feedback on whether they'd guessed correctly, accuracy came out to 55.4%, barely above chance. With feedback after each guess, it rose to 65.1%. The detail worth remembering: people made the most mistakes exactly when they felt most confident in their judgment, performing worse than random guessing at their highest stated confidence level. Confidence and accuracy were not just uncorrelated. They ran in the wrong direction.

Automated detectors have their own serious problem

If human judgment is unreliable, automated detectors aren't a clean fix either, and they fail in a specific, documented direction. A Stanford study tested seven commercial AI detectors, including tools from Originality.AI and OpenAI itself, against 91 real TOEFL essays written by non-native English speakers and 88 essays written by native-speaking eighth graders. The detectors misclassified the non-native essays as AI-written at an average rate of 61.22%. Eighteen of the ninety-one were flagged as AI by all seven detectors at once. The native-speaker essays were almost never misclassified. This is the same finding we pointed to in how to comment in English when it isn't your first language: the tools built to catch AI writing carry a real, measured bias against exactly the writers who are least likely to be using AI to cheat and most likely to be penalized by a false flag.

The part that should worry you more: trust drops even when quality doesn't

This is the finding worth sitting with. A study of nearly 1,500 participants had them read real news articles, some labeled as AI-assisted and some not, then rate the source on trustworthiness, accuracy, and fairness. The trust score dropped in a way that was statistically solid: 5.9 out of 11 for the unlabeled version versus 5.5 for the AI-labeled one. But the accuracy and fairness ratings barely moved at all, statistically indistinguishable between the two groups. People weren't judging the writing as worse. They were trusting the source less simply because AI was named as part of the process, independent of whether the output held up.

A separate, larger study of nearly 4,000 UK adults found the same direction of effect at a smaller magnitude: an AI label reduced perceived accuracy modestly, with no broader knock-on effects on people's views or policy positions. Put the two findings together and the honest read is that the penalty for AI involvement is reputational, not really about whether the resulting text is good.

What this means for a comment you're about to post

Put the findings together and a specific, non-obvious answer falls out. The wording tells ("delve," excess em dashes, generic phrasing) are real but weak signals in something as short as a comment, and low lexical diversity plus a lack of anything specific to the actual post is the stronger one. Detectors are worse than people assume and biased in a way that particularly hurts non-native English writers. And because the trust penalty attaches to the label "AI-assisted" rather than to whether the output is actually any good, a well-reviewed AI-drafted comment and a hastily typed human one carry genuinely different reputational risk, even when they read identically on the page.

That's the entire logic behind building Comment Generator around a caption-specific draft that a person reads before it goes anywhere: specificity is the harder thing to fake, and review is the step that exists precisely because the trust risk sits with disclosure and process, not with word choice alone. For the fuller argument on why that review step isn't optional, see why AI-drafted comments still need a human in the loop, and for what a comment worth posting actually looks like once you've read the draft, how to write genuine LinkedIn comments has the structure.

Felix Melchner

Felix Melchner

I built Comment Generator so commenting genuinely on Instagram doesn’t take forever. I also run RecentReborn, which surfaces the newest posts in your niche for early engagement.