"We needed a way to take AI use away from accusation and police action and make something that rewarded...actual writing."
A nerdy, no-holds-barred conversation with the founder of AI detector Verify My Writing
Update on May 18, 2026: A few weeks after I published the post below, Substacker Ron Charles shared a brief, nifty item about Verify My Writing and “certified human” seals, then added this coda:
P.S. The entire item you’ve just read was written by AI. To create it, I submitted a transcript of my phone interview with Derek Newton to ChatGPT and told it to generate a 500-word story in my voice. I then sent that AI slop to VerifyMyWriting.com and received an Authenticity Certificate for “100% Human Based.” The whole process — including my payment of $10 — took 90 seconds, or about the same amount of time it took for American journalism to die.
So that’s…distressing. I can’t repeat his experiment because I won’t use ChatGPT under any circumstances, but I’m back to believing that unless a written piece has an AI prompt left in it, or the author admitted to using an LLM, you simply cannot be sure. That’s why I think it’s more important than ever to communicate your own stance on AI usage.
I’m leaving the interview below up for transparency. If anything, this all proves my point: AI is leading to massive distrust between creators, patrons, and their industries, and that is very, very bad. —Andi
Derek Newton had good reason to dislike me.
My NYT op-ed about how AI is eroding trust between readers, authors, and the publishing industry linked to his company, Verify My Writing, alongside some other online AI detectors in a paragraph about just how fallible and, in most cases, sketchy the AI-dars1 are.
Instead of lashing out, he reached out to start a dialogue about how, yes, many of these LLM checkers are worthless—but that makes the good ones more crucial than ever. I’ll admit I was wary, but over multiple conversations, Derek has convinced me that—to my surprise and relief—Verify My Writing is the real deal…and I believe that using it is not evil or hypocritical for anti-AI activists like myself.
Over g-chat, Derek and I dug into how the technology works under the hood, why AI-checking-AI won’t necessarily lead to a nightmarish arms race, whether the systems retain your writing or use it for training, and more.
Me: Hello! Thanks so much for chatting with me today
Derek: My pleasure. This is cool.
So we connected after I wrote about AI detectors in my NYT op-ed. Can you start by sharing who you are and what Verify My Writing is all about?
Absolutely. I’m Derek Newton. I’m a writer (journalist) who’s covered education and fraud. And I started Verify My Writing as a way for writers to get our work scanned as verified as authentic and human-written, not AI.
As a google search will tell you, there are now dozens of AI “checkers” available online. Can you say more about how you developed this one and what makes it different?
Yes. There are plenty of free AI-checkers. Most of them are bad. Or worse, outright fraud engines trying to sell people things to fix problems they don’t have. That’s why so many people say “AI detection” does not work. But, as someone who’s covered this space from the start, written about it dozens of time, the research is clear that, while many AI-checkers are junk, a few are good, and a few of those are exceptional.
We use those that have been independently proven as exceptional as the basis for our scans and scores.
OK, so you were like, “We need an AI detector that is actually reliable.” How did you go about developing that?
We do need that. And, more importantly, I think we needed a way to take AI use away from accusation and police action and make something that rewarded the work and craft of actual writing.
So, we contacted the very best companies in this space and signed them up to run our technology.
When we were on a radio panel together, you stated that VMW’s accuracy is incredibly high—98%, something like that?
Yes. It depends on the document length, style, etc. But the underlying accuracy is 1 in 10,000 to 1 in 25,000 which is something like 99.9999%
Maybe just three nines. I’m a writer.
Forgive the dumb-sounding question, but: How can you be sure? When prompted to write “in the style of Andrea Bartz,” ChatGPT gives me sentences that sound just like something I could’ve written. Where does the confidence come from?
There have been many -- maybe two dozen -- university studies to test AI detection accuracy, even under so-called evasion attacks (aggressive editing) and the results hold up.
I had another writer who trained an AI bot to write in his style. Our system flagged the bot-created writing every time.
The other thing to keep in mind that, in some of those same studies, humans are not especially good at spotting AI writing. Even when they are told to look for it.
If you use AI often, you can spot it more reliably. But even then, not at much above 70%.
Are the university studies peer-reviewed or prepublication, and are they sponsored by the makers of the technology? With regard to Pangram (one of the AI-detecting LLMs), the Wall Street Journal wrote: “The company’s website boasts that its product has been ‘reviewed as the proven, most reliable and most accurate AI detection tool in the market by third parties including University of Maryland.’ But the Maryland study is more self-promotion than neutral review. Four of its seven authors are Pangram employees, including Mr. Spero and fellow co-founder Bradley Emi.”
Both.
The WSJ has a viewpoint - that their editorials were not AI created or AI assisted. That’s fine.
And one of the studies in question was co-written by the some of the data scientists at Pangram. But there are many others from other universities, in no way connected to the provider. Either way, they all show pretty much the same thing.
Of course, the Journal only mentioned that one study. Like I mentioned, they have a view.
Got it! Let’s get nerdy about Verify My Writing—under the hood, what is it actually doing when you input a passage of text?
When a writer submits text, we engage with one of our vetted and approved tech providers (Pangram is one of those) and we scan the text. This returns a score, which we calculate as a human score and not an AI score.
In about one in ten of our scans, we engage more than one system, so as to keep an accuracy and validity check on our scores.
Then, we share that score with the writer or person who submitted it. And, if they want, they can get a Certification from us, with that score.
The Certification can’t be altered or faked and is tied to the individual document scanned on that day and time.
The tech providers, including Pangram, are themselves LLMs trained to look for patterns, correct?
Yes. And no. They are closer to machine learning systems that use LLMs as reference points in scoring. They are not LLMs that use that data to create anything. But the machine learning is the operational part. And, with all our partners, we checked that the LLMs they use are all public source or licensed.
These are not the same LLMs as those being built by OpenAI or others.
An analogy - a good AI detection system is looking for the math fingerprints left behind by an AI bot. It does not need a large dataset of existing fingerprints to do that. At most, it needs a set of AI’s fingerprints. It may reference an LLM in training to learn how to spot a fingerprint from a smudge (a human written work). But in day-to-day use, it’s not using a standard, human-populated LLM very often.
I hope that makes sense.
It needs a large body of AI-created text more than a body of not-AI text. May have been easier way to say that.
“The LLM companies can ‘watermark’ their text so it can be spotted very easily. Making ‘was it AI’ conversations clear. But they refuse to. Bad for business.”
Okay, that makes sense, thanks! I have a few questions related to this...
Fire away.
First, machine learning is a subset of AI that’s all about learning patterns from existing data. But it is a kind of AI. So we have AI checking AI. Couldn’t this just lead to an arms race where generative AI models learn what Pangram et al are looking for and get better at changing or hiding it? In other words, couldn’t LLMs start scrubbing away their own fingerprints?
Good question. Based on what I know, no.
So long as a math system (an algorithm) is making it, it will always be possible to see that math behind the scenes.
It may get harder. But it will remain possible. So long as it’s being created with rules.
No rules writing, human writing, is what makes it different.
And, in fact, the LLMs and AI companies could be making their text harder to spot. But they are not. The most recent data shows that AI-produced text is becoming more average, more predictable, easier to spot.
Interesting! My guess is they care way less about individual plagiarism accusations and more about taking over the world, Pinky and the Brain style, haha
Something like that.
They are training to be able to talk like anything -- chemists, pilots, skaters. They are less concerned about general output these days.
Fascinating! I’m resisting the temptation to go into that rabbit hole out of respect for your time and all my additional questions :)
Resist!'
Next one: When I drop my writing into VMW and ask it to scan and analyze it, am I—behind the scenes—enriching an LLM company (like OpenAI) or a company that makes chips to power AI (like NVIDIA)?
Nope. We do not share text with anyone and neither do our partners.
To back up -- the LLM companies can “watermark” their text so it can be spotted very easily. Making “was it AI” conversations clear. But they refuse to. Bad for business.
I don’t know about the chips question -- we use AWS. So, maybe. Just not sure. But the LLMs, that is a no.
OK calling in my machine-learning specialist girlfriend to make sure I word my next question correctly, one sec :)
Love it. Though it may exceed my ability to answer. I’m a journalist, not an AI engineer.
Haha totally fair, I just don’t want to muddy the waters with an imprecise or incorrect question
Good deal. I’m game, either way.
Actually, let me just go back to something you said earlier: “with all our partners, we checked that the LLMs they use are all public source or licensed. These are not the same LLMs as those being built by OpenAI or others.”
sure
So further upstream, Verify My Writing/its partners isn’t using a paid LLM from one of the big 5 companies (OpenAI et al)?
Easy one. No.
Got it. Can you explain more clearly how the system computes a score? Like, is it comparing your text to what the AI version would be and telling you whether it matches or not?
(Machine-learning girlfriend puts it this way, which makes less sense to me but might mean something to you: At prediction time, what happens? What type of machine learning technology is used to assign that score—statistical models or LLMs?)
I can try. Starting backwards, I am pretty sure the answer is statistical models.
More broadly, the system looks at a number of features in the text -- sentence structure, word choice, document structure, patterns in any of those. And it has learned what AI would “write.” So it can say, with high confidence, this is very likely to be AI-created. This, over here, is not. One sentence [of AI-generated text] in a full book manuscript probably still gets a 100% human score. 99.9% at the worst.
A thing to keep in mind with all these is that there no point, in my view, in parsing a 99.4% versus a 99.0%. There’s not much information there. What we want to get to is 20% versus 60% versus 100%.
If something gets a score of, say, 80%, does that mean 80% confidence it’s human-written, or that four-fifths of it looks human and one-fifth looks like AI slop?
The latter. 80% of it, we are confident, is human. 20%, we have (strong) doubts.
We’d call that mixed, in most cases. And it’s up to whomever (editor, reader, agent) to say if that’s an issue for them.
Thank you so much. Newsletter subscriber Wendy Hawkes had this question: “Is there a risk of documents that have gone through their verification process being flagged as AI if a secondary source is used? It may get to the point where publishers run 2 separate verifications--get a second opinion. Does his software leave any trace that would be subsequently diagnosed as AI?”
(I take that to mean that no one’s changing the text or clicking “humanize,” but rather the text has run through another detector first, and someone’s inputting that output into VMW)
Good question. No.
Our systems don’t retain the writing. So it can’t know if something was scanned once or a hundred times. And even if it did, it’s only looking at the words on the page, making a calculation as to whether those words were produced by AI or not.
Multiple scans won’t cause something to be seen as more (or less) likely AI.
The bigger risk is that people use junk detection systems for second (or first) opinions. If a bad detector says something is 90% AI, I’m not sure where you go from there.
That “humanize” button haunts my dreams. Could using a tool like Grammarly or a word processor’s built-in grammar check—and accepting sentence-level suggestions—lead to it getting flagged as AI? Like, where’s the line, or how sensitive is VMW to “light-touch” AI use?
Yes. Grammarly and other AI-powered text replacement can (and does) set off good AI detection. If Grammarly replaces your text, the text is AI-created.
That’s where we are again, was that two or three times in a novel? Not much to worry about, score-wise. You’d likely get a 99.7 or something like that. Or did it swap out a full paragraph or chapter? That’s going to show up
If Grammarly flags a grammar error (like it used to) changing that won’t matter. We can’t know if Grammarly flagged it or if your mother/editor did.
OK, that makes sense! Thank you. Finally, I want to ask about behavior-tracking AI detectors like Puddin’, which purport to look for human patterns in HOW the text makes it way onto the page.
Sure. I know a little about these.
At first glance, this sounds like a great idea! ChatGPT generates words in one belch, while I stare out the window, pause to look things up, etc. Easy way to distinguish between the two. But then I wondered if (to go back to the arms race idea), an LLM could learn to make its word-generation LOOK more human to outsmart these detectors.
Basically, I’m curious what your opinion is on this verification method, as someone in the same line of work!
Yes. There are already dozens of autotypers designed to fool these systems and Google Docs.
They can take text (GPT or otherwise) and key it in to any system, at a set rate, with mistakes, and corrections. All looks very organic. Pretty common in college now.
The only way these work is if the detection system has a sample of your human typing, typed in their system, and then you, from then on, type in that system (browser tab, software, etc). Those are very accurate. But ones without knowing how YOU type, are very easy to bypass.
When it comes to typing, human is easy to fake. Typing like you type, is hard to fake. In other words. But the second version requires a willingness to participate.
Right, then you’re asking a company to constantly surveil your writing...
100%
Costs and benefits. Few good trades out here these days.
What a time to be alive! Okay wow, this has been SUPER enlightening and fascinating, and I so appreciate your patience with all my annoying questions. Is there anything else you want to mention or think the author community should know (about Verify My Writing or AI in the book world in general)?
Thank you! This is great.
Just that it’s tough out here right now and we want to help the real writers who really do the work. That’s our mission and our model.
And on that, we are fully aligned!
No surprise at all. Thanks again.
Derek Newton is a writer whose work has appeared in The Atlantic, Forbes, NBC, USA Today, and many other outlets. In addition to writing, Derek is a leader in integrity and fraud. He was the keynote speaker at the 2025 conference of the International Center for Academic Integrity and writes a regular newsletter on cheating and authentic work, The Cheat Sheet, which has published 400 Issues and has some 5,000 subscribers.
Further reading:
I just made that up, like “gaydar”—what do we think?





What a gentleman to engage in this discussion with you, and to offer such clarity in what has become (increasingly becoming) a murky development. Thanks, also, for including my question (and the shout-out, hello!). So many cogs in this AI Ferris wheel. It's heartening to hear VMW doesn't keep text, nor use LLM trainers like OpenAI, so they're not contributing to the festering stew (so it seems). It was also fascinating to get Derek's take on behavior locAI-tors (hmm, think I like AI-dar better) like Puddin. Great interview, Andrea.