Tools That Erase AI's Writing Tics Are Multiplying
If you need human-sounding text, why not just have a human write it from the start?
BusinessA wave of tools built to erase “AI voice” is showing up all at once
Four things crossed my radar recently, and they were all pointing in the same direction.
On Hugging Face, a model called Hemmingway-1 launched under the banner of “AI that writes like a person.” On GitHub, a Claude skill called im-not-ai — which claims to scrub the telltale signs of AI-generated Korean text — had passed 5,600 stars as of September 22. On the front-end side, developers are using something called Taste Skill to stop AI coding tools from spitting out the same generic-looking screens. And on Facebook, a post made the rounds describing how someone had taken an English-language spec originally written for aircraft maintenance documentation and repurposed it as a prompt for writing Korean-language reports.
The point of intervention differs in each case. Some retrain the model itself. Some edit text after it’s already written. Some impose rules before a single word gets generated. But what they’re all chasing is roughly the same thing: writing that doesn’t give away that AI wrote it.
Looking at this list, though, one question comes to mind first. If you need something that reads like a human wrote it, why not just have a human write it from the start? Today I want to dig into what each of these tools has decided “human-like” actually means — and answer that question along the way.
Three Fronts Are Being Worked at Once
Hemmingway-1, a Model Retrained
Hemmingway-1 is a model built through additional training1 on top of Qwen3.8-27B, with 27 billion parameters2. Its weights are open, and it’s licensed under Apache-2.0, so it can be used commercially too.
The problem it targets is quite specific. The developers use the example of asking for help drafting a text message to a landlord about a broken boiler. A typical model hands back three candidate texts, complete with a greeting and commentary on each option — Hemmingway-1, they say, just gives you the text. By the developer’s own measurements, Fable 5, GLM-5.3, and Kimi K3 buried the sentence you’d actually send inside explanations and options more than nine times out of ten.
The claim that it “sounds human” is also backed by numbers. Across 80 real requests, its answers were paired blind against those of other models, with authorship hidden, and compared. On the share judged to be human-written, it led the next-best model by 26 points. For awkward, hard-to-phrase requests specifically, GPT-6 Astra scored 9% against Hemmingway-1’s 72%.
But the fine print matters here. This evaluation set was built and run by the developer itself, and the judging was done by another AI model, not by humans — a caveat the developer disclosed upfront. There’s also a warning that it’s an English-centric model that can sound confident even when it’s wrong, so it shouldn’t be used for medical, legal, or financial judgments.
im-not-ai, Which Edits Finished Drafts
im-not-ai is a copyediting tool focused on Korean. You install it as a skill3 into coding agents like Claude Code. It sorts the telltale marks AI leaves in Korean writing into 10 broad categories and 70 detailed patterns, scores each by severity, flags them, and fixes them. Typical targets include translation-ese4 phrases like “via ~” or “in regard to ~”, clichés like “in conclusion” or “this carries significant implications”, and mechanical “first, second, third” listings.
What’s worth noting is the boundary it sets for itself. It doesn’t touch facts, claims, figures, proper nouns, or direct quotations. If the change rate against the original text exceeds 30%, it issues a warning; past 50%, it stops entirely. If it breaks every “not A but B”-style parallel construction and erases the writer’s voice along with it, that too counts as a failure. It’s designed to fix only tone, never substance.
ASD-STE100 and Taste Skill, Which Set Rules Before You Write
ASD-STE100 is a controlled-language5 specification first published in 1986. Europe’s aerospace industry created it at the request of airlines. If a mechanic who isn’t a native English speaker misreads a manual, someone can get hurt — so the language itself was narrowed. The current edition, the 9th, came out in January 2025, and starting with this edition it was renamed from a specification to an international standard. It consists of 53 writing rules and roughly 900 approved words, with each word permitted only one meaning and one part of speech. It never strings four or more nouns together in a row, and it restricts usable verb forms to just a handful.
I’ve borrowed this standard for my own work too. When I drew up the writing rules for INLEVEL9 Skills, built for Korean-language Notion users, I referenced ASD-STE100’s approach. After years of using Notion, I’d learned that it’s more effective to fix the rules for vocabulary, sentences, and document structure before you start writing. Whether I was writing myself or handing the job to Notion AI, having clear standards made the results more consistent and easier to review.
So I set it up so that each block holds only one piece of content, the same concept always uses the same wording, and confirmed decisions, proposals, and undecided matters stay clearly distinct. I also had it check, at the end, whether verified facts had gotten mixed up with the writer’s own interpretations. And I added a rule that values missing from the source material — like the person in charge, the deadline, or completion status — should be left as undetermined rather than guessed and filled in.
Taste Skill applies the same idea to screen design. It bundles rules and a list of prohibitions into a skill file so that AI coding agents don’t churn out template-looking screens, and it requires output to pass a checklist before it ships. It calls itself an “anti-slop”6 framework.
The “Human-Like” Ruler Is Shaky
There’s a common thread in how these four tools measure human-likeness: since they can’t measure it directly, they define it in reverse — as the absence of AI tells.
Hemmingway-1’s “human-like” verdict is the rate at which its detection model mistakes text for human writing. im-not-ai’s standard is a list of patterns that show up often in AI writing. The weakness of this approach came to light when im-not-ai underwent outside scrutiny.
In August, the data-quality firm Pebblous pulled apart im-not-ai’s public code and found that a human-written essay from 2020 — before ChatGPT even existed — got flagged at the highest risk tier. The cause was the writer’s comma habit: they placed a comma after connective endings 83.3% of the time, and this single metric dragged the whole verdict along with it. When the development team re-ran 532 human-written pieces, 285 of them — 53.6% — came back flagged as high AI risk.
The team’s response was fast. They capped how much any single metric family could influence the outcome, changed the rule so that two different metric families now have to trigger together before something counts as high-risk, and built separate baselines for the column and report genres. Running the same 532 pieces afterward produced zero high-risk flags. They also noted a limitation in the README: since the baseline was built and tested on the same set of writing, it’s still unclear whether it holds up on new text. But the fact that they published the counterexample and fixed it in the open is, if anything, what makes this project easier to trust.
Still, what this episode revealed doesn’t go away. A good number of the habits classified as “AI tells” were originally human habits. As I discussed in issue #213, chatbot phrasing was learned from the formal register people use in corporate reports and academic papers. And the fingerprint keeps moving. Words that once symbolized AI writing have nearly vanished from the latest models, replaced by other expressions. Today’s list of tells looks less like a timeless catalogue and more like a snapshot of one generation of models’ quirks.
That raises one worry. If thousands of people install the same skill, the sentences that manage to dodge the banned-word list may start converging on each other. If a rule meant to block an obvious, clichéd pattern spreads widely enough, the pattern produced by following that rule could itself become the next generation’s cliché. Human-likeness defined by a blacklist is, by nature, a moving target you can never stop chasing.
What people wanted wasn’t ‘human’ — it was less effort
This time, let’s look at what these tools actually fix. Hemmingway-1 fixes sentences buried in over-explanation, im-not-ai fixes translated-sounding phrasing and hard-to-read sentences, ASD-STE100 fixes instructions that could be misread, and Taste Skill fixes screens that look like something you’ve seen before. None of them checks who wrote the text. All of them are aimed at reducing the effort the reader has to spend.
In that sense, ASD-STE100 is a bit different. It wasn’t designed to make writing look human. It was designed to keep humans from misreading manuals written by humans. It doesn’t place the standard for good writing on the writer trying to imitate it — it places that standard on the reader who could get hurt by a misreading. When you build your rules around the reader, most of the “AI tells” fall away on their own, without ever being targeted directly.
So “be more human” is a slightly imprecise instruction. What people actually wanted, most of the time, was writing that doesn’t leave the reader lost.
So, Should Humans Write the First Draft After All
Let’s return to the question in the subtitle. Line up the manuals for all four tools, and they all leave the same space blank.
im-not-ai puts the principle of not changing a single character of content right at the top. Hemmingway-1 warns you, in its own documentation, that it can state wrong things with total confidence. And the unresolved clause in my own Notion rule, mentioned earlier, is no different — it’s a rule that only works if someone is actually checking whether the material has value.
No tool fills in what to claim, what to verify, or how far responsibility extends. What tools handle is the surface that holds content already decided by those choices. So the answer to the question is only half “yes.” You can outsource the sentences. The claims still have to be written by a human, from scratch.
Reversing that order causes trouble. Take a draft nobody has actually judged, strip out its AI tells, and what you get is prose that reads smoothly but has no one home inside it. Until now, awkward translationese or a predictably bland conclusion served as a signal to readers: “this piece may not have been properly checked by anyone.” A polishing tool erases that signal too. The false-positive problem I covered in Issue 195 was human writing getting misread as AI writing. This time, it runs the other way: writing no human ever checked ends up looking like a human wrote it.
Oswarld’s Lens
I often say that not every task needs the smartest possible model. Hemmingway-1 is a good case in point. A 27-billion-parameter model reportedly outperformed much larger models on the narrow task of everyday messaging. Even granting that this comes from the developer’s own benchmarks, I think it points in the right direction: it shows where a model built for a narrow task can win.
Where I part ways is with the design goal of “human-like writing.” That goal is defined relative to what counts as “sounding like AI,” and that boundary keeps moving every time a new model comes out. Once the standard itself is unstable, you end up with detectors and polishers chasing each other in circles. If it were up to me, I’d anchor the target on the reader instead — who’s reading this, what they need to do after reading it, and where they might misread it. I think that’s exactly why ASD-STE100 has survived 40 years of revisions — because it anchored its standard there. That’s also why INLEVEL9 Skills drew on this specification. What I wanted wasn’t a document that looks human-written, but one that’s clearer and produces consistent work.
So when people ask, “why not just have a person write it?”, here’s how I’d answer: the one thing a human has to supply from the start is judgment. A draft with judgment gets better the more a tool polishes it. A draft without judgment just gets its emptiness better hidden the more you polish it.
Closing
The tools we looked at this time have gotten quite sophisticated at erasing the telltale signs of AI writing. Which means the warning signal those signs used to give readers is fading too. Going forward, the clues we use to decide whether to trust a piece of writing will depend less on its style and more on what it claims and what evidence backs it up.
💬 Reader, what’s the moment where you stop reading an AI-written piece? Was it the tone, or the content? Let me know in the comments.
📨 If you have a colleague who drafts reports or messages with AI, share this piece with them. It could be a good starting point for talking about what to delegate and what to write yourselves.
Already a subscriber? Sign in to join the conversation
References & Further Reading
- Altworld/Hemmingway-1 Model Card (Hugging Face)
- epoko77-ai/im-not-ai: Korean AI-tell remover (GitHub)
- Pebblous, Korean AI Humanizer Teardown (2026-08-20)
- Taste Skill: The Anti-Slop Frontend Framework
- ASD-STE100, About STE
- ASD-STE100, FAQ
- Issue No. 213: The Company Reports That Taught the Chatbot How to Talk
- Issue No. 195: Did AI Write the American Declaration of Independence?
- Issue No. 191: Why Claude’s Writing Always Leaves Fingerprints
📝 Glossary
Footnotes
-
Fine-tuning: further training an already-trained model on data for a specific purpose, to steer its behavior in a desired direction. ↩
-
Parameters: the internal numerical values a model adjusts as it learns. More parameters generally means a bigger model — and more computation needed to run it. ↩
-
Skill: a bundle of instructions and scripts an AI agent reads in order to perform a specific task, typically organized around a file called SKILL.md. ↩
-
Translationese: phrasing that carries over the sentence structure of a foreign language — especially English — so directly that it reads awkwardly in Korean. ↩

Your take shapes the next issue
What resonated most in this issue, or where has your experience been different?