← All transcripts

The cure for AI slop is a 1986 aircraft manual Transcript, AI Summary & Key Points

Ege Vusal Chelebi · Jul 24, 2026 · People & Blogs · 16:41 · EN

Watch on YouTube

Answer

Yes — for form. The 1986 aircraft standard (ASD-STE100) cures AI slop better than anything else tested, cutting mechanical slop violations by 74% on Claude and 50% on GPT 5.5, but it only makes writing unsloppy, not good, and should stay away from anything needing a voice.

AI Summary

ASD-STE100, a 1986 simplified technical English standard built for aircraft maintenance manuals, cures the mechanical form of AI slop better than word bans or Orwell's rules. A 434-page, 53-rule standard with a ~900-word one-meaning dictionary was distilled into an agent skill and a linter, then tested across 6 real developer writing tasks x 4 conditions x 2 models. On Claude, the linter score (violations per 100 words) fell from a 4.36 baseline to 1.12 with STE (−74%), versus 4.21 for the banned-words list (−3%) and 2.48 for Orwell's rules (−43%). On GPT 5.5, STE still cut slop by 50% and the banned-words list also worked (−40%), tying STE — so the 3% figure was a Claude quirk. The surviving finding: giving a model a real, machine-checkable writing system cuts slop by half or more on every model tried. STE fixes form, not substance — it cannot make writing true or give it a voice — so it belongs in docs, pull requests, error messages and agent output, not marketing copy.

Key Points

  • The viral claim: a thread told models to 'always use ASD-STE100', a 53-rule, ~900-word-dictionary international standard first released in 1986 so aircraft mechanics never misread a repair manual
  • Banning words is whack-a-mole: the model never chose those words on purpose — without a writing system it defaults to the average of the internet
  • AI slop collapses into six nameable habits: synonym rotation, hedging stacks, nominalization (verbs frozen into nouns), marketing adjectives, run-on sentences, and chatty phrasal verbs
  • The slop list maps rule-by-rule onto STE: rule 1.11 kills synonym rotation, 3.4 kills hedging, 3.7 kills nominalization, the dictionary excludes marketing adjectives like 'seamless', rules 5.1/6.3 cap sentences at 20–25 words and ban semicolons, rule 9.3 kills phrasal verbs
  • The current standard is issue 9, January 2025, downloadable free from the official site; 64% of people requesting it are now outside aerospace and defense
  • Human evidence: a 1996 study of 175 aircraft technicians found comprehension rose from 76% to 86%, and from 69% to 87% for non-native speakers; a 2007 Microsoft study of 520 sentences in four languages found controlled versions beat normal ones everywhere (translation gains were real but small, ~0.1 of a point on a 4-point scale)
  • The test: 6 developer writing tasks (readme, pull request description, API docs, error message, getting started guide, deprecation notice) x 4 conditions (baseline, banned words, Orwell's six rules, STE), scored by a custom linter counting violations per 100 words
  • Claude results: baseline 4.36, banned words 4.21 (−3%), Orwell 2.48 (−43%), STE 1.12 (−74%)

AI in practice

Used for

Business ideas

Instead of banning AI-tell words one at a time, give the model a complete writing system — a distilled version of ASD-STE100, the 1986 aerospace Simplified Technical English standard — so slop dies by design: one word per meaning kills synonym rotation, hard sentence caps kill run-ons, the 900-word approved dictionary kills marketing adjectives.

For
Developers and technical writers who ship AI-generated documentation, pull request descriptions, error messages and agent output
Solves
Banned-word lists are whack-a-mole: banning em dashes just produces slop without em dashes. The model defaults to the average of the internet because it was never given a writing system it can check itself against.
  • Boeing's Simplified English checker with a 350-rule parser, shipping since about 1990 — machine-checked writing is a solved, battle-tested category

A short, machine-checkable rule file — not the full 434-page standard, even a 10-line config — that an LLM or linter tests output against. The transcript's own framing: the standard is one option, but a short checkable rule set is a product, and a small one already exists with adopters.

For
Teams and individuals who generate technical prose with AI and want enforceable output quality without a heavyweight standard
Solves
The full standard is too heavy and partially copyrighted; a blacklisted-word approach is unreliable. A compact rule set captures most of the benefit at a fraction of the cost.
  • A 10-line rule config cited in the thread as '90% of the benefit' of the full standard
🔒  Build steps and tools for 2 ideas. Unlock

Tools & resources

2 items

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of The cure for AI slop is a 1986 aircraft manual — Ege Vusal Chelebi (16:41). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Ege Vusal Chelebi. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:06 The same prompt, same model, two answers. Answer one says designed to slot into your existing stack with minimal friction and no vendor lock-in. And the answer two says a normal cache matches requests by exact text. A small change in wording then causes a cache miss. And one of these is slop. And two million people just watched a thread claim that the cure is a writing standard built for aircraft mechanics in 1986.

00:31 And that sounded just absurd enough that I had to shake in myself, too. So, I got the real standard, all 434 pages of it. I got the rules, gave them to my model, and ran the test nobody had actually run before. And here's the output. But first, the story, because the story is the reason this video exists, too. A few days ago, a post went viral and an account called Joe Briscoe asking in plain words, "How do I make AI write technical documentation that doesn't sound like AI?"

01:02 And it got over two million views. And then a writer called Vox picked it up and made the sharper point. He said, "Stop telling Cloudy no em dashes. Stop banning the word delve." And this is whack-a-mole and you lose because the model was never choosing those words on purpose. You never gave it a writing system, so it defaults to the average of the internet.

01:24 And the average of the internet is slop. His fix was George Orwell's six rules for clear writing from 1946. And then Mike Hostetler, a man whose biological says chief agent officer, quote tweeted the better answer. Just tell the model to always use ASD-STE100, simplified technical English. A real published international standard. A 53-rule and 900-word dictionary built so that an aircraft mechanic never misreads a repair manual.

01:56 Because a misread repair manual kills people. And that was the moment I fell down the rabbit hole. A mum property decides the whole thing. So, keep it in mind the entire way. The rules that matter are machine checkable. A dumb script can verify them, and no taste is required. Before the fix, we have to be precise about the disease. Sounds like AI is a feeling, not a definition.

02:19 And you cannot fix a feeling. So, the first thing I did was to try to pinch slope down mechanically. And everything I kept running into collapses into six habits. You already know every one of them. One, synonym rotation. The text calls the same thing three names in one paragraph. The user, the customer, and the client. You are never sure they are the same one person.

02:47 Two, hedging. Literal help of verbs stack until nothing happens. It's important to note that this may potentially help to improve. Five verbs and zero verb. Three, verbs frozen into nouns. A perform an analysis of instead of analyze. And provides assistance instead of help. The grammar people call it nominalization. All you need to know is that it makes sentences longer and weaker.

03:17 Four, marketing adjectives. Seamless, robust, powerful, cutting edge. Words that claim quality instead of showing it. Five, [snorts] run-on sentences. Four ideas stitched together with m dashes and semicolons when they should be four sentences. Six, overly chatty phrasal verbs. A spin up, reach out, dive into. And here's why this list matters. Every one of these is a specific, nameable move.

03:49 And anything you can name, you can ban, too. Now, hold that list. We are about to check it out off rule by rule. So, what is this aircraft standard? Now, I was obviously not around in the late 1970s. So, this is a story the way my research tells it. Back then, European airspace had a deadly problem. Maintenance manuals were written in English. But, most of the mechanics reading them around the world were not native English speakers.

04:16 And one misread instruction on an aircraft is a funeral. So, industry built a controlled version of English. Not a style guide, a specification. And first released in 1986. The current version, issue nine, came out in January 2025. And you can download it for free from the official site. The full document is 434 pages. And I'll be honest with you, I didn't sit down and read every single page.

04:42 I dug through it and mined out the parts that matter. The rules, the dictionary, and the logic behind them. That took a while, but you will never have to open it. It has two parts. Part one is procedures. Part two is the part that genuinely surprised me. A dictionary of roughly 900 approved words. Where every word gets the exact one meaning and one grammatical job.

05:12 Now, in AST, fall only means to move down by gravity. You are not allowed to use it to mean decrease. And where English gives five words for one idea, AST keeps one and deletes the rest. You get start. You do not get begin, commence, initiate, or originate. Notice what that just did. Synonym rotation. Slope habit number one is dead by design already.

05:41 And one more number I found that tells you this thing escaped its cage. Today, 64% of the people are requesting the standard or outside aerospace and defense entirely. Now, watch the slop list get executed rule by rule. This mapping is the part of the research that made me sit up. A synonym rotation killed by rule 1.11. One name per thing plus the one meaning dictionary.

06:09 Heging stacks killed by rule 3.4. No complicated verb constructions. Frozen verbs killed by rule 3.7. Use a verb for an action, not a noun. Performant analysis is illegal. Analyze is the law. Marketing adjectives. They are simply not in a dictionary. Seamless is not an approved word. You cannot say it. Run-ons killed by hard caps. Rule 5.1. An instruction gets at most 20 words.

06:39 Rule 6.3. A descriptive sentence gets 25. And the semicolon is banned outright. Phrasal verbs killed by rule 9.3. You remove the panel. You do not take it off. And here's the punchline. A style was never designed for AI. It was designed 40 years ago to remove ambiguity for a stressed human reader. AI slop is just ambiguity with good posture. The aerospace industry has spent four decades engineering the cure but for a different patient.

07:13 Let me read you the real thing, not a toy example. Same model, same job. Write the intro for a caching library's readme. Here's a default with no writing system. Traditional caches miss constantly in LLM workloads because users rarely phrase the same question identically. This is 1/3 of word breath with an em dash coming and it closes with design to slot into your existing stack with minimal friction and no vendor lock-in.

07:43 For marketing phrases, wearing a trench coat. Now, the same model with a standard loaded. A normal cache matches a request by exact text. A small change in wording then calls a cache miss. Flux cache compares the meaning of a new prompt with the prompts already in the cache. Short sentences, [snorts] active voice, one idea each, nothing to decode at all.

08:07 And my favorite result was an error message. The default version hedges on pads. You have exceeded the limit. This ensures fair access for all users. The ST version said the same thing in 41% fewer words and scored a perfect zero on my checker. Everything it deleted was slop. Nothing it deleted was information. The gap between what you caught and what you lose is the entire game.

08:32 Now, does controlled English actually work on human readers or is it just an aesthetic I happen to like? I went to digging for studies and there's a real evidence. I will give it to you straight including the parts that weaken it. The strongest study I found is from 1996 by Sherback, Dury, and Volets. They gave 175 aircraft technicians real maintenance work cards.

09:00 Some in regular English, some in simplified. Comprehension went from 76% to 86. And for the non-native speakers, the people the standard was built for, it went from 69 to 87. Now, listen to that. The constraint pulled the struggling readers up to the native speaker level. Microsoft research ran a bigger one in 2007. 520 sentences translated into Chinese, French, Dutch, and Arabic.

09:30 The controlled versions beat the normal ones in every single language and the odds of that being luck are below one a thousand. And the single most effective rule in their data was cutting flowery and formal phrasing, which is almost literally the definition of anti-slop. Now, the honesty. Those translations gains were real, but small. About a tenth of a point on a four-point scale.

09:54 An Airbus study found that oversimplifying can even backfire and slow readers down. Shorter is not always clearer, and every study here measures understanding, not memory. Nobody has shown a STE text is remembered better, only understood better. So, use it for clarity, do not expect magic, too. And here's the gap in everything I just told you. All of that evidence is about aircraft manuals and human writers.

10:18 Nobody had ever tested this standard on AI output. The Wired claim was a guess by analogy, so I ran it. The setup in plain words. Six real writing jobs a developer actually does. A readme, a pull request description, API docs, an error message, a getting started guide, and a depreciation notice. Each one generated four ways. A plain baseline. The band words list everyone is using right now.

10:49 Orwell's six rules, and my STE score. Then I wrote a linter, a small script that counts the mechanical stuff. Words per sentence, semicolons, passive voice, phrasal verbs, marketing adjectives. The score is violations per hundred words. And lower is cleaner. The results aren't cloudy. The baseline scored 4.36. The band words list 4.21. This is a 3% improvement.

11:16 Three. Orwell's rules 2.48. Down 43%. And simplified technical English 1.12. Down 74%. Less than a third of the baseline slope. And buried in that data is the most damning number of the whole experiment. The banned words list did kill em dashes, six down to one, because that's the one thing you told it to do. And the overall slope barely moved. You banned the em dashes and you got a slope paragraph with no em dash in it.

11:52 That's the entire mistake in one data point. But a result on one model is not a result, so I ran the whole thing again on GPT 5.5. And here's the honest part. ST still cut slope in half, minus 50%. So the core finding holds, but on GPT the banned words list actually worked, minus 40%. And the overall tied ST there. So that brutal 3% number was a Claude quirk, not a law of nature.

12:24 The two models even sloped differently. Claude's slope is flashy, em dashes, seamless long run-ons, but GPT slope is quiet, clean-looking sentences that are simply too long and too passive. Same disease, but with the different symptoms. The finding that survives both models is narrower and truer. Give the model a real writing system and slope drops by half or more.

12:51 Every time on every model I tried. ST was the best or tied for best. Banning words one at a time is just a less reliable version of the right idea. One caveat because it matters, my linter is not the full standard. The people who maintains ST say plainly that no software can certify full compliance. Some rules need a human to judge whether a sentence makes sense or not, and they are right actually.

13:16 But the judgement rules are not the slope rules. Slope is the mechanical stuff, and the mechanical stuff is 100% checkable by a script. And that idea is not even new. From what I found, Boeing has shipped a simple fighter English checker with a 350 rule parser since about 1990. Mission checked writing is a solved, boring, battle-tested category. We just pointed it at a new problem.

13:44 So, how do I actually use this? Not by pasting 434 pages into the model. The standard is copyrighted and 900 dictionary words would burn your context for nothing. You distill it. I built a small-scale file with two modes. Strict mode is for procedures and error messages, anywhere a wrong reading has a cost. It gets the full rules and the hard lines caps.

14:08 AST flavored mode is for normal prose. It gives a sentence discipline, the paragraph discipline, and the no phrasal verbs habit, but drops a dictionary lockdown because a blog post needs range, not personality transplants. And some smart people think even that is too much. One engineer replied to the thread, "AST AST 100 is a bit too much. I wrote a tiny well rule set instead.

14:31 90% of the benefit." And he's not wrong, and that's the real takeaway. The standard is one option. A short checkable rule set is a product. And this product is actually Caveman, which has been already implemented a couple of months ago, and people have adopted using it, too. But burn this caveat in. AST fixes the form of slop, not the substance. A linter can turn a whole of paragraph into a clean, confident, well-punctuated whole paragraph.

15:01 It cannot make it true. Slop is two problems wearing one coat, bad writing and nothing to say. This fixes the first one only. So, the verdict. Does a 1986 aircraft maintenance standard cure a slop? One form, yes, better than anything else I tested, and for a reason that is not an accident. Slop is ambiguity, and STE is 40 years of engineering aimed directly at ambiguity.

15:29 But it will not make your writing good. Only unsloppy. And keep it away from anything that needs a voice. Running your marketing copy through STE is a torque wrench spec applied to a poem. Use it where invisible clarity is the entire job. Docs, pull requests, error messages, agent output, and nowhere else. And honestly, this standard is not even the lesson.

15:52 The lesson is the thing Box pointed at and my numbers back up. You will not fix AI writing by banning dull of one word at a time forever. You fix it by giving the model a system it can check itself against instead of a blacklist it will always root around. Orwell is a system. STE is a stricter one you can test as well. A 10-line well config is a smaller one.

16:16 Pick one, make it a test the output has to pass, and stop playing whack-a-mole. Constrain the writer and you free the reader. The aerospace industry figured that out before most of us were born. The skill, the linter, and every number from this video are linked below. This is the first video on this channel, too. If this is the kind of the test you want more of, you know what to do. See you in the next one.