The assembly line behind AI scripts that don't sound AI
Ask ChatGPT for a viral TikTok script and you get slop. The assembly line we built instead, with real transcripts, overproduction, and three AI critics.
Ask AI to "write me a viral TikTok script" and you'll get the same thing every time. A chirpy hook, three tidy points, and a summary nobody asked for. It sounds like AI because it was written like AI, vibed from nothing in one pass, and reviewed by no one.
We write scripts that real creators can't distinguish from their own, and better prompting isn't the answer. You need what Henry Ford realized 100 years ago, an assembly line. It takes raw materials in one end (ideas and examples), throws out most of what gets produced in the middle, and has a few checks at the end that catch what everything else missed. Here's a walk down the line.
The raw material is transcripts, never topics
The root cause of AI-sounding scripts is trying to write just from a topic. A topic gives the model nothing to imitate, so it fills the gap with the average of everything, and the average of everything is slop.
So the assembly line starts with references. The top 5 real, currently-performing videos in the format are transcribed and actually read. The transcript gives the model the voice. Everything is downstream of this. Spend the most time finding the right samples for the right voice. Write them by hand if you have to. If you have bad references, don't even bother with the rest of the steps because they won't work.
Before anything writes scripts, those transcripts get turned into reference packs. The reference pack describes the hook style, the writing style (rhythm, tropes etc) and anything else about the style of writing it can determine.
Calibrate before you scale
Then it starts generating scripts. We read every single one and give feedback. This doesn't make sense or this hook was weak because it doesn't generate curiosity. These feedback sessions are essential because they become principles for the future agent writer to adhere to.
Do this for 100 scripts and you won't have to give it any more feedback. Taste compounds when you write it down.
Overproduce, then hide the losers
We tell it to write roughly 3x more scripts than we need. We also tell it to write 5 hooks per script. Everything gets ranked by the writer. Does it match the reference, does it have one memorable moment, would a human actually say it out loud and would it make sense to someone with no context on what this video is about? Only the top slice survives, and nobody ever sees the rejects.
The hook gets written before the rest of the script and goes through its own dedicated review steps because the first sentence carries more of the outcome than everything after it combined. Five options minimum, and the winner is the one that leaves the widest open loop, not the cleverest sentence.
The critics don't know the writer
The same model does not check its own scripts. A model asked to check its own work rationalizes things as good even if they suck, the same as people. So checking is done by three separate critics who had nothing to do with the writing, and each one owns a different failure mode.
One checks reference fidelity, and it carries the best test on the whole line. Could this draft have been written by someone who knew the format's name but never read the references? If yes, reject. One checks spoken quality. Would this specific person say this exact sentence to a friend without sounding rehearsed? This is critical for natural sounding UGC videos. The last checks the boring stuff, hooks, lengths, format rules. All three have to pass, failures go back with notes, and after three round trips a human decides.
The dumb gate
The final gate isn't a judgment call at all, and that's why it works. It's a mechanical scan, and it catches what the smart critics rationalize past. Em dashes get rejected, that's AI's fingerprint. Openers like "Okay so" get rejected. Any 8-word phrase appearing in two scripts gets rejected, because that's a template. Near-verbatim copies of a reference line get rejected too.
And the gates themselves get audited. A random 20% of everything the critics approved is re-verified, because checkers get lazy too. Trust nothing you haven't spot-checked, including your own quality gates.
What the assembly line is enforcing
Underneath the machinery sits a short list of rules about how people actually talk, and they're worth stealing even if you never build any of the rest. Mix your contractions, real speech uses "I'm" and "I am" in the same breath, and a script that's all full forms reads translated. Budget your fillers, one "honestly" sounds human and five sound generated. Never write "No X, no Y, just Z," the most AI sentence shape in existence. And name real things. "Caymus Cabernet" beats "wine brand" every time, because specificity is what people who actually belong sound like.
A human still reads every script that survives the line. But by the time one reaches a human, it's been through references, ranking, three critics, and a regex, and the human's job has changed from fixing slop to choosing between good options. That was the real discovery.
You can't prompt your way to quality. You have to build the assembly line that produces it.
