AI slop: four checks that make AI-built software safe
· 5 min read
If you are a developer who thinks AI writes rubbish, you are mostly right. If you are a business owner who has heard a developer say so, they have a point. Left alone, an AI coding tool writes code that looks fine but copies what already exists, skips the unusual cases and adds features nobody asked for. That is slop. The sceptics are right about the problem, but the problem can be managed, and four habits do most of the work: a written plan, clear rules, tests and automatic checks.
The sceptics have the evidence
GitClear studied 211 million changed lines of code written between 2020 and 2024. From 2021 to 2024, copy-pasted lines rose from 8.3% to 12.3% of changes. Lines where existing code was tidied up for reuse, which developers call refactoring, fell from 25% to under 10%. In 2024, for the first time, more code was pasted than reorganised.
DORA, a long-running study of software teams, found the same pattern in its 2024 report. AI made individual developers more productive and happier in their jobs. But it was linked to less stable software and slower delivery of changes. DORA's own advice is the useful part: basics like small changes and strong testing still matter most.
AI writes more code, faster. Without discipline, that means more copying, bigger changes and more things to break. Most of that is a failure of process rather than of the AI, and process can be fixed.
Defence one: write it down before any code exists
Most slop starts with a vague request. "Build a booking page" leaves the AI room to guess, and it guesses generously. The fix is boring, and it works. Write a short brief first. We write ours in markdown, which is plain text with a few simple marks for headings and lists. People and AI can both read it easily. The brief says what the problem is, who it is for, what "done" looks like and what is left out.
The most useful part is the list of what is left out. "No user accounts in this release. No payments. No admin screen." An AI cannot invent what it has been told not to build. A reviewer can reject anything that is not in the brief.
Clear pass-or-fail conditions matter just as much as the description. "A booking cannot overlap another booking for the same room" can be checked. "Bookings should work well" cannot. Every requirement worth building can be written as a sentence a test can prove.
Defence two: give the AI rules it cannot talk its way around
A new developer learns a team's habits over weeks. An AI starts every session knowing nothing. A written rules file, kept with the code, closes that gap. It says how to run the tests, which ready-made components to use and which to avoid, where things go, and what must never happen.
Good rules are specific and checkable. "Write clean code" changes nothing. These change behaviour:
- Every change points to a requirement in the plan, or it does not get made.
- No new outside component unless the plan names it.
- If a requirement is unclear, stop and ask. Do not fill the gap with a guess.
- Passwords and keys never go in code, logs or test data.
The last two rules do more to stop slop than any clever wording of a request. An AI that is allowed to say "I am stuck" invents far less than one that thinks it must always finish.
Defence three: tests are the contract, and nobody marks their own work
Tests written after the code tend to describe what the code does, not what it should do. With AI this gets worse. The same AI that wrote a mistake will happily write a test that passes with the mistake in place.
Two habits fix most of this. First, write the tests from the pass-or-fail conditions before the code, so each test states what was asked for. Second, do not let the same AI session write both the code and the tests that judge it. A second AI session, or a person, writes or checks those tests.
Then test the ways it can fail, not just the smooth path. Wrong input, missing permission, another customer's records, a button pressed twice. The copying GitClear measured shows up here too. Tests that cover shared behaviour catch the copy that someone forgot to update.
Defence four: automatic checks nothing gets past
Developers call this continuous integration: a set of automatic checks that every change must pass before it is accepted. The checks do not care how confident the AI sounded. Every change faces the same tests: tidy layout, basic error checks, the full set of tests, a trial build, a scan of outside components for known security holes, a scan for leaked passwords, and a check for common security mistakes. If any check fails, the change is rejected.
Keep the changes small. DORA's point about small changes applies directly. An AI will write a thousand-line change as easily as a fifty-line one, and nobody checks a thousand lines properly. Each change should cover one requirement and be small enough for a person to review properly.
Releases get the same treatment. A private test copy comes before the live system. The new version must report that it is healthy before customers reach it. Going back to the previous version takes a single command. If something slips through, undoing it takes minutes.
What this does not fix
These defences remove most slop. They do not replace judgement. Someone still has to decide what is worth building, choose a design that will last, and notice when a requirement is wrong rather than simply unmet. That is where senior developers earn their keep, and it matters more now that the typing is cheap.
What to do with this
- Before the next AI-assisted change, write a one-page markdown brief with pass-or-fail conditions and a list of what is left out.
- Keep a rules file with your code. Include how to run the tests, which components are allowed, and "stop and ask when unclear".
- Write the tests first, and have someone other than the author check them.
- Set up automatic checks that block any change failing the tests, error checks, component scans or password scans.
- Keep each change small, and have a person review every one.
At Smartible, every project starts by writing the problem down properly in a Specification Sprint, so the code that follows has something firm to be measured against.