Can you ship AI-generated code to production?

The question gets asked as "is AI-generated code good enough for production?" — but that is not really the question, and answering it does not help anyone. Code is good enough or it isn't in the same way any code is: it depends what it does and who is relying on it.
The question underneath is harder and more useful. You did not write this. So how would you know?
When a developer on your team ships something, you have a lot of signal before you even read the diff. You know roughly how they think. You saw the PR description. There were tests, or there weren't, and you noticed. With generated code you have none of that. You have an app that works when you click it, and several thousand lines you have never read.
"It works" is the weakest possible evidence
Every AI builder demo ends with a working app. That is the easy part now, and it stopped being a differentiator some time ago.
Working means the happy path runs on your machine with your data while you are watching. It says nothing about the second concurrent user, the malformed input, the expired token, the query that is fine at fifty rows and pathological at fifty thousand, or the endpoint that checks authentication but forgets authorisation — so any logged-in user can read any other user's records by changing a number in the URL.
That last one is worth dwelling on, because it is the single most common serious flaw in generated applications. The code looks right. It has an auth check. The check answers "are you logged in?" when the question was "is this yours?"
You cannot see that by clicking through the app. You see it by looking for it.
What actually has to be true
Before an app carries real users and real data, a short list has to be true, and none of it is exotic:
- Authorisation is enforced server-side, on every endpoint, not hidden in the interface. A hidden button is not a permission.
- Secrets are not in the repository. Not in a config file, not in a committed
.env, not in a frontend bundle where anyone can read them. - Input is validated at the boundary, and queries are parameterised. Injection is a solved problem that keeps being unsolved.
- The schema has constraints. Foreign keys, uniqueness, not-null. A schema without them will accept data that breaks your application later, at a distance, in a way that is very hard to trace.
- Migrations exist and run forward. If the only record of your schema is the current state of a database, you have no way back and no way to a second environment.
- Something fails loudly. Errors that are swallowed in a
catchblock are worse than crashes, because the system keeps running while being wrong. - There are tests for the things you would be embarrassed to break. Not coverage targets — the specific behaviours a customer would notice.
Notice that this is the same list as for hand-written code. AI does not add requirements. It removes the incidental knowledge you would normally have about whether they were met.
Reading code you did not write
The instinct is to read all of it. Do not — on a few thousand lines you will skim, and skimming finds nothing.
Read in this order, and stop when you find something:
- The schema. It is the shortest file and the most load-bearing. Every mistake here becomes structural.
- One write endpoint, end to end. Route to validation to authorisation to query. If the pattern is sound here, it is usually sound everywhere, because generated code is consistent — which is one of its genuine advantages.
- Anything touching money, permissions or personal data. Three places where being wrong is expensive rather than annoying.
- The configuration. What is hard-coded, what is environment-specific, what is a default that should not have survived.
That is an afternoon, not a week, and it tells you more than reading everything would.
Where CodeSky stands on this
We took the position that a builder which cannot tell you what is wrong with its own output is only doing half the job — so CodeSky scores every project rather than declaring it finished.
The readiness report runs checks across security, quality, delivery, documentation and mobile, and grades the result: A at 90 and above, B at 75, C at 60, down to F. The grade is not the interesting part. A single failing critical check caps the project at C no matter how well everything else scores. That cap is deliberate. An app can be beautifully documented, fully tested, fast, and still have an endpoint that leaks other people's data — and an average would hide that behind a good-looking number.
Behind the score are agents that look for specific things rather than offering general opinions. The security audit reviews auth, injection, secrets and the OWASP top ten. Performance analysis looks for N+1 queries, hot paths and render thrashing. The accessibility audit checks focus order, contrast, semantics and keyboard navigation against WCAG. QA audits the app for broken and unwired flows — features that exist in the interface and do nothing. Unit tests are generated against the project's own testing framework, not a generic one.
The report is also shareable as a link, which matters more than it sounds. It is the artefact you hand to a client, an investor doing technical diligence, or the engineer you are about to hire — evidence rather than assurance.
And the code is in your GitHub repository, in a language and framework your team already knows, with migrations and real structure. That is the part that makes every check above possible. You cannot audit an export.
What this does not solve
Being straight about the limits, because a claim of total safety would be the least trustworthy thing in this article.
Static analysis finds categories of problem, not all problems. It will not catch a business rule that is subtly wrong — code that correctly implements the wrong requirement passes every check, which is exactly why the planning stage matters. It will not replace a load test against realistic data, or a penetration test before you handle payments, or the judgement of someone who knows your domain.
What it does is close the gap between "it works" and "I have looked", which is where most generated applications quietly ship from.
The bar
The honest bar is not "was this written by a human or a model". Plenty of hand-written code in production fails every item on that list.
The bar is whether anyone checked, whether the checks were specific enough to fail, and whether you can still tell in six months. AI-generated code can clear that bar. It just will not clear it by accident — and neither would yours.