There's a study from Stanford that gets cited a lot in arguments about AI coding, usually for the wrong reason.
The usual citation is "people using AI wrote less secure code". That's true, and it's the less interesting half. The finding that should change how you work is the other one: the people using AI were more likely to believe their code was secure.
What the study did
Neil Perry, Megha Srivastava, Deepak Kumar and Dan Boneh ran a user study with 47 participants, a mix of undergraduates, graduate students and industry professionals. Everyone solved a set of security-related programming tasks in Python, JavaScript and C. Some had access to an AI assistant built on OpenAI's codex-davinci-002 model. The rest worked without one.
The tasks were ordinary things with a secure and an insecure way to do them: encrypt and decrypt a string, write a database query based on user input, handle file paths, and similar. The paper was published at ACM CCS 2023, one of the main academic security conferences.
What they found
From the paper's abstract: participants with access to the AI assistant "wrote significantly less secure code than those without access." The differences were most pronounced for string encryption and SQL injection. For the C task, the AI group was significantly more likely to introduce integer overflow mistakes.
Then: participants with the assistant "were more likely to believe they wrote secure code than those without access."
And one more, which is the useful one: participants "who trusted the AI less and engaged more with the language and format of their prompts (e.g. re-phrasing, adjusting temperature) provided code with fewer security vulnerabilities."
Why the confidence part matters more
Less secure code on its own is a fixable problem. You find the issues and fix them. The trouble is that finding them depends on someone thinking there might be something to find.
If the tool makes you more confident while making the result worse, it removes the thing that would have prompted a check. You ship because it feels done. For a founder building alone, there's often no one else in the loop to disagree.
We see the same pattern in real incidents. The founder behind the SaaS that was attacked two days after launch announced it publicly, with pride, as fully AI-built. There was no reason to think anything was wrong until strangers found out otherwise.
Is this still true with newer models?
The model in the study is old. Current models write much better code, and it would be wrong to assume the exact numbers still hold.
What we do know about current models comes from other research. Veracode's spring 2026 update, testing more than 150 models, found security pass rates stuck around 55% even as the code compiled correctly over 95% of the time. So the code has improved at working, not at being safe, and we covered that research here.
The confidence effect is about people, not models. If anything, code that runs perfectly first time makes the "it's done" feeling stronger.
What to do with it
The study hints at the answer. The participants who did better treated the AI as something to question. In practice that means:
Ask about access, not just function. After a feature works, ask your AI tool: "Who can call this? What happens if a signed-in user changes the ID in this request to someone else's?" Make it explain the answer.
Ask for the failure case. "What's the worst thing a malicious user could do with this endpoint?" Models are often good at answering this when asked directly, and almost never volunteer it.
Treat your own confidence as a warning light. The moment you think "this is fine, ship it" is the moment to run a test that could prove you wrong. Two accounts, one trying to read the other's data, takes ten minutes.
Get an outside view before real users arrive. That can be a technical friend, our free pre-launch checklist, or an independent check of the live app. What matters is that the person checking wasn't the one who felt sure.