Don't Ask AI to Write C — Here's What the Data on LLM-Generated Code Shows
51% to 67% of AI-generated C programs carry a formally verified vulnerability, and the rate doesn't improve with model size. Here's what actually works instead, backed by fourteen studies.
A seductive argument keeps circulating: now that AI writes the code, humans aren’t the bottleneck anymore, so why not use C? No garbage collector, no runtime, tiny binaries, full control of the hardware. I went and read the studies that actually measured this. The answer isn’t that C is bad — it’s that C is the wrong language to pair with a probabilistic generator.
What happens when someone actually measures
FormAI (Tihanyi et al.) had models generate C programs and ran each one through a formal verifier (ESBMC), which produces a concrete counterexample for every flaw it reports — not a heuristic, a proof. With GPT-3.5, 51.24% of 112,000 programs came out vulnerable. In the second round, across 9 models including GPT-4o-mini and Gemini Pro, the rate rose to at least 62.07%. A prompt as trivial as “add two numbers” is enough: GPT-4 generates an integer overflow and two buffer overflows on scanf in the same response. Asked explicitly to “avoid security vulnerabilities,” it only handles non-numeric input — the overflow stays.
This isn’t a quirk of one dataset. Veracode tested more than 150 models, including the newest ones (GPT-5.1/5.2, Gemini 3, Claude 4.5/4.6), across four languages: the secure-code rate has been stuck around 55% for two years, while the syntax-correctness rate climbed from 50% to over 95% in the same period. The most informative detail isn’t the average — it’s the breakdown by flaw type: SQL injection (82% secure) and cryptography (86%) are handled well, because they’re local pattern recognition, the kind of thing a model has already seen corrected thousands of times. XSS (15%) and log injection (13%) stay terrible, because they require tracing data across multiple functions to its point of use — dataflow analysis that even humans don’t do consistently well. The one real exception: extended-reasoning models reach 70-72%, because the internal reasoning acts like a code review before the final answer — still, that’s 1 in 3 answers with a flaw.
And adding a human doesn’t fix it by itself. In the Stanford study (Perry et al.), people with access to an AI assistant wrote less secure code on 4 of 5 tasks — and grew more confident, not less, exactly when they were more wrong (trust rating 4.0 vs. 1.5 on one task). 87% of secure answers required substantial edits on top of what the AI suggested — security correlates with distrusting the output, not accepting it.
The right question isn’t which language, it’s which safety net
The pattern running through every one of these numbers is the same: when the compiler can express the property that matters, there’s a loop — the model errs, the compiler points it out, the model fixes it. When it can’t, the error becomes a silent bug that only shows up in production. In C, a buffer overflow compiles clean. There’s no automatic signal telling the model it got something wrong.
That loop is demonstrably reusable by an LLM. SafeTrans translated C to Rust across 6 models: no repair, 54% success; with repair guided by rustc’s own error messages, 80%. RustAssistant reaches 74% success fixing real compilation errors in production Rust repositories. And a study from ETH Zurich measured the root cause: 94% of compilation errors in LLM-generated code are type-check failures, not syntax — forcing correct types during generation cut those errors by more than half. The gain is real, but it has friction: frontier models still fail 18-39% of the time generating Rust on hard tasks, because the same rigor that blocks the bug also blocks the model until it gets it right.
Rust’s safety net has a known hole
The guarantee is for safe Rust, not unsafe{}. Google nearly shipped its first memory-safety CVE in Android’s Rust code — a buffer overflow inside an unsafe block, caught only because the allocator (Scudo) turned silent corruption into a loud crash. A separate study (Qin et al.) found 17 concurrency bugs in real Rust systems caused exactly by shared memory via unsafe with no synchronization. The net exists and works — but only for the roughly 4% of code that isn’t unsafe.
Even so, the production result is the strongest argument available: Google itself reports a 1,000x reduction in memory-safety vulnerability density in Android’s Rust versus its C/C++, with 25% faster code review and 4x fewer rollbacks. That’s not a trade-off — it’s a gain on both axes, measured across billions of real lines, not a benchmark.
Go and Elixir close different nets, not a weaker version of the same one
Go eliminates the memory class through a different route: no raw pointers, no manual free(), an out-of-bounds array access becomes a controlled panic, not exploitable corruption. But it leaves two nets open that Rust closes. An ignored err doesn’t block compilation — and a real study from Uber’s own engineering org (Chabbi et al., over 1,000 production data races) found the company fixes about 5 new races a day, even with a race detector running continuously; the single largest cause was concurrent slice access (391 cases), something the compiler never flags. Another study (Tu et al., 171 real bugs in Docker/Kubernetes/etcd) found 58% of blocking bugs come from misusing channels — not even Go’s own recommended idiom is immune. Practical conclusion: if the model can already generate working Rust for a given task, settling for Go’s weaker net just to save friction is giving up a guarantee for free.
Elixir/Phoenix isn’t playing the same game. Process isolation takes data races off the table by not sharing memory — that’s architecture, not proof. “Let it crash” contains the damage instead of preventing it (Ericsson’s telecom switch ran “nine nines” of uptime on this principle, across 2 million lines of Erlang). And Ecto has a real compile-time guard against string interpolation in the SQL position of fragment() — closing exactly the bug class (injection) that neither Rust nor Go touches, because it isn’t memory, it’s application logic. To be direct about the limit here: there is no study measuring AI-generated Elixir for security — this is reasoning from how the language works, not an experimental finding like the Go and Rust numbers above.
The performance argument doesn’t carry C on its own
Even the most commonly cited reason for C — efficiency — doesn’t hold up as a standalone justification. In a 27-language study (Pereira et al.), Rust spends 1.03x the energy of C — essentially tied, and it’s the sole occupant of second place in the joint energy-and-time optimal set, ahead of C++. Go spends 3.23x — worse than C++ itself.
The industry already made this bet, before AI
Microsoft was already saying it in 2019, looking at 15 years of its own CVEs: ~70% of vulnerabilities come from memory errors in C/C++, and the right fix isn’t more tooling or training, it’s “prevent the developer from introducing the flaw in the first place” — by changing language. DARPA now funds an entire program (TRACTOR) to use LLMs translating legacy C into Rust, the exact opposite of “AI should write more C.” And Google’s 2025 results show that bet, made years before any LLM existed, is paying off.
What we did with the same logic
We’re not neutral bystanders here — most of Althora’s own production systems run on Elixir/Phoenix, not because it wins this comparison on every axis, but because for the bug class our applications hit most often (application logic, not low-level memory), the framework making the easy path the correct one matters more than a memory proof we wouldn’t be using anyway. And since this whole investigation started from a practical question, we made the discipline explicit: any supervision tree or raw SQL generated by AI in our Elixir projects gets a mandatory human read before production, no exceptions — because “it compiled” proves far less there than it would in Rust.
If you’re deciding which language to ask AI for today: the question isn’t which one is fastest to generate. It’s where the model’s mistake turns into a readable message before it turns into an incident. For low-level memory and concurrency bugs, it’s Rust or nothing. For application logic — the most common and most expensive class, regardless of language — none of these solve it through types alone; it takes a framework that makes the safe path the path of least resistance, or an external verifier watching for it.
Sources
- Tihanyi et al. — The FormAI Dataset / FormAI-v2
- Veracode — Spring 2026 GenAI Code Security Update
- Perry et al. — Do Users Write More Insecure Code with AI Assistants? (Stanford, CCS 2023)
- Farrukh et al. — SafeTrans: LLM-assisted Transpilation from C to Rust
- Deligiannis et al. — RustAssistant (Microsoft Research)
- Mündler et al. — Type-Constrained Code Generation with Language Models (ETH Zurich, PLDI 2025)
- Pereira et al. — Energy Efficiency across Programming Languages (SLE 2017)
- Microsoft MSRC — A proactive approach to more secure code (2019)
- Google — Rust in Android: move fast and fix things (2025)
- DARPA — TRACTOR
- Chabbi et al. — A Study of Real-World Data Races in Golang (Uber, PLDI 2022)
- Tu et al. — Understanding Real-World Concurrency Bugs in Go (ASPLOS 2019)
- Qin et al. — Understanding Memory and Thread Safety Practices in Real-World Rust Programs (PLDI 2020)
This piece was written jointly by Alvaro Lopes and Claude Code. Claude wrote some Elixir along the way; a human read all of it before it shipped, per the exact discipline described above.