We stand with Ukraine
Go Wombat logo

Is Vibe Coding Bad for Software Quality?

Article by

Updated on June 25, 2026

Read — 4 minutes

In February 2025, Andrej Karpathy coined "vibe coding" to describe something many engineers were already doing: you tell the model what you want, accept what it gives you, and keep moving. The term spread fast. So did the questions that followed it, most of them circling the same concern: does this approach actually produce software that holds up in production?

This article answers it with data: quality metrics from an AI-assisted industrial B2B SaaS delivery our team completed. It is one project, so treat the numbers as a case study, not a benchmark.

What vibe coding actually is (and what it is not)

Vibe coding sits on a spectrum. At one end, a developer describes a feature in plain language, accepts whatever the model generates, and never reads the output closely. At the other end, an experienced engineer uses AI-assisted development as a deliberate accelerant: prompting carefully, reviewing outputs critically, running tests before committing, and catching the cases where the model built something plausible but incorrect.

Most professional AI-assisted development sits somewhere in the middle.

The term has become shorthand for AI-assisted development in general, which muddles conversations about quality. With measurement and review, AI assistance speeds delivery up. Without them, it builds up debt, mainly because the review step that catches poor code gets skipped.

So the useful question is whether a team has the measurement and governance discipline to manage the defect pattern that AI-assisted code tends to produce.

What the data from a real AI-assisted delivery actually shows

What the data from a real AI-assisted delivery actually shows

In a recent industrial B2B SaaS delivery our team completed, we tracked every issue logged on the project board from the first sprint through go-live. At the point of final QA, the board held 229 total issues: 101 bugs, 122 tasks, and 6 epics.

The key figures:

  1. Bug rate.

44.1% of all issues raised were bugs. Nearly one in every two items on the project tracker was a defect.

  1. Defect density

The project averaged 0.86 bugs per development task. The backend engineer carried 1.18 bugs per task; the frontend engineer carried 0.68.

  1. Severity

19 bugs were Critical or Highest severity. All 19 reached a "Ready For Production" status before go-live.

  1. Traceability

Only 32 of the 101 bugs had an explicit parent task link in the tracking system. The remaining 69 were raised without a reference to the feature that generated them.

  1. Data-integrity catch

A bug that silently stripped leading zeros from numeric fields in CSV exports was identified during active QA testing. The bug would not have been visible to users until they ran a batch import and found their product data corrupted.

By contrast, the DevOps engineer on the same project carried zero bugs across ten tasks. Infrastructure-as-code has fewer of the ambiguous integration surfaces where AI-generated application code tends to fail.

Where vibe-coded projects accumulate bugs (and why)

In the delivery above, five of the seven most bug-dense task areas were in authentication and user management: login flows, registration, password policy enforcement, admin user CRUD operations, and approval workflows. The two remaining high-defect areas were data import/export and catalogue redesign features added late in the build.

This clustering is not a coincidence. Language models have likely seen a great deal of authentication code, and much of it is probably tutorial-grade: it covers the happy path, handles the common errors and stops there. Real production auth involves token expiry timing, concurrent session handling, permission inheritance and state transitions that such code does not represent well. The model produces something plausible, and a basic review passes it.

The import/export bugs follow a slightly different logic. Features added near the end of a build are always the active QA frontier, regardless of whether AI is involved. The leading-zero strip described above is a classic format-handling edge case: technically wrong, not immediately visible, and destructive under specific production conditions.

In both areas, the AI-generated code handled core logic adequately and fell short on edge cases, especially in authentication state, data transformation and cross-component integration.

The four risks that turn vibe coding into a liability

The four risks that turn vibe coding into a liability

The complexity cliff

Bug intake in the delivery above was steady across the first four weeks, ranging from 4 to 12 bugs per week. Then 38 bugs arrived in a single week during the first full-system QA pass. Quality debt had built up unseen while the team moved fast, and it surfaced all at once when integrated testing began.

Concentration risk

One engineer was assigned 64.4% of all bugs on the project, including every Critical and Highest severity item. This follows from how AI-assisted work is organised: the developer who ran the generation sessions keeps the context needed to fix the output. And because nobody else reviewed the code line by line on the way in, nobody else can easily take over.

The traceability gap

When code is generated in large chunks rather than written task by task, bugs rarely trace back to a specific parent feature in the tracker. 68.3% of bugs in this project had no explicit task link. The workflow simply did not produce an audit trail.

The data-integrity tail

The last bugs to surface in an AI-assisted build are often the quietest and the most damaging: format errors, precision loss and edge-case data corruption. They throw no visible errors, so someone has to look for them deliberately.

What good vibe-coding governance looks like in practice

None of these risks is exclusive to AI-assisted development. Teams without strong governance run into all of them on traditional builds too. Vibe coding compresses the timeline: the happy path ships fast and the debt arrives sooner, so governance has to keep pace.

Five practices that keep an AI-assisted build on track:

  1. Weekly defect density tracking from sprint one

If you wait for QA to surface the quality picture, you will see it at the worst possible moment. A weekly bug intake figure costs nothing to track and shows the complexity cliff before it becomes a crisis.

  1. Explicit sign-off records for all Critical and Highest bugs

The meaning of "Ready For Production" depends on the project stage: for a pre-production build it signals readiness; for a live system it implies verified regression coverage. Each critical fix should carry a named reviewer, a regression test result and a timestamp.

  1. Distributed ownership reviews for AI-generated modules

Before a major module is merged, a second engineer reads the output critically, so that more than one person understands what was built and can support it later.

  1. Production environment validation as a hard gate

The production environment must be validated before bug resolution counts as complete. Bugs cleared in a staging environment that diverges from production are not truly cleared.

  1. Post-production monitoring is defined before go-live

Error logging thresholds, alerting rules, and a first-week triage cadence should be agreed upon before deployment, not assembled reactively in the days after.

None of this is exotic. It is what an AI-assisted delivery looks like when the team treats it as engineering rather than a shortcut to a working prototype.

What leaders should remember

On this project, the quality profile of AI-assisted code was predictable: bugs clustered in specific areas, arrived at specific moments and responded to specific governance practices. Predictable problems can be planned for.

The figure that tends to unsettle people is the bug rate: 44.1% of all tracked issues in our delivery. Track it weekly from sprint one, so that a spike like the 38-bug week shows up as soon as it starts. Teams that do not measure defect density during an AI-assisted build do not get cleaner software; they just find out later.

Focus QA effort where defects are most likely or most costly instead of spreading it evenly. On this project, that would have meant front-loading testing on authentication and data import/export, the areas with the densest bug clusters.

The four structural risks (the complexity cliff, concentration risk, the traceability gap and data-integrity edge cases) can all be reduced with a few habits applied consistently from sprint one.

Frequently asked questions

Is vibe coding bad for production software?

Not inherently; governance decides. In our delivery, AI-assisted development produced a recognisable quality profile: a high bug rate and edge-case failures clustered in complex feature areas. Teams that measure defect density, track bug intake weekly and apply structured QA can deliver production-ready software. Teams that use vibe coding instead of engineering discipline find the gaps later, often in production.

What bug rate should I expect from an AI-assisted build?

Our data comes from one delivery, so treat it as a reference point, not a benchmark. There, 44.1% of all tracked issues were bugs, and the project still reached production because most of those bugs were found and fixed before go-live. A more useful measure is defect density per development task (0.86 on average in our case) combined with the severity distribution. A high bug rate with a managed severity profile is a process outcome; an unknown bug rate is the real risk.

How do you run QA on a vibe-coded project?

The same way you run QA on any project, with sharper attention on two areas: high-complexity features (authentication and user management are a common example, though the hotspots vary by project), where AI-generated code often fails on edge cases, and data-handling features, where format and precision bugs surface late. Weekly defect density tracking, explicit sign-off records for critical bugs and a validated production environment are the minimum governance we recommend.

What are the biggest risks of vibe coding for a B2B product?

The four that matter most are the complexity cliff (a spike in bug intake when integrated QA begins), concentration risk (one engineer holding all the context and all the critical bugs), the traceability gap (bugs not linked to parent features in the tracker), and data-integrity edge cases that only become visible under specific production conditions. For a deeper look at concentration risk, read our article on the bus factor of vibe-coded projects.

Can vibe-coded software be delivered to a professional standard?

Yes, with governance. On the delivery described here, 44.1% of tracked issues were bugs, and all 19 Critical and Highest severity bugs reached "Ready For Production" status before go-live. This is one project without a non-AI baseline, so it shows what one governed AI-assisted delivery looked like, not how AI-assisted development compares with traditional development.

How can we help you ?

How can we help youHow can we help youHow can we help you