Can't AI Test It?

If AI can build websites and apps, why can't we just get AI to test everything?
Artificial intelligence head image with glowing brain connections
Written by
Tom Batey
Published on
September 22, 2026

Many of our clients are using AI to assist them when building websites and web applications. AI can obviously generate code very quickly, new features and whole applications take shape in the blink of an eye. The speed of this progress can all feel a bit like magic.

That is, however, until it's time to test what the AI has produced. That magic feeling soon fades.

Another idea doing the rounds is having AI generate test cases and then execute those test cases to test a website or application. I”ve seen many boasts on LinkedIn and other places about not needing any human involvement in the testing of a project.

If we want to go faster, we can just get AI to test it, can't we? Read on to discover why we shouldn't.

AI Horror Stories

We've heard the horror stories that often stem from an over-reliance on AI and giving it permissions or an ability to run in a production environment that it probably should not have had.

The AI then goes on to delete code, drop database tables, or do something that it should definitely not do but has a major negative impact on the application and the business.

I'm not really focusing on those aspects here, this article is about whether using AI to test AI-generated code or human built stuff is a good idea or not.

Faster Faster Faster

Generating code with AI or using an AI to assist with writing code enables software teams to build faster. New features are churned out faster, the roadmap moves along at an ever quicker pace, new releases are more frequent and the momentum builds, 'look how fast we are moving'.

Testing, traditionally seen as holding back teams that want to go faster, can now just be speeded up by using AI. No more annoying humans getting in the way of releasing dozens of super new features as quickly as possible. Just get the AI to test the stuff that the AI generated - perfect, right?

Take a step back for a minute, is faster always better? Is churning out piles of half-baked new features really benefitting your product or is it causing your users to grow increasingly tired and frustrated of the frequent problems and bad execution that contributes to a worsening experience?

How many times do you release a new feature, only to find after the release that there were problems with it - something didn't work, users started complaining or people didn't use it, didn't understand it, it wasn't what they wanted or expected? Releases are reworked, iterated, hot fixed, over and over - is this achieving anything by going faster? Are you going faster but more in a circular direction?

Is it possible for the business to go faster by actually slowing the pace of releases down, allowing for more careful and thoughtful work to take place and doing some thorough testing?

Usefulness Of AI

I”m not suggesting for one minute that AI is not a very useful tool. We use AI in an increasing amount at WebDepend, as a tool to help us in the testing of websites and applications.

And that”s what AI is, it”s a tool that people use to help them in their work. AI does not replace testers or testing., just as it does not replace developers, engineers, product managers or designers.

We are also building testing tools with AI. Partly to learn for ourselves what AI is good at and what it”s not good at. And also to build things that are genuinely useful for our team, after they have been tested thoroughly to find any issues.

What Is Testing Anyway?

A quick reminder as to what testing actually is. Hint - testing is not an automated thing carrying out some arbitrary checks.

Software Testing, as defined by Wikipedia, 'is the act of checking whether software meets its intended objectives and satisfies expectations.'

You can give the AI tool the intended objectives of the thing you want it to check, and the AI will most likely also pick up some signals from other material it has been given or has access to, which enable it to carry out tests that seem entirely reasonable and may find some legitimate bugs or issues that require investigation.

However, in the above Wikipedia definition, the part the AI tool cannot do is whether the software satisfies expectations. The AI tool has no way of assessing whether the project being tested satisfies expectations or not, as it has no method of understanding and so cannot gain the understanding of whether a specified expectation or an implied expectation has been satisfied.

An AI tool following a series of prompts in the form of test objectives, or a set of test cases. This is providing some basic value but is not an assessment of the application nor is it testing. Furthermore, AI tools may raise 'issues' that are not issues at all because the AI tool does not understand the implied expectations of the project being tested.

In using AI to build and carry out checks for WebDepend Labs projects, we're gaining experience of how an AI tool interprets some of the prompts we give it and what issues it comes back with. Some of them are indeed relevant whereas other issues are missed entirely.

Similar to test automation, having AI carry out a series of checks can be a part of testing but should not be the whole test effort.

Responsibility

Who is actually responsible for what gets built and released to production?

If an AI generates the code, then the AI goes on to check that code works as intended or whether the generated code meets the objectives and then the AI pushes that code to production, is the AI responsible for what ultimately ends up in production?

The answer is, of course not, the AI cannot be responsible as it is only a tool. AI is as responsible as your Jira or your GitHub. The humans using the AI are responsible.

When AI is used in the development process and also used to check what has been built, there is a responsibility to check that what has been built by the AI is what the business wants.

There is also responsibility towards the users and the customers of the business to check that what has been built is something they can adequately use, that satisfies their expectations (remember the definition of software testing?) and does not cause problems for them. Otherwise, users are liable to move to a different application, service or website that does satisfy their expectations.

There is also responsibility to check that what has been built complies with any rules and regulations that the business may operate under.

Because AI cannot be responsible for anything, does the responsibility therefore reside with the development team?

The answer is that the development team is also not responsible, as they are instructed to work using the tools and processes provided and defined by the management. The management of the business are ultimately responsible for what gets deployed to production, whether AI generated all the code and did all the testing or whether people were involved.

But Everything Seems To Work Fine

If everything seems to work fine with the AI doing the testing then what is the problem?

Well, the whole point of testing is to cut out the 'seems to' part. Anything can 'seem to' work correctly in the first impression. Testing is about digging around underneath the glossy exterior to look for the problems that lie there - and problems do indeed lie there, lots of them.

What AI is good at is announcing confidently that all checks passed, or all test cases are green. AI does not understand what it is testing or why, it has no concept of expectations that the website visitors or application users have when they try to use the application. It does have an ability to mimic what understanding looks like and that can seem like the AI is testing thoroughly and doing a good job at it.

There is a duty for someone to go beyond what the results are from the AI's checks to carry out testing and look for issues.

Conclusion

So should you use AI to conduct all the testing for a project or not?

For simple things that don't matter that much, such as personal projects then feel free to use AI as a checker. I would still recommend that you also do some testing though to make sure you have carried out some kind of assessment. Overall, there is less risk for personal projects that if the AI missed something or got something wrong, then the only reputation that is damaged is your own.

But for a business, which has a responsibility to its users and customers and to rules and regulations, AI can be used as a tool to do some checks but should not be used to test everything. There needs to be someone involved in carrying out testing that goes beyond what the AI is checking and also challenges the AI's findings.

At WebDepend, we're starting to use AI to carry out some base-level checks, things that are well defined and fairly easy to report back on. One example is an accessibility scanner tool that we've developed, which scans a website and reports back on accessibility issues.

Even for this, where the accessibility tool is expected to find less than half of the accessibility issues, there should be a person checking the report over, investigating the issues and then carrying out the rest of the accessibility testing that the automated tool cannot cover.

Photo by Viktor Hanacek from picjumbo.com

Newsletter
No spam. Just the latest testing information and tips, interesting articles and Testing Manager updates in your inbox.
Read about our privacy policy.
Thank you for signing up! Your first newsletter will be sent soon.
Oops! Something went wrong while submitting the form.

Testing On Demand

Flexible testing when you need it, no minimum amounts and no contracts.

Related Articles

Project managers are busy people and can get burnt out from having to test projects
Web Project Management
November 3, 2025

Agencies - are your project managers getting burnt out from testing?

Project managers can get burnt out from additional responsibilities, such as testing.
Read post
Estimating web development projects is hard - don't forget to estimate testing as well
Web Project Management
September 1, 2025

Tackling a problem for Digital Agencies - effectively estimating testing

Digital agencies struggle to estimate testing, but now there is a better way.
Read post
The ability to estimate testing - how much time is required for testing and what should be included in the scope of testing - has always been a problem
Web Project Management
August 19, 2025

New Testing Estimator feature in version 1.10

A new feature in Testing Manager to make it easier to effectively estimate testing.
Read post