Get in touch

The Baker Doesn’t Eat the Bread

Published: June 5, 2026

Updated: May 30, 2026

A sourdough loaf with a scored ear, fresh from the oven

I bake my own sourdough. Not from a kit, not a shortcut. I make the starter, I maintain it, I mix and fold the dough, I shape it, I bake it. I’ve done it enough times that I don’t use a recipe anymore.

Last week I cut into a loaf that looked perfect.

Good rise. Crisp crust. Even that little ear on top you hope for but can’t quite plan. I thought, this is it. Finally got it right.

Then I tasted it.

Flat. Dense. Just off.

I had done everything right. Followed the process. Checked the signals. And still missed the outcome.

That’s what I keep thinking about when I look at what’s happening in software quality right now. A lot of engineering teams are cutting into what looks like a perfect loaf and calling it done. The metrics say green. The AI ran a thousand tests. The coverage dashboard looks great.

Then a user tries to do the thing the product is supposed to do.

And it doesn’t work.

The same problem appears in AI testing: passing checks does not always mean delivering software quality.

The float test problem

When I make sourdough, I drop a bit of starter into water before I bake. If it floats, it’s ready. Simple. Binary. Feels almost scientific.

But after doing this enough times, you learn something uncomfortable. The float test is just a signal. Sometimes it’s right. Sometimes it lies. You stop relying on the test alone and start reading the whole picture. The smell. The texture. How long it’s been fermenting. Whether it was a cold night.

We have the same float tests in software. Automated test suites. Coverage reports. Dashboards full of green. They give us confidence. They tell leadership the system is ready.

But they are signals. Not the truth.

This is where things get tricky for anyone managing engineering teams right now. AI can generate more tests, more coverage, more signals than any QA team could, faster and cheaper. That sounds like a solution. In practice it can be a very efficient way to produce false confidence. This is one of the challenges modern AI software testing still struggles to solve.

Same problem.

What AI is actually good at

This is not an argument against using AI in testing. That ship has already sailed.

AI is good at generating test scenarios quickly, exploring edge cases at scale, speeding up feedback loops, and finding patterns in large outputs. Real capabilities. They save time and catch things that would otherwise slip through.

The mistake is assuming that because AI can do more, humans need to do less.

That’s not how it works with sourdough either. A better thermometer doesn’t make you a better baker. It just gives you more accurate temperature readings. You still have to know what to do with that information.

We’ve written more about where AI fits in a testing practice here.

The part that doesn’t automate

Here’s what I’ve learned from running a software testing firm for over two decades. The work that matters most is not the execution. It’s the judgment layered on top of it.

Knowing what not to test. That’s a real skill. AI will generate tests for everything. Experienced testers decide what actually carries risk. Those are different activities.

Recognizing when a signal is misleading. The tests pass. Does that mean the system works? Sometimes yes. Sometimes the tests are testing the wrong things. It takes experience to know the difference.

Understanding how real users behave. AI can simulate personas. It can complete flows. What it can’t do is feel frustrated, lose trust, or abandon a process three steps in because something felt wrong. Real users do all of those things, usually without filing a bug report.

Making tradeoffs under pressure. Speed versus confidence. Coverage versus cost. Release now or hold. These aren’t technical decisions. They’re judgment calls that require context, business knowledge, and someone willing to own the outcome.

That last one matters more than people admit. AI has no skin in the game. Someone still has to say “this is good enough” or “this will fail users.” You cannot outsource that to a model.

The feedback loop nobody talks about

Here’s the sourdough thing that maps most cleanly to testing.

With bread, every loaf gets eaten. Every time. You bite into it and you know immediately whether it worked. The feedback loop is tight. You fail, you learn, you adjust, you bake again.

Software doesn’t work like that. The real testers are your users. But most of them never tell you when something is wrong. They adapt. They work around it. Or they leave. Quietly.

Silence looks like success. It usually isn’t.

So we build testing as a substitute for that missing feedback. QA cycles. Staging environments. Automated test suites. Now AI-generated scenarios. All of it exists to answer one question before users do: will someone struggle with this?

AI expands our ability to ask that question faster and at a greater scale. That’s valuable. But it still doesn’t close the loop. It doesn’t know what actually matters to the user sitting on the other end. It doesn’t get annoyed. It doesn’t give up. It just completes the flow because that’s what it was asked to do.

Where to now

If you’re leading an engineering team right now, you’re either already using AI in your testing and wondering how much to trust it, or you’re about to, and trying to figure out where the guardrails go.

The honest answer is that the balance requires judgment. Not just configuration.

The teams that get this right are not the ones using the most AI. They’re the ones who know where to insert human expertise into the process. When to trust the signal. When to override it. When to slow down even though the dashboard says green.

That’s not a tool problem. It’s a knowledge problem, and the confidence you have based on that knowledge.

I’ve watched teams shortcut this and ship fast. Some got lucky. Others found out the hard way what their float test was missing.

The goal isn’t to use less AI. The goal is to know where you still have to show up yourself.

If you’re not sure where those lines are in your own process, that’s probably worth a conversation. It’s what we spend most of our time thinking about at XBOSoft these days. Not replacing the tools, but making sure the judgment is there when it needs to be and ensuring our confidence level is based on accurate information. Not a pass or fail decision based on test cases that passed and didn’t.

Because in the end, the oven doesn’t decide if the bread is good.

The person eating it does.

The XBOSoft Perspective

At XBOSoft, we’ve seen the same pattern across organizations adopting AI software testing: execution scales faster than judgment. AI testing tools can generate more tests, expand coverage, and accelerate feedback loops, but software quality still depends on understanding risk, context, and user impact.

The challenge for engineering leaders is not deciding whether to adopt AI in testing — that decision is already being made. The challenge is deciding where human expertise must remain in the loop. Teams that successfully scale quality are not replacing judgment; they are applying it more intentionally.

Our perspective is simple: AI should amplify testing expertise, not replace it. The goal is not more signals. The goal is better decisions.

Next Steps

See where AI fits in testing This article explores why AI testing tools need human judgment. For a broader framework on using AI effectively in QA, read more about how AI fits into a modern testing practice.

Explore AI-Informed QA: Going Beyond the Hype

Talk through your situation Every organization’s testing process is different. A conversation can help clarify where AI can add value, where human expertise is still essential, and how to build confidence in your release decisions.

Contact XBOSoft

Related Articles and Resources

Looking for more insights on Agile, DevOps, and quality practices? Explore our latest articles for practical tips, proven strategies, and real-world lessons from QA teams around the world.