Claude Code made me question whether Chad needed to exist. Twenty minutes later, it showed me exactly why it should.
I was working on the ZeroLayer website and decided to test how well coding agents could handle its technical SEO and AI visibility. Chad is being built to identify what prevents a company from being discovered, decide what should be fixed, and help execute the work — so if a founder can just point Claude Code or Codex at a site and ask it to fix the SEO, it's fair to ask what's left for a product like Chad to do. So I did exactly that.
The first result was uncomfortably good
I gave a coding agent access to the ZeroLayer repository and asked it to run a technical audit. It found real issues — it inspected the architecture, metadata, and page structure, explained what needed to change, and connected the recommendations to the underlying code.
My first reaction was: this is really good. Not "good for an AI" — actually good. Faster than reviewing the repository by hand, with plausible reasoning, and capable of moving straight from diagnosis to a fix. For a moment I wondered whether this had commoditised an important part of Chad. Maybe the future really was that simple: point a powerful coding agent at the site, let it find the problems, let it fix them, move on.
Then I changed the model.
Sonnet gave me one interpretation of what needed fixing. Opus found a different set of issues and reordered the priorities. Codex 5.6 surfaced additional problems again. The foundations generally overlapped — there were basic technical facts every agent agreed on. But the further the analysis moved from observation into judgement, the more the answers diverged: different severities on the same findings, different implementations, different assumptions about how aggressively to change the site. Even within the same platform, changing the model or the reasoning level changed the recommendation.
None of the answers was obviously wrong. That was the problem.
Several intelligent versions of the truth
When a weak system gives you a bad answer, the decision is easy — you ignore it. When several capable systems give you different, well-reasoned answers, the decision gets harder. Each model could explain itself. Each sounded confident enough to act on. Each noticed something the others had ignored or weighted differently.
I had several intelligent opinions, and I was still the one responsible for turning one of them into an operating decision.
That's the strange consequence of more capable AI. We assume more intelligence produces more clarity — sometimes it does, but once intelligent analysis becomes abundant, it can create a new kind of uncertainty instead. One model says the positioning is unclear. Another says positioning is fine and the real problem is discoverability. A third thinks the site is being discovered but failing to convert the attention it gets. Every answer might contain something useful, and the founder has to reconcile them: which conclusion is evidence, which is interpretation, which issue actually matters most, which change is safe, and has something genuinely changed or did a different model just form a different opinion?
The bottleneck isn't access to an intelligent answer anymore. It's deciding which intelligent answer becomes the action.
Evidence is not the same as interpretation
The test exposed a distinction AI reports often blur. Some findings can be verified: a page is indexable or it isn't, a canonical tag is present or missing, a sitemap contains a URL or it doesn't, structured data passes validation or it fails. Other findings require judgement: is the positioning clear enough, is a technical problem material enough to fix now, would a new page create more value than improving an existing one, does the expected impact justify the effort.
A model can help with both kinds of questions. It shouldn't present both with the same certainty. There's a real difference between "we verified this is broken" and "based on the available context, we believe this is the most valuable thing to change" — and a confident model response can make those sound equally factual when they aren't.
This is where putting a more powerful model inside a product stops being enough on its own. The model can be excellent and the system still needs a consistent way to separate evidence from interpretation, compare competing recommendations, and decide what happens first.
Models provide intelligence. Chad should provide the governance — the same observe, decide, act discipline every ZeroLayer agent runs on.
What the experiment changed
This test changed how I think about Chad's value. The moat can't be "uses a smart model to audit a website" — anyone can do that now, and the models will keep improving. Chad needs to operate above the individual models: able to use whichever one is best for a given task without letting the company's priorities reset every time the underlying intelligence changes.
That means separating verified findings from model judgement. It means making agreement visible when evaluators agree, and disagreement visible when they don't. It means judging recommendations by expected impact, confidence, and effort — not by how convincingly they were written. And it means remembering why earlier decisions were made: a new coding-agent session can inspect a site as it exists today, but it doesn't know why a previous choice was accepted, what's already been tried, or which risk the team deliberately chose to tolerate. Without that context, every new model can confidently reopen the whole website. Continuity doesn't mean refusing to change — it means preserving the reasoning behind a settled decision and reopening it only when new evidence changes the answer, not whenever a different model happens to have a different opinion.
Founders don't need more recommendations
Technical founders already have near-unlimited analysis available — ChatGPT, Claude, Codex, an SEO crawler, an AI-visibility audit, a performance test, hundreds of recommendations before a single piece of work gets done. The scarce resource isn't advice anymore. It's clarity: what's definitely wrong, what's probably wrong, what actually matters, what happens first, and what can safely be ignored. Once a decision is made, the work still has to get done and the result still has to get measured — otherwise AI has just created another management responsibility.
That's why I'm sceptical of products that promise a founder an entire team of AI agents: one for SEO, one for social, one for content, one for analytics, one for strategy. The founder still has to decide which agent to activate, reconcile what they recommend, and work out whether any of it moved the business forward. That's not removing the marketing department. It's creating an AI marketing department the founder now has to manage.
I started this test wondering whether coding agents had made part of Chad unnecessary. I finished it believing the opposite. The agents were genuinely impressive, and Chad should use that capability rather than compete with it — but intelligence alone didn't give me continuity, didn't create a stable standard across models, and didn't stop the priorities shifting every time I changed the evaluator. Most of all, it didn't remove me as the person responsible for reconciling every opinion and deciding what happens next.
That's the job Chad needs to do. Not be the smartest model in the room — turn intelligence from whichever models are best into one operating decision, and make sure that decision gets executed.
Also published on Medium.