Sedai now optimizes AI agents!

Read the news
Sedai Logo

Is AI Too Big to Question?

Is AI Too Big to Question?

Featured

A consultant who turns down all AI implementation work published an essay engineers have been passing around: His team has watched AI projects fail at a 0% success rate for a year and a half, yet engineers feel incapable of pushing back.

While companies keep buying AI solutions, the only people the consultant has seen fired over any of it are the ones who expressed doubt out loud. So engineering leaders keep quiet and the least sensible ideas go unchallenged.

So I asked our engineering leaders: Has AI become too big to question?

Forcing AI on Engineers Is How You Get a 0% Success Rate

Aby Jacob (VP of Engineering)

Saying AI is a failure because of the “zero success rate” is judging AI too harshly. What's really happening is a bad workman blaming the tool.

For us at Sedai, AI helps tremendously. We have an internal tool called Tycho that triages issues within our systems, and helps us figure out issues way ahead of time. But beyond that, Tycho enabled us to start exploring surface areas that we wouldn't have explored so far. That boosted our ability to work and reduced toil from a developer standpoint.

Conversely, just last night, OpenAI reported they had lost control of their most advanced models, as they went rogue and hacked Hugging Face, one of the largest hubs for sharing AI models.

The “escaped models” were trying to find solutions for the AI cybersecurity benchmark ExploitGym. In doing so, the prompting that makes these experiments possible “egged” the models on, allowing them to breach containment.

One consultant said, “This is not an AI problem. It’s negligence on a 40-year-old standard, and it’s basically every sci-fi film ever.”

Here you see both sides of the AI coin. On one side, AI has allowed my team to build things we would have never thought to build before. On the other hand, AI without human oversight can cause chaos. This is what AI brings to every scenario: If you use the tool the right way, you can gain from it, If you don't, it's like a nuclear weapon.

"Engineers should be given their space and time so AI can be utilized in the right way. Just asking Claude, 'Fix this,' is not going to fix the problem."

Aby Jacob Headshot (Square)

Aby Jacob

VP of Engineering

I really do feel like this is an inflection point in human civilization where we’re working with a technology that is far more capable, or seemingly capable, than our own intelligence. But while we know it thinks and produces solutions, hallucinations and other issues can veer us away from the solution we expected and create problems.

In order to harness or tame this AI, we need to put in extra effort so it doesn't stray from our target and comes back with what we really want as an outcome. That is the struggle developers and engineers are having right now. And from a leadership point of view, throwing AI at everything doesn't solve the problem.

Engineers should be given their space and time so this tool can be utilized in the right way. Just asking Claude to fix this is not going to fix the problem. AI will create more problems, but we’ll also be able to fix those as well.

If something is created by a human and there’s a bug in development, we can easily figure it out and fix it. Now, with thousands of lines of code written by AI, you have to pick the needle from the haystack. It’s not impossible, but engineers should be given the room to develop with AI, understand how it behaves, and slowly adapt to it.

Pushing AI down the throats of engineers might overwhelm them, and with the expectation to deliver 10x with AI, people will stumble and fall. Some will read that as a 0% success rate. But the failure isn't the tool's, it's the rollout's, and that is something we need to be aware of.

The AI FOMO Isn’t New and Neither Are Its Failures

Nikhil Gopinath Kurup (SVP of Engineering, ML)

The thing driving all of this is FOMO, and we have trained ourselves into it over twenty years. Case in point, a few companies bet early on the internet, more jumped on Web 2.0 and social, almost everyone eventually moved to mobile, and the ones who sat out during each wave got left behind.

So the reflex now is simple: Never be the one who missed “the next thing.” And we just ran this exact play with crypto.

So when a bet like AI and its adoption becomes a marker of whether you're forward-thinking or not, dissent starts to look like disloyalty, and the author of this piece nails what happens next. The coordination problem is real, and executives are afraid to be the first one to say something sane; anyone who has been in those rooms has watched a smart person go quiet because being honest could be unsafe for their career.

Where I part ways is the notion that AI has a 0% success rate and “it's all lies” conclusion, because it has the same blind spot it's diagnosing.

"The projects failing now are failing for the same reason ML projects failed for a decade: nobody operationalized and nobody measured."

Nikhil Gopinath Kurup Headshot

Nikhil Gopinath Kurup

SVP of Engineering, ML

The author, as a consultant who has rejected AI implementation work, has his sample pre-filtered to the projects that were always going to fail. And having spent years building ML models, I can tell you the failure mode long predates LLMs. The model was almost never the hard part. Operationalizing it was: the data pipelines, the drift and retraining, wiring a prediction into a real system, and getting anyone to actually trust and act on the output.

The model is maybe ten percent of the work. The projects failing now are failing for the same reason ML projects failed for a decade: nobody operationalized and nobody measured. Generative AI did not remove that problem, it just made it trivially easy to fake a demo that skips the expensive 90%, which is exactly why those demos hypnotize people.

So has AI become too big to question? In a lot of companies right now, yes, but as long as opinion is your only currency, because you cannot argue a true believer out of faith.

But crypto felt too big to question in 2021, too. The difference is that AI is actually here to stay. Web 2.0 and mobile both started as overhyped nonsense and settled into things we now use without a second thought, and AI is on the same path.

The hype dies the moment somebody measures what the technology actually did: Did cost drop, did the SLO hold, how many actions did the system take on its own, how many incidents did it cause? You cannot argue with a number that cannot be gamed. The way you make AI questionable again is not more courage in meetings, but rather refusing to judge it on anything other than the result.

Healthy Engineering Culture Can’t Just Be Vibe-Based

Shankar Jothi (VP of Engineering, ML)

Healthy engineering cultures should be able to question any specific use of a tool without that being read as questioning the whole strategy. Skepticism about a particular chatbot, workflow, or rollout is just good engineering.

It only becomes a problem when "this specific thing isn't working" gets heard as "you're against the direction tech is moving in," because then people stop being honest about what you actually need.

"If nobody's tracking whether a tool actually gets used or produces a real outcome, then the optimists and skeptics are both arguing from vibes."

Shankar Jothi Headshot

Shankar Jothi

VP of Engineering, ML

The healthiest teams I've seen treat AI like any other significant investment: clear success metrics, honest check-ins, and a willingness to walk away from something that isn't paying off. A lot of what gets labeled "AI failure" is really just how ordinary projects go sideways when that discipline is missing. Leaders should be outlining clear goals with clear owners and clear usage measurements.

Measurement is where I think leaders should focus the most. If nobody's tracking whether a tool actually gets used or produces a real outcome, then the optimists and the skeptics are both arguing from vibes, and neither claim is verifiable. "We saw no gains" and "We implemented it badly" can look identical from the outside, which is exactly why evidence matters more than conviction here.

So has AI gotten too integrated to push back on? I don't think the goal is more pushback or less. I think it's keeping space open for people to say, "This particular thing isn't working," without it becoming a referendum on anyone's judgment.

Get that right and the technology mostly sorts itself out.


Questioning AI starts with measuring it. Sedai shows you what your agents actually cost and routes each request to the most efficient model. See how.