Categories
business technology

Consistency is a challenge in AI

A big source of concern around AI in business applications revolves around trust. Can you trust AI to give you the right answer; or will you be constantly worrying about hallucinations? This will always be a challenge because of another characteristic that I think has as big an impact on business applications: a lack of consistency. Each session is an independent event and each agent is an independent agent. The same context can lead to different results from AI even if in many cases they are small and negligible. However, given that the promise of AI means embedding them into critical systems for business and public sector, even small and negligible drift will cause headaches, and that’s assuming the mistakes stay small and negligible. 

Let’s take the example of email agents, a very common use case for personal productivity. A user writes a set of guidelines for AI agents, gives them access to their inbox, and lets the agent autonomously do their work. Many people build human-in-the-loop into the process so they can review but many more likely prefer to just take everything off of their hands.

For this use case, you would write conditional logic in traditional programming. If an email has such-and-such topic and comes from so-and-so, then provide this message in response. There is no judgment. It is predictable and it is consistent. But it is also limited. Even if you build a more sophisticated engine to characterize emails, you still have a discrete set of cases with a discrete set of outputs.

Using AI for the automation of email responses is a completely different paradigm. There are no algorithms and logic statements are for direction not for machine processing. With AI, you provide a prompt with guidance on how to categorize emails, the types of responses to send, when to send the responses, etc. Your system prompt guidance is likely going to be pretty detailed but very different in character than a discrete set of logic statements. The AI model then uses judgment to decide what to do. Testing the process is also different. You can test as much as you’d like but repeatability is no longer the goal; instead, your measures of success shift from a pass/fail certainty to confidence levels.

Now let’s consider business applications and that most ubiquitous use of AI: the virtual assistant, aka the chatbot. At its most basic, a chatbot passes your message to the LLM, along with any relevant additional context (e.g. policy documents for a customer service bot or an API library for a transactional HR bot). In this case, each message is actually its own independent session so it is impossible to predict how a model will respond. For basic use cases, the risk is fairly low as the instructions are straightforward, the solution is simple, and the response options are limited. It doesn’t matter if the model formulates the phrasing of the response a bit differently each time. For more complex interactions and, beyond chatbots, applications with more sophistication and ambiguity, the lack of consistency can transition from innocuous to critical.

AI is embedded in everything from software that checks application hardening against Federal security controls to weapons guidance systems to resume and candidate screening. These are use cases where the model is making more and more judgment calls. Consistency matters but consistency is not guaranteed. Organizations and implementers must take this into account. They must view AI as inherently unpredictable and build harnesses, safety nets, and governance models around it to protect their business and mission goals. Where AI was once thought to be a death knell for SaaS applications the most foresighted have integrated AI into their products cleverly by making it part of the ecosystem instead of making the ecosystem about AI. Tools like watsonx orchestrate should make a lot of sense to business and IT leaders: these tools not only help with managing integrity, safety, and cost but also consistency within common processes and workflows. 

The consistency challenge becomes even more interesting across models. Not too long ago I had to undertake an activity with one of my AI applications (Shelf, a book recommendation engine, it’s great so check it out) that I never expected. When excitedly describing it to a friend and giving her the link, we discovered the site was down. It turns out the issue was that Anthropic had deprecated the model I was using. I now had to deal with one of the greatest banes of an IT professional: the upgrade.

When you upgrade a traditional application, whether truly old school or SaaS, the core functions more or less stay the same. When you change your AI model, it’s possible everything changes. For me, a newer Claude Sonnet model took 4x as long, interpreted my system prompt differently, and returned more hallucinations. During a debugging session, I made a comment about how inefficient the new model was. Here is Claude’s response:

It’s not that it’s “less efficient” — it’s that extended thinking is genuinely more capable on newer models, which tends to mean more thorough reasoning. The old model was likely under-thinking the scoring algorithm. Sonnet 4.6 is actually doing the work more carefully…

One can’t help but think that Anthropic has a lot of consulting materials in its training set as this thought completely ignores that everything was working exactly the way I wanted in the previous model. In a business context, this would be even more frustrating as the implication is that a more advanced model will always be better but anyone familiar with technology knows this is often untrue. Left with no choice, I ended up moving up to the more expensive Opus model, dealing with new parameters and more “independent” behavior, where oftentimes it would ignore the directions in the system prompt. I also ran into this doozy while debugging:

But Opus will sometimes guess rather than say it doesn’t know

That’s the type of unpredictable and difficult-to-control behavior that will give any programmer, IT architect, project manager, or business leader fits. In my case, the stakes are low: it’s an application I built as a hobby. Put in the context of a business application and the upgrade costs become significant. Not only is the new model 2.5x more expensive than the prior, you have to build more code to manage the unpredictability and different approach to judgment of the more advanced model. The upgrade process is different but perhaps it is comforting to know that even with a technology that few of us understand, you can still experience the same headaches as you did with legacy IT.

The challenge then is handling the lack of consistency and that’s where the harnesses I mentioned earlier come into play. Prompt engineering is critical but experience has shown me that it is not as reliable as we hope: models will oftentimes ignore or miss guidance in the prompt. Reducing the calls to an LLM also helps. There is something to be said for the predictability of logic statements, which is why all commercial chatbot builders have integrated deterministic AI to understand certain intents so you can provide discrete responses. Error handling is also critical. The difference is instead of a catch block like in traditional code, you are building logic that anticipates potential inconsistencies, such as ignoring output formats dictated in your system prompt. 

When using AI for personal use, drifts in consistency are not as important; in fact, they may even be helpful for things like building itineraries for a holiday as you get different ideas in each session. For a business application, consistency is critical and any organization deploying an AI-integrated tool to a customer base risks disaster if they do not account for consistency along with the more traditional concerns of cost, trust, and safety.

– – –

Note: This post is also available on my AI-focused site at https://dorukai.dorukakan.com/.

Categories
business culture

AI’s advanced models are bad business

The general atmosphere around AI has become much more alarmist over the past couple of months. With news of rogue AI agents escaping into the wild to wreak havoc or Anthropic claiming they have blocked several attempts to develop biological weapons, we are now squarely in the “This is just too big” era of AI. This is the same space occupied by climate change, a problem so big and complicated, with so many vested interests preferring the status quo, that it seems impossible or at least improbable that there will be any action to satisfy the urgent need for regulation or protection or oversight or, really, anything that minimizes the risk to humanity and the world at large.

Opinions on the ethical and financial value of AI are wide-ranging, touching on everything from NIMBYism and global markets to ethics and the human condition. Worries about the governance and environmental impacts of data centers are the current political hot button in our country. The dialog around job destruction has subsided for the moment as some evidence points to AI creating new jobs in the short term – though I fear the long term will be more dire for employment. The moral and philosophical debates on its impact on art and humanity continue and they are amongst the most important of all discussions but perhaps too esoteric for regular dinner table conversations. And the bullish excitement about rocketing stock market values combined with an anxious, hope-for-the-best and don’t-bother-planning-for-the-worst hand-wringing about a potential bubble is an emotional roller coaster for anyone with a retirement fund. 

It’s an all-encompassing topic and the future keeps arriving faster than we can handle. With AI frontier labs churning out more and more advanced, and more and more energy-hungry, models with more and more concerning consequences, I can’t help but wonder what is their purpose? Anthropic states on their home page that “[a]t Anthropic, we build AI to serve humanity’s long-term well-being.” Their purpose: “We believe AI will have a vast impact on the world. Anthropic is dedicated to building systems that people can rely on and generating research about the opportunities and risks of AI.” Meanwhile, Open AI “is an AI research and deployment company. Our mission is to ensure that artificial general intelligence benefits all of humanity.” Despite both prioritizing the idea of research, they are also companies, with revenue and operating costs and margins and so forth. They are both planning IPOs within the next six months, which will bring even greater scrutiny into their financial performance and lead to a reckoning of their balance between markets and mission.

I see the frontier labs continued drive to build more advanced models as bad business practice. Adoption of AI across public sector and commercial organizations remains relatively slow, with military and intelligence organizations being the most avid. Moreover, the current state of models is more than good enough to meet almost all of the use cases that companies would benefit from. Putting it in terms of Anthropic’s Claude models, which are the ones I most use, Sonnet is sufficient for business use with Opus, let alone Fable, adding little additional value. So the frontier labs are essentially spending huge amounts of capital on models that have limited practical value for their enterprise and business customers, organizations whose continued adoption of (and payment for) AI is financially existential to the continued existence of the labs.

How will the markets respond? Can the frontier labs, once public, balance the pressure for increasing revenue with their ambitious appetites, and outsized costs, for cutting-edge research? Are the advanced models flooding the market with too much supply for something that has little demand? Perhaps I lack the necessary imagination but I consider myself fairly fluent in AI usage, certainly more so than most of my peers in the corporate world, and I find Opus to be rather disappointing. If and when I am in a position to procure technology for an organization, I know that Sonnet is fully sufficient for almost every use case and Opus would be a buy for perhaps a small number of individuals to experiment with – certainly not needed for the enterprise, especially at 2.5x the cost. Businesses can derive immense value from Sonnet and they should be using it (or other similar models from other labs). Opus? Pass.

Of course, many companies have been able to successfully combine business and research in all sorts of industries, like pharmaceuticals and manufacturing. I have worked most of my career at IBM and, across its various incarnations, it has a legacy that continues to this day of running an IBM Research division, which is currently exploring the cutting edge of quantum. The difference is in the stakes. The amount of money that is tied up in the frontier labs means they must be financially successful and because of their operating costs the target for financial success is a huge number. They have already become too big to fail and that is before becoming public. The consequences of failure for other companies is implosion; for the AI labs, it is explosion, a domino effect with an outsized impact on the world economy.

There have been so many eloquent and well-reasoned arguments to slow down the pace of AI development that focus on the societal, geopolitical, and environmental consequences. I tend to agree with most of them. But, counterintuitively, I also think we should slow down the pace of AI development for economic reasons as well. The pace is too fast even in our short-term, quarter-to-quarter world, making it increasingly difficult for the business world to effectively adopt and deploy the technology. We are so inured to the capitalistic drive for growth, growth, growth but this is a classic example of the need to walk before running. If you don’t allow customers to walk first they may just throw their hands up and say no mas – and that’s not a good thing for the margins these companies must maintain.

One of my favorite ideas I’ve read is to split these companies up into a profit-making enterprise and a research arm, with the latter being absorbed by the academic community. Of course, with our universities under unparalleled attack by myopic political forces, this idea is likely just a pipe dream. I still find it most compelling. The profit-making side could focus on tweaking models and building tools for their customers, and bringing more advanced models to market more thoughtfully, in the same way that technology companies release upgrades and new versions. And the research side would reside within a world that is more accountable and more collaborative, distributing control of the future of AI whereas in the current state it is concentrated in the hands of a very few and increasingly unstable men. 

Some day, activist investors will likely force a spinoff of the research arms of the labs anyway, so why not get ahead of it with a win-win proposition of benefits for business through more focus and a more controlled and transparent development for AI, where it can actually benefit humanity. The alternative, absent any serious regulation or geopolitical compromise, is a light speed trip to financial disaster or a world where the oligarchs end up making use of those underground bunkers they have been building…

– – – 

Moments after my initial draft of this post, I read two interesting and relevant articles. The first covers how an OpenAI model solved one of the Millenium Prize Problems (in mathematics). This report covers a lot of interesting angles. Two thoughts I want to draw out as they relate to what I wrote above. First, if you are a CIO, your reaction would likely be that’s cool… but so what? Is a model powerful enough to solve the Navier-Stokes model really what you need to improve your business process and go-to-market capabilities? 

Second, the article notes that the estimated cost for solving the problem was $15M. This is an example of the tension between business and research. This may be $15M well spent for the advancement of knowledge (though many of the mathematicians in the article make excellent refutations to that claim) but it is $15M poorly spent for business. A traditional McKinsey consultant would slash this line more quickly than underfunded school districts axe their arts programs.

The second covers an essay published by Dario Amodei, the CEO of Anthropic, urging an AI slowdown. This report (behind a paywall but other sources widely available I’m sure) summarizes his points. According to the summary (the essay itself is 3800 words long, twice as long as this post), his main argument is safety, one of the most important and compelling arguments. But one can’t help but wonder whether Dario Amodei and his company may be the worst actors of all. We hear Anthropic continually publish warnings about the dangers of AI and make the most frequent calls for action on regulation, but they always externalize the need instead of actually doing anything themselves. Once widely seen as an admirable counterweight to the psychopathy of Elon Musk and emotional instability of Sam Altman, Anthropic has proven itself to be just as ego-driven as their competitors, throwing in fistfuls of pious hypocrisy into the mix as well.

– – – 

Note: This post is also available on my AI-focused site at https://dorukai.dorukakan.com/.