Categories
business technology

Consistency is a challenge in AI

A big source of concern around AI in business applications revolves around trust. Can you trust AI to give you the right answer; or will you be constantly worrying about hallucinations? This will always be a challenge because of another characteristic that I think has as big an impact on business applications: a lack of consistency. Each session is an independent event and each agent is an independent agent. The same context can lead to different results from AI even if in many cases they are small and negligible. However, given that the promise of AI means embedding them into critical systems for business and public sector, even small and negligible drift will cause headaches, and that’s assuming the mistakes stay small and negligible. 

Let’s take the example of email agents, a very common use case for personal productivity. A user writes a set of guidelines for AI agents, gives them access to their inbox, and lets the agent autonomously do their work. Many people build human-in-the-loop into the process so they can review but many more likely prefer to just take everything off of their hands.

For this use case, you would write conditional logic in traditional programming. If an email has such-and-such topic and comes from so-and-so, then provide this message in response. There is no judgment. It is predictable and it is consistent. But it is also limited. Even if you build a more sophisticated engine to characterize emails, you still have a discrete set of cases with a discrete set of outputs.

Using AI for the automation of email responses is a completely different paradigm. There are no algorithms and logic statements are for direction not for machine processing. With AI, you provide a prompt with guidance on how to categorize emails, the types of responses to send, when to send the responses, etc. Your system prompt guidance is likely going to be pretty detailed but very different in character than a discrete set of logic statements. The AI model then uses judgment to decide what to do. Testing the process is also different. You can test as much as you’d like but repeatability is no longer the goal; instead, your measures of success shift from a pass/fail certainty to confidence levels.

Now let’s consider business applications and that most ubiquitous use of AI: the virtual assistant, aka the chatbot. At its most basic, a chatbot passes your message to the LLM, along with any relevant additional context (e.g. policy documents for a customer service bot or an API library for a transactional HR bot). In this case, each message is actually its own independent session so it is impossible to predict how a model will respond. For basic use cases, the risk is fairly low as the instructions are straightforward, the solution is simple, and the response options are limited. It doesn’t matter if the model formulates the phrasing of the response a bit differently each time. For more complex interactions and, beyond chatbots, applications with more sophistication and ambiguity, the lack of consistency can transition from innocuous to critical.

AI is embedded in everything from software that checks application hardening against Federal security controls to weapons guidance systems to resume and candidate screening. These are use cases where the model is making more and more judgment calls. Consistency matters but consistency is not guaranteed. Organizations and implementers must take this into account. They must view AI as inherently unpredictable and build harnesses, safety nets, and governance models around it to protect their business and mission goals. Where AI was once thought to be a death knell for SaaS applications the most foresighted have integrated AI into their products cleverly by making it part of the ecosystem instead of making the ecosystem about AI. Tools like watsonx orchestrate should make a lot of sense to business and IT leaders: these tools not only help with managing integrity, safety, and cost but also consistency within common processes and workflows. 

The consistency challenge becomes even more interesting across models. Not too long ago I had to undertake an activity with one of my AI applications (Shelf, a book recommendation engine, it’s great so check it out) that I never expected. When excitedly describing it to a friend and giving her the link, we discovered the site was down. It turns out the issue was that Anthropic had deprecated the model I was using. I now had to deal with one of the greatest banes of an IT professional: the upgrade.

When you upgrade a traditional application, whether truly old school or SaaS, the core functions more or less stay the same. When you change your AI model, it’s possible everything changes. For me, a newer Claude Sonnet model took 4x as long, interpreted my system prompt differently, and returned more hallucinations. During a debugging session, I made a comment about how inefficient the new model was. Here is Claude’s response:

It’s not that it’s “less efficient” — it’s that extended thinking is genuinely more capable on newer models, which tends to mean more thorough reasoning. The old model was likely under-thinking the scoring algorithm. Sonnet 4.6 is actually doing the work more carefully…

One can’t help but think that Anthropic has a lot of consulting materials in its training set as this thought completely ignores that everything was working exactly the way I wanted in the previous model. In a business context, this would be even more frustrating as the implication is that a more advanced model will always be better but anyone familiar with technology knows this is often untrue. Left with no choice, I ended up moving up to the more expensive Opus model, dealing with new parameters and more “independent” behavior, where oftentimes it would ignore the directions in the system prompt. I also ran into this doozy while debugging:

But Opus will sometimes guess rather than say it doesn’t know

That’s the type of unpredictable and difficult-to-control behavior that will give any programmer, IT architect, project manager, or business leader fits. In my case, the stakes are low: it’s an application I built as a hobby. Put in the context of a business application and the upgrade costs become significant. Not only is the new model 2.5x more expensive than the prior, you have to build more code to manage the unpredictability and different approach to judgment of the more advanced model. The upgrade process is different but perhaps it is comforting to know that even with a technology that few of us understand, you can still experience the same headaches as you did with legacy IT.

The challenge then is handling the lack of consistency and that’s where the harnesses I mentioned earlier come into play. Prompt engineering is critical but experience has shown me that it is not as reliable as we hope: models will oftentimes ignore or miss guidance in the prompt. Reducing the calls to an LLM also helps. There is something to be said for the predictability of logic statements, which is why all commercial chatbot builders have integrated deterministic AI to understand certain intents so you can provide discrete responses. Error handling is also critical. The difference is instead of a catch block like in traditional code, you are building logic that anticipates potential inconsistencies, such as ignoring output formats dictated in your system prompt. 

When using AI for personal use, drifts in consistency are not as important; in fact, they may even be helpful for things like building itineraries for a holiday as you get different ideas in each session. For a business application, consistency is critical and any organization deploying an AI-integrated tool to a customer base risks disaster if they do not account for consistency along with the more traditional concerns of cost, trust, and safety.

– – –

Note: This post is also available on my AI-focused site at https://dorukai.dorukakan.com/.