Guest Post: AI Doesn’t Eliminate Technical Debt. It Inherits It.

By Don Boxley, CEO and Co-Founder, DH2i (www.dh2i.com)

AI has been dominated by one question for the last two years: How smart is the model?

As a place to start, it’s understandable. Capabilities that were difficult to imagine just a few years ago have been unlocked by better models. Better reasoning, lower costs, faster inference, or more natural conversations are all now promised by every new release.

But, I think… we’re reaching an interesting turning point. The question is changing as AI moves from pilot projects into everyday business operations. Organizations aren’t asking whether AI can generate a better answer. They’re now asking whether they can depend on it. A very different problem, right?

Every Successful AI Project Eventually Becomes an Operations Project

The first version of an AI application is usually built by developers. The second version is owned by operations. That’s true of almost every technology we’ve adopted over the last thirty years. Building something is one challenge. Running it every day is another.

Once an AI app starts supporting customers, approving transactions, helping clinicians, or assisting employees, reliability becomes part of the product. Users don’t separate the model from the application. They simply expect it to work. If it doesn’t, nobody blames the language model. They blame the business.

That’s why I believe the next phase of AI won’t be defined by model improvements alone. It’ll be defined by how well organizations operate the infrastructure underneath those models.

AI Doesn’t Replace Existing Infrastructure

One assumption that keeps showing up is that AI somehow gives organizations an opportunity to start over. It doesn’t.

Most enterprises are introducing AI into environments they’ve spent years refining. They’re not building greenfield environments. Critical databases already exist. Business applications already exist. Windows servers, Linux systems, Kubernetes clusters, public cloud services, private cloud infrastructure, and edge deployments all exist. AI has to live with all of it.

That means the challenge isn’t replacing existing infrastructure. It is making existing infrastructure work together in ways it wasn’t originally designed to.

The Hardest Problems Aren’t AI Problems

The answers rarely have anything to do with model accuracy, when you ask an operations team what worries them. They worry about downtime. They worry about planned maintenance becoming unplanned outages. They worry about databases staying available. They worry about security. They worry about recovering quickly when something breaks. Those concerns haven’t changed because of AI. However, they have become more important.

The less tolerance there is for operational failure, the more business decisions depend on AI.

If an AI app isn’t available during peak business hours – that isn’t an AI problem. Availability is the problem. If an AI app can’t securely access the data it needs – that isn’t an AI problem. Architecture is the problem.

Architecture tends to outlive individual technologies – and that’s an important distinction. Today’s models will eventually be replaced. Good infrastructure should make that tech replacement almost invisible.

Hybrid Isn’t a Compromise

Hybrid infrastructure was treated like a temporary stop on the way to something else for years. I’m not convinced that’s true anymore. Different workloads have different requirements, and that’s why organizations are choosing hybrid.

Development teams may prefer Kubernetes. Production databases may remain on VMs. Sensitive data may stay on-prem. Inference may run closer to users.

Those aren’t signs that modernization has failed. They’re signs that organizations are making practical decisions instead of ideological ones. Infrastructure should support that flexibility rather than fight it.

The Conversation We Should Be Having

When AI discussions begin with model selection, infrastructure often becomes an afterthought. It should be the opposite.

Organizations should understand how they’ll keep the app available, how they’ll protect the data feeding it, how they’ll automate recovery, and how they’ll support the workload as it inevitably moves between environments–all before deciding which model to deploy.

Of course, those are not glamorous conversations. But, they are the ones that determine whether AI becomes a dependable business capability or another isolated technology project.

What You Can Do Tomorrow

If your organization is planning an AI initiative, don’t begin your next meeting by asking which model to use. Instead, ask your architects and operations teams a few different questions:

  • What happens when the database goes offline, especially, if this AI app is business-critical? (The ideal answer: It shouldn’t stay offline. With minimal disruption, critical workloads should automatically fail over to a healthy node. Availability should be built into the architecture. It should not depend on someone manually restoring service.)

  • If an app, cluster, or site fails, how quickly can we recover? (The ideal answer: Recovery should be measured in seconds or minutes – not hours. Automated detection and failover should reduce downtime and eliminate manual intervention wherever possible.)

  • Without having to do a major redesign, can this workload move between environments? (The ideal answer: Yes. Applications should be portable across physical servers, VMs, Kubernetes, cloud, and edge environments. Without requiring the application itself to be rewritten. Infrastructure should adapt to the workload. Not the other way around.)

  • From day one, are we designing security, availability, and automation into the architecture, or are we planning to add them later? (The ideal answer: They must be integrated from the beginning, because retrofitting resilience and security after deployment is more expensive, more complex, and introduces unnecessary risk.)

  • Which parts need to evolve? And, which parts of our existing infrastructure already solve these problems/challenges? (The ideal answer: Keep what already works. Modernize only where it creates business value. AI should leverage existing investments whenever possible instead of forcing wholesale infrastructure replacement.)

Again, agreed… those conversations may not generate headlines. But they’ll have a far greater impact on whether your AI initiatives succeed over the next five years.

Technology will continue to change… Models will improve… New platforms will emerge… That’s the easy part.

However, organizations won’t lead their respective markets and create lasting value from AI because they chased every new breakthrough. They’ll be the ones that build operational foundations that let them adopt new technology without disrupting the business.

I think we can all agree, that’s actually what good infrastructure has always done.

AI simply gives us another reason to get it right.

Leave a Reply

Discover more from The IT Nerd

Subscribe now to keep reading and get access to the full archive.

Continue reading