Why every enterprise AI use case will keep a human in the loop

As enterprises push AI from pilots into production, the debate over autonomy versus oversight is getting louder. Atul Arya, Founder & CEO of Blackstraw AI, has spent a decade building AI system for Fortune 500 clients, and he argues there is no hard line between the two. Human-in-the-loop, he says, is part of deployment itself and gives enterprises the confidence to go live. In this conversation with CIO&Leader, Arya explains why accountability for agents rests with vendors and enterprises, not with the technology, and how poorly grounded data can break guardrails. He also discusses the case for a regulatory framework built on possibility, why on-prem infrastructure suits India’s data, and why electricity and practical talent will define AI’s next phase.

AI
Atul Arya, Founder & CEO of Blackstraw AI

CIO&Leader: Where do you draw the line between AI autonomy and human oversight in mission-critical systems?

Atul Arya: There isn’t a hard line there. We have been working on AI systems now for the last 10 years. And even before that, in various companies that I worked for, human in the loop, or HITL as we call it, human in the loop was always there.

You will always have human in the loop. Separating these into use cases that can be 100% no-touch and those that will have heavy human-in-the-loop is a misnomer.

Every use case will have a human in the loop for the foreseeable future, and it should. Because that’s how enterprise gets confidence that these use cases are viable and can eventually really see the light of day.

So I do not know there is a hard line. It should be considered as part of AI deployment. Human-in-the-loop should be considered part of AI deployment. I will still have a challenge.

CIO&Leader: Blackstraw works with Fortune 500 enterprises across financial services, healthcare, manufacturing, retail and logistics. What’s driving the shift from GenAI pilots to production-grade, autonomous AI systems in these sectors?

Atul Arya: Every ecosystem is a little bit different. So if you think about retail, CPG, and market research, these are industries that have been the longest users of AI.

Because, as you can imagine, versus a banking system or a healthcare system or an airline system, these are industries that are more prone to experimentation. So retail, CPG, and market research are slightly more prone to experimentation. And hence, AI has always been slightly more advanced in terms of use cases and capabilities.

What’s really happening at this point in these industries is that people are ready to go to the next level of autonomy. Let’s say I was in a store and I had earlier robots that were detecting out-of-stock items and saying, “Can you just tell me what’s out of stock?” And then a replenishment team would come and replenish the product.

They’re going to the point where they’re saying, “Hey, I just don’t want to know what’s out of stock. I would also like to know where the product is and if there’s an automated way of getting it to the shelf.”

So I think what’s really happening in these industries is that the shift is happening just from analytics to execution using AI. And that’s driving a lot of use cases.

So it’s becoming part of an end-to-end workflow rather than siloed individual workflows. So that’s the situation with, let’s call it, retail, CPG, and market research—one segment.

The other situation is where, in financial services and healthcare, it’s very clear that you cannot have AI right now. So while there was resistance in those ecosystems over a period of time, saying, “Hey, these are very core ecosystems, and people should be careful before automation happens,” for them, in the last couple of years, it has been very clear that these verticals will be left behind if AI is not infused.

There’s a start of getting that data in order, starting to do some basic use cases, building confidence in the enterprise, and then taking it to the next level; different industries are in different—different verticals are in different phases.

Some are like, “Hey, because they’re more used to it and it’s less risky, they’re going much, much faster with their use cases.” Some have imbibed the idea that “Hey, you know, AI is here to stay, and we need to make it meaningful.”

They are getting on the bandwagon and getting new use cases under their belt.

CIO&Leader: When an AI agent independently plans and executes a multi-step task, which is accountable if something goes wrong: the enterprise, the vendor, or the AI system itself?

Atul Arya: You cannot make the agent or the model responsible. It doesn’t exist. In the end, the responsibility always lies with the person who was involved, or eventually involved, in bringing that agent.

It’s a shared responsibility between the vendor and the enterprise, right? There is no accountability with the agent itself.

The vendor and the vendor team that is building it, and the enterprise that is using it or has chosen the vendor, are all collectively responsible for making sure that there are guardrails, there is technical efficiency in building it, there’s a human in the loop, and there’s enough testing before it is rolled out.

The vendor and the enterprise are equally responsible for making this happen. But definitely not the agent. It doesn’t exist. You can’t jail the agent, right? So that’s how it is.

CIO&Leader: What does trustworthy enterprise AI architecture actually look like in practice beyond the buzzwords of “explainability” and “governance”?

Atul Arya: What we discussed, right, a little bit about trustworthiness decreases cost. You always have to have—yes, you have to have explainability models, which you mentioned are more like buzzwords, but those are important too.

In the end, I do feel that confidence in the enterprise comes from what is appropriate—and appropriate, underline appropriate—because you cannot have an extreme amount of human in the loop, because then it’s as good as, you know, what was there before.

An appropriate amount of human in the loop is absolutely needed to build confidence, and that’s as important as explainability, right?

Because explaining those two things—explainability and human-in-the-loop—to me are extremely important for building confidence.

CIO&Leader: What guardrails have you seen work or fail when AI agents operate in high-stakes, mission-critical environments?

Atul Arya: The guardrails that fail when you’re trying to use cases that haven’t been tried or tested before are around the fact that you essentially start having hallucinations on core processes that are not grounded in training data from the enterprise, right?

When you see—so think about how a model is trained or deployed, right? It has two potential data sources. One is your enterprise data, which you can use to ground your model. The other is internet data, third-party data, whatever, right?

Now, if you ground your models more in internet or third-party data, the hallucinations will be much more than they would be with grounding in your enterprise data. The first guardrail that really starts failing is that, hey, you know, I created my weights correctly, I did my training correctly, but because I had a bias in how the data was sourced for my training, I’m going to get a level of hallucinations.

And you would still see explainability, because the model is generating explanations based on the data it was trained on. You’re going to start seeing hallucinations, and you’re going to see a lot of explainability, which will make you feel it is true. And that’s where human-in-the-loop is extremely important, right?

That’s why it’s important that, hey, you know, when that guardrail breaks because you sourced it from data that itself was not pristine, then you need a human-in-the-loop to make a judgment call saying, “Hey, you know, this is a breakage of my guardrail, and we need to do something about it.”

CIO&Leader: Could a clearer regulatory framework actually accelerate enterprise AI investment in India, rather than slow it down?

Atul Arya: Is there anything that brings in more clarity? Because, you see, which, again, I go back to what I started with, there are two parts of this: creating the frontier model, and then there’s AI adoption, The AI adoption by enterprises.

Enterprises are not going to go fast and furious until they have regulatory backing behind them, understanding as to, you know, what’s happening, and a clear-cut way that, hey, whatever they’re doing is not suddenly going to be taken away.

So, generally, there’s confidence when the government gets involved, and frameworks are laid. People start trusting and using them slightly better. So it’s not a bad idea to have governments and frameworks.

The only thing is, it should not go to the other extreme, where you make it so stifling that whatever is possible suddenly comes to crashing speeds, you know, crashing rates. Saying, “Okay, hey, you can’t do this, can’t do this, can’t do this, can’t do this.”

It has to be a framework that comes from the standpoint of: if you need to do this, then what should you do? Rather than creating a bunch of guidelines on what cannot be done. That’s important. Frameworks, guidelines, all of that is great, but it has to come from a mindset of possibility rather than curbing people from doing something.

CIO&Leader: How should India balance innovation with data sovereignty and privacy as it builds out its AI governance regime?

Atul Arya: Data is critical. India-centric data may not be available anywhere else in the world. You know, the volume of data, the type of data, the situations of data. And in India, you would have to figure out how to build data centers to house and manage India-specific data.

Having large cloud hyperscaler systems may not be the right approach for the volume of data we’re talking about in India and its use cases.

Privacy, once you have it housed internally and you’re on-prem, it automatically provides for that. Obviously, you’ll need your standard security protocols built on on-prem servers.

And India is heading in that direction. There are many data center needs, and an understanding that data must be housed within the ecosystem.

Keeping it primarily on-prem manages both costs and privacy concerns. So that would be the way to go for India.

CIO&Leader: What does the next phase of AI adoption demand in terms of talent, infrastructure and technology investment?

Atul Arya: Yeah, from a technology standpoint, people have to get more focused on, you know, for the longest time, when you think about AI, people are like, “Okay, how do I measure the model? How good or bad is it?”

That conversation should go away, and people and talent should be trained on industrial use cases, real use cases, how to deploy these use cases, or how to deploy these models.

And less on the theory and more on the practicality of adoption is where a lot of training needs to happen when it comes to talent.

And colleges, early stage, you know, when people get into early-stage companies, we are all responsible for getting them real data and real use cases to train on, and not just theory. That’s extremely, extremely important. Infrastructure-wise, the limiting factor to this is electricity. Electricity is the key to this.

You have to figure out alternative ways to generate electricity, or you go to the extreme, as with what SpaceX is doing, where you’re trying to move the entire data center into space.

That is extreme, but for the foreseeable future, we really have to find alternative ways to generate electricity.

You know, data centers are going to be important. Alternative ways of generating electricity are definitely needed, along with renewed investment.  Technology is enough; enough is happening already. In fact, it’s going too quickly.

What we really need is, you know, using that technology. You see, technological advancement will automatically happen. We don’t need to worry about it.

You know, there are the big players—OpenAI, Anthropic, and, you know, what’s it called, Grok. There’s a whole lot of stuff. They’re spending billions and trillions of dollars trying to drive technology advancement.

We need to make sure that whatever is available, which already surpasses what’s needed for tech deployment, is utilized by enterprises and delivers ROI.

As long as we can do that, we’ll be fine. More investment is needed in talent and infrastructure, and technology will take care of itself.

Share on