Field Note

When AI Makes Your Data Problem Worse, Not Better

AI tools accelerate execution without fixing portfolio structure. If your data is already fragmented, AI makes the fragmentation faster and more expensive.

Lior BarakBy · Data Portfolio Advisor
7 May 2026·5 min read

Every AI tool your data team deploys lands on top of an existing data infrastructure. If that infrastructure has friction in it, and almost every one does, the AI does not remove the friction. It multiplies it. This is the pattern I keep seeing, and it is expensive enough to be worth understanding before the next AI initiative gets approved.

The scenario that keeps repeating

A company builds an AI agent to help business teams analyze data independently. The pitch is compelling: reduce dependency on analysts, speed up decision-making, give marketing and finance direct access to insights.

Six months later, the marketing team is spending more time on data than before the agent was deployed. Not because the tool is broken, technically, it works. But because the data it draws on has quality issues that were always there, and the AI surfaces them in ways that require human verification before anyone trusts the output. The team has added a QA layer on top of the AI layer on top of the existing data layer.

The tool created work instead of removing it.

A second version of the same problem

Self-service analytics built on a language model is another version. The idea: let business teams ask natural language questions directly against the data warehouse. No SQL required. No analyst in the loop.

What actually happened at one organization I know of: users started running queries they did not fully understand. The LLM generated responses and also generated database tables, thousands of them, associated to queries that produced no actionable output. Storage grew. Compute costs grew. Token costs for the LLM context grew. The IT team spent weeks cleaning up tables that had accumulated with no owner and no purpose.

The cost was not just the tool subscription. It was the hidden infrastructure cost of an AI system operating on data that was not structured for it.

Why this keeps happening

AI tools are amplifiers. They amplify what is already there, including the problems.

When a business team spends fifteen hours a week manually quality-checking data before they can use it, that is a signal. The data product does not fit how they actually work. An AI layer placed on top of that product does not fix the fit problem. It adds a new surface area for the same friction to appear.

The organizations that get real value from AI on their data are the ones that have already done the upstream work: clean data models, documented business logic, products that are actually adopted and trusted by the teams using them. AI in that environment moves fast. AI on top of unresolved infrastructure debt moves backward.

The question to ask before any AI data initiative

Before approving an AI tool or AI-driven analytics project, there is one question that will tell you more than any vendor demo: what does the business team currently do with the data output before they make a decision?

If the answer involves spreadsheets, manual reconciliation, a second check against another source, or a conversation with an analyst to verify the numbers, stop. That is the problem to solve first. Adding AI to that workflow will make the workflow longer, not shorter.

The measure of whether data infrastructure is ready for AI is whether the people using it trust it enough to act without a verification step. Most organizations are not there yet, and the path to get there has nothing to do with AI.

What the cost looks like

The cost of premature AI deployment is not always visible as a line item. It shows up in:

  • Business team hours spent on QA that predates the AI tool but intensifies after it
  • Infrastructure costs that were not in the AI project budget (storage, compute, token spend)
  • Data team hours spent cleaning up outputs the AI created without intent, tables, queries, cached results that accumulate with no governance
  • Deferred trust, where a team's confidence in the data declines because the AI surfaced errors they had previously worked around

None of this shows up on the AI tool invoice. It spreads across the existing operational cost and becomes invisible to the budget conversation.

The honest calculation

If your data team is still spending meaningful time on reactive support, incidents, ad-hoc requests, quality issues, and if your business teams have workarounds built into their daily workflows to compensate for data that does not quite fit, AI will not change that. It will add to it.

The sequence that works is: reduce the hidden cost of the existing portfolio first, establish that business teams trust and use the products already built, then extend with AI into a system that is already functioning. That order matters.

Going the other way, using AI to leapfrog infrastructure problems, is an expensive way to find out the problems are still there.

If you recognize this pattern

If you have already deployed AI tools on your data and you are seeing more friction rather than less, a 30-minute conversation can usually identify where the infrastructure gap is and what it would take to close it before the next phase of investment.

Book that conversation here. No slides, just a direct look at what is actually happening.

Further reading