Your company runs 23 AI tools. Which ones work?

Your company runs 23 AI tools.  Which ones work?

Somewhere in your company right now, someone in marketing is running a workflow through an AI tool that IT has yet to discover. Somewhere else, a developer just added a second coding assistant because the first one missed a deadline, the professional equivalent of hiring a backup umbrella.

Multiply that pattern across every team, and you get the real state of enterprise AI in 2026: sprawling, expensive, and mostly unmeasured…


The number is bigger than anyone budgeted for

Larridin’s State of Enterprise AI 2026 research puts the average enterprise at 23 distinct AI tools in active use. Only 38% of those companies maintain a complete inventory of what is actually running, leaving the majority managing a stack they can only partially see.

Agents make the picture messier still. Salesforce’s 2026 Connectivity Benchmark Report, built on a survey of 1,050 enterprise IT leaders, found organizations run an average of 12 AI agents today, a figure projected to climb 67% by 2027.

Half of those agents operate in isolation, disconnected from every other system meant to give them context.

Ask a CIO why the count keeps rising, and the honest answer usually traces back to procurement approving only a slice of it, with the rest arriving through department budgets and personal logins.

Bridging the gap from supercomputing to AI factories

A comprehensive industry report on modernizing high-performance computing for production AI, featuring insights from NVIDIA and WEKA leaders.

Shadow AI made the org chart optional

Gartner’s 2025 research found 69% of organizations already carry confirmed or suspected shadow AI somewhere in the business.

A PagerDuty survey of 1,250 office professionals at companies earning $500 million or more found 66% had used an AI tool at work despite believing it violated company policy.

Verizon’s 2026 Data Breach Investigations Report recorded a fourfold jump in shadow AI detections inside a single year. Employees pick tools based on what solves today’s problem, and a formal review process rarely enters that decision.

That is a rational response to a slow procurement pipeline, and it’s the argument behind efforts to turn shadow AI into a managed, agentic workforce rather than banning it outright.

Either way, it explains how a stack reaches 23 tools while security can vouch for only a handful of them. The org chart, it turns out, was more of a suggestion.


Approval is a different question than performance

Bring every shadow tool into compliance and a harder question remains: does the tool actually do the job, and does the agent using it pick the right one for the task in front of it?

Gartner estimates that among the thousands of vendors marketing agentic AI, roughly 130 offer genuine agentic capability. The rest wrap an existing chatbot or RPA product in fresh language.

Even genuine agentic tools fail in a specific, measurable way. An August 2026 arXiv paper from researchers Atul Anand and Sourav Chattaraj tested eight models against 120 tasks using canary tools, deliberately planted traps built to catch tool-selection errors.

Susceptibility to those traps varied by roughly 36 times across the eight models, and the most susceptible hosted model landed in the middle of the capability rankings rather than at the bottom.

That finding punctures a comfortable assumption in procurement meetings everywhere: a higher price tag or a stronger benchmark score buys safer tool selection on its own. The data says otherwise, which is awkward news for every vendor slide with a leaderboard on it.

Your AI agent’s skills are lying to you about why they work

Skills don’t teach your agent much of anything, according to a new 8,135-trial study: only 4.5% of skill use is actual knowledge injection. The rest is mostly the agent using the skill file to stay on track. And the more skills you add, the worse it gets at finding the right one…

What the sprawl actually costs

A stack that has yet to be fully inventoried carries costs beyond the subscription line:

  • Integration debt compounds fast. Every disconnected tool needs its own data pipeline, which is exactly the kind of fragmentation that threatens operational stability in mission-critical ML systems. Salesforce’s benchmark found only 27% of the average enterprise’s 957 applications are actually integrated with each other.
  • Compliance exposure grows with every unreviewed vendor. Data flowing to a tool that skipped review is hard to document under GDPR Article 30 or the EU AI Act’s high-risk category rules.
  • Redundant spending hides in plain sight. Two teams often pay for functionally identical tools because visibility into what the other team already owns stays limited.
  • Governance arrives after an incident forces it, which is the most expensive moment to build it and one of the recurring mistakes AI leaders make with agentic deployments.

A working audit beats a policy memo

Fixing this starts with visibility rather than a fresh approval form destined for the same drawer as the last one. A few habits separate companies with a real handle on their stack from companies still guessing:

  • Build the inventory first. Rationalizing a tool starts with logging it, and usage data should drive that list ahead of a survey.
  • Score tools by outcomes, ahead of usage volume. A tool with heavy adoption and weak task completion is a habit worth reconsidering, rather than a result worth protecting.
  • Test tool selection under pressure, the way canary-tool research does. An agent’s benchmark rank says little about how it behaves once a tool description oversells itself.
  • Set a review cadence with teeth. A tool that fails a quarterly check loses its budget line, regardless of how attached a team has grown to it.

The real competitive edge is boring

Every company in your market has access to roughly the same AI tools. The gap between the companies extracting real enterprise value and the companies accumulating subscriptions comes down to whether anyone actually measures what each tool does once the demo ends.

That work stays unglamorous, and it is also the entire job now. Few people put an audit spreadsheet on a highlight reel, but the spreadsheet is the reason the highlight reel exists at all.


Where the infrastructure question gets answered

Your company runs 23 AI tools.  Which ones work?

Every tool and agent in that stack still runs on compute somewhere, and a stack built on infrastructure designed for a different era struggles to scale no matter how well it gets governed. 

The report Bridging the Gap from Supercomputing to AI Factories, drawing on insights from NVIDIA and WEKA leaders, digs into what modernizing that layer for production AI actually takes.

  • A clear picture of where legacy HPC architecture buckles once agentic workloads start hitting it at scale.
  • Direct insight from NVIDIA and WEKA leaders on the shift from supercomputing-era design to AI factory throughput.
  • A practical framework for the retrofit-versus-rebuild decision sitting under every AI infrastructure roadmap right now.

Get ahead of the crowd. Get your copy today

Scroll to Top