Agentic BI Without the Hangover: A Software Architect’s Field Guide
Vendors are selling something. Executives are buying something. Whether those two somethings are the same thing is anyone’s guess. Somewhere in the middle, an architect is being asked to make it work without anyone waking up with a headache eighteen months later.
This is roughly the state of agentic business intelligence right now. Past the marketing, the technology works, and the teams doing it well are already getting results. The interesting question for architects is what specifically separates the projects that pay off, and what separates them turns out to be the scaffolding around the model rather than the model itself.
What “Agentic” Actually Means Here
Agentic BI is a specific claim. Instead of a human dragging fields onto a dashboard, an autonomous system reads the question, decides which metrics matter, queries the data, interprets the result, and does something about it. Not “answers.” Acts.
That last word is the whole game. A chatbot that summarises last quarter’s revenue is a useful party trick. An agent that detects a margin slide, flags it to finance, generates the supporting analysis, and opens a Jira ticket against the responsible business unit is a different category of thing entirely. The first is a feature. The second is an actor in your operating model.
This distinction matters because the failure modes are different. A wrong dashboard wastes time. A wrong action moves money. Architecturally, the difference is one we know how to handle. Patterns are clear, standards are converging, and most of the heavy lifting sits in territory architects are already good at.
The Numbers, In Context
The numbers tell a useful story. Deloitte’s State of AI in the Enterprise 2026 report, based on a survey of more than 3,000 director and C-suite leaders across 24 countries, found that nearly three-quarters of companies plan to deploy agentic AI within the next two years. Only 21% have a mature model for governing autonomous agents. Gartner adds the corollary: more than 40% of agentic AI projects are expected to be cancelled by the end of 2027. The cancellations Gartner attributes to escalating costs, business value that never materialises, and risk controls that turn out to be inadequate. None of those reasons are about the technology itself.
Both analysts say the same thing. Deloitte calls it a governance maturity gap; Gartner frames it as hype meeting reality. Either way, the problem isn’t model capability.
Translated from analyst-speak: the LLM is fine. The architecture around it is what separates the projects that ship from the ones that get quietly cancelled in 2027. Good news for us, because that’s the bit we get paid to design.
The Semantic Layer Is the Architecture
Here’s the bit that vendors are now in violent agreement about, even if they’re each selling a different version of it. Agentic BI needs a governed semantic layer. Without one, the agent invents definitions on the fly, and you find out which definitions it picked only when the answer turns out to be wrong.
The reason is mechanical. When a user asks an agent “how’s revenue trending in EMEA,” the agent has to translate that natural-language question into a precise query. If your organisation has fourteen versions of “revenue” scattered across BI tools, spreadsheets, transformation pipelines, and a wiki page nobody’s touched in years, the agent will pick one. It will sound completely confident about the answer. The answer might even be right. You won’t know which, and neither will the agent.
That’s the eighteen-months-later headache. Not a model that hallucinates wildly enough to be obviously broken, but an agent that’s been quietly applying a stale revenue definition for the last quarter, and the variance only shows up when finance is reconciling at year-end. The technology hasn’t failed. The meaning layer underneath it has.
The semantic layer solves this by being the single place where business definitions live. Revenue gets defined once, EMEA once, the fiscal calendar once. One canonical model with governance attached to it from day one rather than bolted on afterwards. Every consumer, whether human or agent, gets the same answer to the same question.
Governance in practice is concrete. Each metric has a named owner, typically in finance or the relevant business function, who signs off on the definition. Changes go through a review process the same way code does. Lineage tracks every metric back through its transformations to the source data, so when an agent makes a decision based on revenue, you can reconstruct exactly what number it saw and how that number was derived. This isn’t new territory. dbt’s semantic layer, Cube, AtScale, LookML, and Databricks Unity Catalog have all supported variations of this for years.
What’s new is that these layers have gone from “nice to have for analytics consistency” to load-bearing for AI safety. For architects, that reframes the procurement question. You’re no longer choosing a BI tool. You’re choosing what defines truth for every autonomous system you’ll deploy in the next decade. That decision wants more than a vendor demo and a quarterly steering committee.
MCP and Skills: The Connective Tissue
The other big shift is that connecting agents to data has quietly standardised. Two open standards now do most of the work, and they cover different problems rather than competing.
The Model Context Protocol is the connectivity layer. Anthropic released MCP in November 2024. Within a year it had been adopted by all the major frontier model providers and donated to the Linux Foundation’s Agentic AI Foundation as neutral infrastructure. That’s the closest thing the industry has to OpenAPI for AI agents. Power BI, dbt, Looker, and most of the major semantic layer platforms ship MCP servers. If you’re building integration architecture for an agent that needs to read or write external systems, MCP is now the default.
Agent Skills, introduced in October 2025 and open-sourced in December, sit alongside MCP and solve a different problem. A skill is a folder containing a SKILL.md file plus any supporting scripts the agent might need. The agent loads only the metadata up front, and pulls the rest into context when the skill is actually triggered. Tell the agent a skill exists; only spend the context on it when the skill gets used.
This matters operationally for two reasons, both of which any architect who’s done integration work before will recognise. The first is context economics. MCP loads every connected server’s tool definitions into the context window up front, and a feature-rich server can take a meaningful bite out of the context budget before any conversation starts. Skills sidestep that. You can give an agent dozens of skills without paying the same context tax. The second is authentication. Each MCP server handles its own auth, so a real deployment ends up with a small zoo of OAuth flows and credential rotations to babysit. Skills run inside the agent’s own execution environment and inherit whatever auth context the agent already holds. They don’t replace MCP, but for procedural knowledge that operates on data the agent can already see, they avoid an entire class of credential-management problem rather than solving it.
MCP teaches an agent what tools exist and how to call them. Skills teach an agent how an organisation actually does things. One is the plumbing, the other is the playbook. A mature deployment uses both, with MCP reaching the governed semantic layer and skills encoding the workflow patterns specific to a team’s reporting cadence or finance close process.
For architects, the integration layer is settling on open standards, and quickly. Decisions you make on this in 2026 should still be the right ones in 2030. That’s more shelf life than most architecture bets get.
Three Things That Catch Teams Out
A few patterns come up often enough across published deployments to be worth knowing about up front. None of them are showstoppers. They’re just counterintuitive enough that mature teams design around them rather than discovering them by accident.
Stale data fails differently with agents. A human reading a dashboard from this morning notices the timestamp and adjusts. An agent doesn’t. Its confidence in a wrong answer is identical to its confidence in a right one, and it acts on both at the same speed. The architectural implication is that your data pipeline SLAs are now agent SLAs, and the tolerances you set for human-readable dashboards are probably looser than what an autonomous actor needs.
Context is a budget, not a feature. Every tool definition you load up front, every schema you expose, every piece of system prompt you stuff in to “help” the agent costs you context. A maximalist approach (give the agent everything, just in case) leaves it short of room for the actual conversation. This is the problem Skills exist to solve, but it generalises beyond Skills. Treat the context window like memory in an embedded system, not like disk space.
Bounded autonomy is the design, not a temporary safety net. “Bounded autonomy” is the phrase the research community uses for agents with explicit operational limits and clear escalation paths. It sometimes gets read as a phase you grow out of once you trust the agent properly. That framing is wrong.
Read-only agents and write-capable agents are different risk profiles, and write-capable agents that touch financial systems are in a different bracket altogether. Treat the boundary the way you’d treat a privilege escalation in a security review: not a phase, an architectural choice with lasting consequences.
The Architect’s Brief
Four points, ordered by how hard each is to reverse later.
Pick a semantic layer with intent. Treat it as a load-bearing wall rather than a furniture choice, and prioritise open standards. The Open Semantic Interchange initiative is worth tracking. Vendor-locked semantic layers in 2026 are a longer bet than they look.
Build governance in from day one. Named owners for every metric, change control for definitions, and lineage that traces every agent action back to a specific version of the model that informed it. If the auditor can’t reconstruct what an agent saw and decided, you’ve got more work to do.
Adopt MCP and Skills together. MCP gives your agents reach. Skills give them procedural memory without the context-window penalty. A vendor that supports both is a vendor whose architecture you won’t have to redo in eighteen months.
Measure something from day one. Time to insight, decision latency, action accuracy, fallback rate when the agent escalates. Whatever you pick, pick it before you deploy. Reverse-engineering ROI from a live deployment is a tax nobody wants to pay.
The Bottom Line
Agentic BI is a real category, and the technology works. The protocols are standardising faster than anyone expected. Architectural patterns are clear, and most of the work plays to architects’ strengths.
The hard part isn’t the AI. It’s that an autonomous decision-making system in your business is only as trustworthy as the meaning layer underneath it, and that meaning layer is built by humans who disagree about what revenue is. Solve that, and the rest is protocols and plumbing. We’ve been doing plumbing for decades. Get the meaning layer right and the agents you build on top of it will be trusted with anything that matters.
- State of AI in the Enterprise 2026: The Untapped Edge — Deloitte AI Institute, January 2026
- Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 — Gartner press release, June 2025
- Google Next ’26 Just Validated Agentic BI Needs a Semantic Layer — Strategy.com, April 2026
- Why Agentic Analytics Starts with a Well-Governed Data Layer — Databricks, April 2026
- Model Context Protocol — official site and specification
- Linux Foundation Announces the Formation of the Agentic AI Foundation — The Linux Foundation, December 2025
- Equipping Agents for the Real World with Agent Skills — Anthropic, October 2025
- Agent Skills vs Model Context Protocol: How Do You Choose? — Ravikanth Chaganti, February 2026
- The Agentic Future Demands an Open Semantic Layer — Salesforce, January 2026