Nine layers. 170+ companies. 8 standards. Updated July 2026.
How to read this
Every entry on the graphic is here, with what it does and why it earned the slot.
Two things before you start.
This is not a quality ranking. I ranked by news cycle share, public valuation, and enterprise footprint, which are proxies for relevance and not for whether the software is any good. Some excellent tools are missing because nobody wrote about them this quarter. Some mediocre ones are here because everybody did.
Category leader also means something narrower than it sounds. It means that if a mid-market company had to pick one vendor in this layer today, with limited time to evaluate, this is where I would tell them to start. Not best product. Not market share leader. Three of my leader picks are open-source projects with no sales team.
The graphic uses four colors for four buying motions, three shapes for three entity types, and a set of markers for deployment model and name confidence. If you have the poster in front of you, the color is the useful part. It tells you which budget a thing comes out of.
I write for companies between roughly $50M and $500M in revenue. Everything below answers the question of what a company that size actually does with this.

Band 1: Infrastructure spend
Power, capacity, and tokens. Metered, volatile, and the largest line item.
Layer 1: Silicon and Hardware
The physical floor. Every token you generate is a matrix multiply that happened on one of these chips, and the cost of that multiply flows straight through to your API bill.
You are not buying chips, though. You are buying the downstream effects of chip supply. Competition here is the largest single driver of falling inference prices, and inference pricing decides whether your AI feature has a gross margin at all. So when AMD lands gigawatt-scale commitments from OpenAI and Meta, that is not really tech news. It is your 2027 budget.
Category leaders
NVIDIA. Fiscal 2026 revenue of $215.9B, up 65 percent, with $193.7B of it from data center, and roughly $1 trillion in locked Blackwell and Vera Rubin orders through 2027. Vera Rubin is a rack-scale system, not a chip: the NVL72 pairs 72 Rubin GPUs with 36 Vera CPUs, claiming around 10x Blackwell's inference throughput per watt at a tenth the cost per token, shipping in production this fall. CUDA is still the moat, and every alternative on this list gets measured against it.
AMD. Earned the promotion this year. Helios entered full production at Advancing AI in July: 72 MI455X GPUs, 31TB of HBM4, 2.9 exaflops of FP4, all-Ethernet networking instead of proprietary interconnect, with a claimed 30 percent advantage in tokens per dollar over Rubin NVL72. The spec sheet is not what convinced me. The customer list is. OpenAI at 6GW with warrants for up to 10 percent of AMD stock, Meta at 6GW, Anthropic at up to 2GW plus as much as $5B in equity, Microsoft anchoring Helios for Azure inference, Oracle planning a 50,000-GPU cluster. That is a backlog, not a roadmap.
Broadcom. The company most buyers have never thought about, quietly designing everyone else's chips. Roughly 70 percent of the custom AI accelerator design services market. Google's TPU partner since 2016 under an agreement running to 2031. Co-designer of OpenAI's 10GW custom accelerator program. Now selling assembled Ironwood racks directly to Anthropic. If you want to know who wins when hyperscalers stop buying merchant GPUs, this is the answer.
Notable players
Google (TPU). Ironwood, the seventh generation, is the first custom ASIC to reach seven-figure deployment at a single customer. Anthropic runs more than a million of them. Google now sells systems externally instead of only renting them, and the eighth generation splits into separate training and inference chips at TSMC 2nm.
AWS Trainium. Trainium3 anchors a large share of Anthropic's training and gives AWS a real cost story against NVIDIA inside Bedrock. Less flashy than TPU, more embedded in enterprise procurement.
Marvell. The second custom-silicon design house. On the list mostly as the hedge against Broadcom concentration risk, which is real given how much of the ASIC market runs through one vendor.
Groq. Correct the record here. NVIDIA licensed Groq's LPU technology for roughly $20B, it did not acquire the company, and Groq raised $650M in June to rebuild as an independent inference cloud. The tech now ships inside Vera Rubin as Groq 3 LPX.
Cerebras. Wafer-scale, still the fastest thing available for certain single-model inference workloads, still narrow in model coverage. Expanded its AMD relationship in July.
Microsoft Maia. Maia 200 shipped in early 2026 with Maia 300 in design. Behind Google and Amazon on custom silicon, though Microsoft has the demand to close that gap.
Meta MTIA. In production on ranking and recommendation workloads, with a custom MI450-generation accelerator in development alongside AMD.
MediaTek. The dark horse. Won Google's cost-optimized inference silicon by delivering Ironwood I/O modules 20 to 30 percent cheaper than the alternatives, and has reportedly asked TSMC for a sevenfold increase in advanced packaging capacity.
Qualcomm. Owns the phone, and by extension the default answer for on-device inference at consumer scale.
Apple Silicon. The cheapest credible way for a small company to run a capable open-weight model with zero data egress. A workstation-class Mac handles a surprising amount of internal work at no marginal cost.
Hailo. Industrial and camera inference. Here because edge AI is a real mid-market use case that the frontier conversation ignores completely.
Intel. Still struggling, with no credible answer to Helios or Rubin before 2027 at the earliest. I included it so you know it is not on the shortlist, instead of wondering whether I forgot.
Layer 2: Cloud, Capacity and Power
Where you rent the chips. Increasingly, where you rent the electricity.
Most of your AI money goes here, and so does your lock-in. Model choice is reversible in an afternoon. Data gravity is not. The single most useful planning fact I can give you: Alphabet raised 2026 AI capex guidance to somewhere between $195B and $205B, then told investors it would buy third-party capacity because its own buildout cannot keep pace. The four largest hyperscalers have committed roughly $700B for the year. Capacity is the constraint now, not capability.
Category leaders
AWS. Bedrock is the broadest multi-model catalog in enterprise procurement, SageMaker covers the build side, and Trainium3 gives AWS a cost lever nobody else has at that scale. The advantage is not that AWS is best at AI. It is that AWS is already in your contract, and Bedrock is the only place to run Claude with AWS enterprise indemnification.
Microsoft Azure. Deepest OpenAI relationship, Foundry as the model and agent platform, and Provisioned Throughput Units for teams that need predictable billing instead of a surprise invoice. Microsoft has also committed roughly $60B across CoreWeave, Nebius, and Nscale, which is how it adds capacity without adding capex. If your company lives in Microsoft 365, the integration argument is close to unanswerable.
Google Cloud. Q2 2026 cloud revenue of $24.8B, up 82 percent, against a $514B backlog. Gemini runs natively on Google's own silicon with nothing in between, which is why GCP keeps undercutting the field on sustained inference. Best price-performance of the three, with the caveat that Google is capacity-limited by its own admission.
Notable players
CoreWeave. The largest neocloud. 2026 capex guided as high as $35B against a $66.8B backlog, a $21B Meta commitment, and Anthropic added as a customer in April. Also carrying roughly $29B of debt, with interest expense approaching 27 percent of revenue. The growth is real. So is the leverage.
Nebius. Q1 revenue up 684 percent year over year to $399M, capex guidance raised to somewhere between $20B and $25B, and a Meta deal worth up to $27B. Fastest grower in the category, off a smaller base.
Oracle Cloud Infrastructure. Consistently underrated in AI conversations and consistently present in the large deals, including a planned 50,000-GPU AMD cluster. Worth a look if you already run Oracle Fusion, NetSuite, or an Oracle database estate.
Lambda. GPU cloud with a developer-friendly front door. A reasonable middle ground between hyperscaler pricing and running your own hardware.
Crusoe. Built its position on stranded-energy siting, which tells you most of what you need to know about what this layer is really selling.
Nscale. Second-tier NVIDIA Cloud Partner with Microsoft commitments behind it. Included as evidence of how much hyperscaler capacity now sits on other people's balance sheets.
Fluidstack. Handles on-site deployment for Anthropic's own TPU racks. A capacity operator that also does the physical work, which is rarer than it sounds.
Meta Compute. Announced in early July: Meta selling its excess AI capacity. Every neocloud stock dropped double digits on the news. If a hyperscaler with 6GW of AMD commitments becomes a seller, the whole neocloud margin story gets harder.
Adjacent supply
These two sit in their own sub-row on the graphic because they are not peers to the neoclouds above, and pretending otherwise would mislead you.
Hetzner. Still the best price-performance in Europe for inference workloads, and still the right answer for anyone who does not need enterprise-grade everything. I run production agent infrastructure there. It does not compete with CoreWeave and I am not going to claim it does.
IREN. Power and infrastructure supply, converting mining infrastructure to AI capacity. A power play, not a cloud platform.
The on-premises math, updated
The old rule of thumb said owning hardware beats cloud past $2,000 to $3,000 a month. That number should come down for inference-only workloads and go up for anything involving fine-tuning. The honest 2026 framing: buy hardware for privacy and predictability, rent for peak and for frontier capability. Cost is now the third reason on that list, not the first.
Layer 3: Inference and Model Serving (NEW)
The gap between a model existing and a model answering your users at a price you can live with.
In March you could reasonably file model serving under MLOps. Not anymore. Between January and July, Fireworks went from a $4B valuation to $17.5B, Together closed $800M at $8.3B, Baseten went from $2.15B to $13B in about twenty months, and Modal raised $355M. Roughly $45B of enterprise value created in one calendar year, in a category that did not have a name in 2024.
For you, this layer determines gross margin. Pick wrong and you pay more for tokens than you charge for the product. Pick right and open-weight models absorb 80 percent of your traffic at a fraction of frontier pricing, while you escalate only the hard requests. Together reports customers cutting inference costs by up to sixty-fold versus closed models. Several vendors here will also run inside your own VPC, which settles a data residency argument that no frontier API can settle.
Category leaders
Fireworks AI. Raised $1.505B at a $17.5B valuation in July, with ARR crossing $1B, up roughly 5x year over year, reportedly processing more than 10 trillion tokens a day. Customers include Cursor, Perplexity, Notion, Shopify, Uber, and DoorDash. This is the proof that open-weight inference is a real business and not a hobbyist preference.
Together AI. Closed $800M at $8.3B in July with annual bookings around $1.15B. Broader than Fireworks: hosting, fine-tuning, training, and inference in one place. The right pick when what you want is an open-model cloud, not just a token endpoint.
Baseten. Went from roughly $200M to $600M in annualized revenue inside a single quarter, then closed $1.5B at $13B in June. Asset-light, orchestrating GPU capacity across about twenty providers. Self-hosted and hybrid VPC deployment is the differentiator, so if you have a compliance function, shortlist this one first.
Notable players
Modal. Python-native with roughly one-second cold starts and a $355M Series C. Best developer experience in the category, and increasingly it straddles inference and agent infrastructure both.
Groq. Post-NVIDIA-license, rebuilt as an inference cloud with $650M raised in June. Still the latency leader for supported architectures. It appears in Layer 1 as well, because it genuinely is both.
Cerebras. Wafer-scale speed, narrow model coverage. Right answer for one specific high-volume latency-critical workload, wrong answer for a general platform. Also in Layer 1.
Replicate. The developer and community end of the market. Best for experimentation and multimodal breadth.
vLLM. The open-source serving engine underneath a large share of everything else on this list. If you self-host, you are probably running it whether you know that or not.
SGLang. The other major open serving runtime, strongest on structured generation and complex prompting patterns.
Ollama. Local inference, and the default way a developer now runs a model on a laptop. A lot of mid-market prototyping starts here.
OpenRouter. One endpoint, hundreds of models, one bill. A practical answer to a market where four frontier models shipped in three weeks.
LiteLLM. Open-source gateway and proxy. Same problem as OpenRouter, solved on your own infrastructure.
Portkey. Commercial gateway with guardrails, caching, and observability attached. The enterprise version of the routing argument.
Amazon Bedrock. The hyperscaler serving tier. Nearly 100 models, enterprise indemnification, and prompt caching that delivers up to 90 percent savings on repeated context. More expensive than the specialists, and already in your MSA.
Microsoft Foundry Models. More than 11,000 models with instant benchmarking and intelligent routing. The Azure-native path of least procurement resistance.
Google Vertex AI. The GCP equivalent, strongest if your data estate already lives there and you want Gemini without an intermediary.
Band 2: Data and model stack
Where switching cost accumulates. Your data has gravity. Your model choice should not.
Layer 4: Data and Context Infrastructure
Where the answers come from. Models do not know your business, and this layer is how they find out, plus how you prove afterward where an answer originated.
The structural change since March is that standalone vector databases stopped being a category and became a feature. Postgres won the default. Snowflake bought Crunchy Data for about $250M, Databricks bought Neon for about $1B, and both platforms shipped native vector search. The question is no longer where data sits. It is what context an agent gets handed at request time, and with whose permissions.
This is still where most AI projects die, and rarely because the model was bad. Retrieval is bad. Permissions are wrong. Nobody can explain where an answer came from. Glean's pitch is instructive on the economics: its context graph reportedly cuts enterprise AI token spend by about 30 percent. Better context is cheaper, not only more accurate.
Category leaders
Databricks. Unity Catalog for governance, Mosaic AI Agent Bricks for building agents next to the data, native vector search, and Neon underneath for serverless Postgres on agent workloads. The most capable platform in the layer and the most demanding. It assumes you already have engineering maturity.
Snowflake. Cortex AI is in active use across more than 9,100 accounts. Cortex Analyst for text-to-SQL, Cortex Search for vector search over your own tables, Cortex Agents for multi-step work, plus hosted Claude, Llama, and Mistral. The pitch is one sentence long and it lands in regulated industries: the data never leaves.
PostgreSQL + pgvector. Not a company, which is exactly the point. Two platform giants spent $1.25B combined betting that Postgres is the default database of the agentic era. For a mid-market company already running Postgres, adding pgvector is a weekend of work instead of a procurement cycle. This has been my default recommendation for a while and I have not seen a reason to change it.
Notable players
Pinecone. The easiest managed vector database to stand up. The trade is that your embeddings live on someone else's servers.
Qdrant. My pick for a standalone self-hosted vector store. Open source, runs on one box, keeps your data on your infrastructure, with a managed cloud available if you change your mind.
Weaviate. Best hybrid search in the category, managed or self-hosted.
Milvus. For genuinely large distributed deployments. More operational overhead than most mid-market teams should sign up for. Zilliz is the commercial entity behind it.
Chroma. Prototyping and embedded use. Fastest path from idea to a working retrieval demo.
LanceDB. Quietly became the right answer for multimodal and on-disk workloads, where the corpus is too big for memory and too small for a cluster.
Turbopuffer. Object-storage-backed vector search, which changes the cost curve for large cold corpora. Interesting, not yet proven.
Neon. Serverless Postgres, now inside Databricks. It is on the list because it explains why Postgres suddenly matters, not because you would evaluate it standalone.
Crunchy Data. Enterprise Postgres, now inside Snowflake. Same reasoning.
Fivetran. Managed ingestion. Unglamorous and load-bearing.
dbt. Transformation and semantic modeling. Nothing above works if your definitions are inconsistent, and in most mid-market companies they are.
Atlan. Governed metadata as an agent context layer. Early, but it articulates where this layer is going better than anything else I have seen.
Layer 5: Foundation Models
The base intelligence, and the fastest-moving part of the stack. Also the part that deserves the least loyalty.
Model choice sets your ceiling on quality and your floor on cost. The more important 2026 development is that the gap between frontier and open weights compressed to a matter of months on most enterprise tasks. Frontier is for hard, ambiguous, high-stakes work. Everything else should run somewhere cheaper.
July alone brought Grok 4.5 on the 8th, GPT-5.6 and Sol, Kimi K3 at 2.8 trillion parameters, and Claude Opus 5 on the 24th. Four frontier models in three weeks. The right response is not to pick one. It is to build so that picking one is cheap to undo.
Category leaders
Anthropic. Run-rate revenue past $30B, up from roughly $9B at the end of 2025. The current family is Mythos 5, Fable 5, Sonnet 5, and Opus 5. Opus 5 shipped July 24 at $5 per million input and $25 per million output, half the cost of Fable 5, and tops Artificial Analysis on both the Intelligence and Agentic indexes. My primary enterprise recommendation. The pattern worth noticing is that Anthropic now competes with itself on price per task instead of on benchmark deltas, which is what a maturing market looks like.
OpenAI. GPT-5.5, then the GPT-5.6 family, then GPT-5.6 Sol as the reasoning flagship, which takes the top spot on LiveBench Mathematics and Reasoning and hit 93 percent on ARC-AGI-2. Largest ecosystem, most aggressive enterprise push through Frontier and AgentKit, and a free tier that now defaults to a model that would have been frontier eighteen months ago.
Google DeepMind. Gemini 3.5 and 3.6 Flash are generally available and remain the best price-performance option in the stack for high-volume work, with native multimodality across text, image, video, and audio in one model. It runs on Google's own silicon, which is why the pricing keeps working when competitors' pricing does not.
Notable players
xAI. Inside SpaceX since the February merger. Grok 4.5 shipped July 8 at $2 per million input and $6 output, more than 60 percent below Opus 4.8 and GPT-5.5, scoring 83.3 percent on Terminal-Bench 2.1. The efficiency claim is the interesting part: roughly a fifth the output tokens of Opus 4.8 for comparable work, which matters enormously once agent sessions run long.
Meta (Llama). Llama 4 has the deepest Western ecosystem and the longest context, but the community license restricts EU use and it is not an OSI open source license. Meta's strategic energy has visibly shifted toward infrastructure. No longer the default open answer.
Mistral. Mistral Large 3, Small 4, Medium 3.5. The European sovereignty answer, and the only one most EU procurement teams will accept without a fight.
DeepSeek. Anchors the price floor. V4 Flash lists at $0.14 per million cache-miss input tokens and $0.28 output with a million-token context. DeepSeek's own claim is that its open weights trail the closed frontier by months.
Alibaba (Qwen). The widest family and the most permissive tiers. Qwen3-Coder-Next at 80B total and 3B active under Apache 2.0 is the most interesting self-hosted coding agent model available. Watch the label, though: Qwen3.7-Max, the top performer, is proprietary and API-only. Two different things under one name.
Moonshot (Kimi). K2.6, K2.7, and K3, the last at 2.8 trillion parameters. Native multimodal agentic models built for sub-agent parallelism. The autonomous-execution claims are aggressive and I would test them against your own workload before believing them.
Z.ai (GLM). GLM-5.2 for long-horizon coding, million-token context, MIT license. The license is the differentiator for anyone planning to fine-tune.
MiniMax. Low-cost throughput with a credible multimodal family. Also behind Hailuo on the video side.
Cohere. The enterprise and regulated-deployment answer, with Cohere North for on-prem. Less consumer visibility than anyone else here, and more willingness to run inside your building.
Sakana. The genuinely different architecture bet, with the TRINITY orchestration engine. Included because a market this concentrated needs somebody trying something else.
Poolside (code). A code and developer model specialist, not a general frontier lab. Laguna S 2.1 shipped in July: a 118B mixture-of-experts with open weights and a million-token context.
A caveat that belongs in writing
As of their 2026 releases, neither DeepSeek nor Kimi K3 has published a third-party security audit of weights or inference stack. If you self-host Chinese open weights, verify checksums against official releases and run new deployments isolated first. That is not a geopolitical opinion, it is basic supply chain hygiene, and it applies to any unaudited weights from anywhere.
Layer 6: Model Dev, Evals and Observability
How you find out whether the thing works, and how you find out when it stops.
The name changed for a reason. In 2023 this layer was about training and fine-tuning. In 2026 it is overwhelmingly about evaluation and tracing, because almost nobody in the mid-market trains anything, and everybody deploys agents they cannot fully explain.
Agent failures do not look like software failures. They show up in multi-step causal chains: a tool call that silently returned nothing, a context window that truncated the important part, a loop that never terminated. Standard APM cannot see any of it. Running agents in production without trace-level observability is not operating a system, it is hoping.
Category leaders
Hugging Face. The center of gravity for open models and where every open-weight decision starts. Expanded its AMD relationship in July, a small signal that the ROCm ecosystem is finally getting real support. Be precise about what it is, though. Hugging Face is a hub and a distribution ecosystem, not an evaluation tool. It leads this layer because it is the front door, not because it competes with LangSmith.
LangSmith. The deepest agent-engineering platform if you are on LangChain or LangGraph. Node-by-node state diffs, full execution graphs, replay against new model versions. Nothing else debugs a broken agent this well. Proprietary, and self-hosting is enterprise-tier only.
Braintrust. Evaluation and observability in one workflow, with CI/CD gating, which is the correct architecture. If your bottleneck is regression testing instead of debugging, start here. The free tier at 1M spans and 10K eval runs means a small team can start today without a purchase order.
Notable players
Langfuse. The open-source baseline. MIT licensed, self-hostable on Postgres and ClickHouse, framework-agnostic through OpenTelemetry. Now owned by ClickHouse, which is worth knowing, because the independent open option in this category just acquired a corporate parent.
Phoenix. Arize's open-source tracing and evaluation project. The self-hosted entry point into the Arize ecosystem.
Arize AX. The commercial platform, OpenTelemetry-native, with the strongest ML-grade eval primitives. Right answer for organizations running classical ML and LLM workloads side by side.
W&B Weave. Strong eval harness, now inside CoreWeave. Best fit for teams already standardized on Weights and Biases.
MLflow. Still the open standard for experiment tracking and model versioning, and it still runs anywhere.
Databricks Mosaic AI. Fine-tuning and evaluation inside the lakehouse, governed by Unity Catalog. Evaluate it as part of the Databricks stack, not as a point tool.
Galileo. Hallucination and quality detection at scale, plus agent governance work that overlaps into Layer 8.
Promptfoo. Best CI/CD integration in the open eval space. Acquired by OpenAI in March 2026, which changes the neutrality calculus and belongs in your thinking before you standardize on it.
Patronus AI. Evaluation and guardrails with a research bent. Good for teams that want to measure specific failure modes instead of general quality.
DeepEval. The open-source framework from Confident AI. Pytest-style assertions for LLM output, which is the right mental model for engineers.
Fine-tuning
Separated on the graphic because these are training and adaptation tools, not tracing or evaluation, and lumping them in confused the layer.
Axolotl. Configuration-driven fine-tuning for teams that actually fine-tune. Handles complexity you would otherwise write yourself.
Unsloth. Made single-GPU fine-tuning genuinely accessible, and is the reason a mid-market company can consider fine-tuning at all without a research team.
Standards and guidance
OpenTelemetry GenAI semantic conventions. Not a vendor. Instrument to this and you can change observability vendors without rewriting your instrumentation. It is the most useful architectural decision available in this layer and it costs nothing.
Band 3: Runtime control and governance
What agents are allowed to do, and how you prove it afterward. This now drives cost and risk as much as the model layer does.
Layer 7: Orchestration and Protocols
What turns a model that answers questions into a system that does work.
Protocols won this year. MCP was donated to the Linux Foundation's Agentic AI Foundation in December 2025 with OpenAI and Google as co-founders, and has since passed 100 million monthly SDK downloads. The reported effect on integration work is the number that should get your attention: wiring a new SaaS tool into an agent dropped from around 18 hours of custom function-calling code to roughly 4. The two-layer pattern is now the default, with MCP handling vertical tool access and A2A handling horizontal agent-to-agent coordination. A joint specification is expected in Q3.
This is still where the ROI lives. LangChain's State of Agents 2026 puts 57 percent of organizations running agents in production, and other 2026 surveys put it far higher. But the failure modes caught up with the enthusiasm. Most agent incidents trace to tool call failures, context truncation, and runaway loops, not to model errors. The frameworks that win are the ones that make those failures visible and recoverable.
Category leaders
LangGraph. The most production-ready orchestration for stateful workflows: checkpointing, branching, human-in-the-loop, and an enormous battle-tested ecosystem. It is the only framework where I have watched mid-market teams ship genuinely complex agents and still be able to debug them six months later.
OpenAI Agents SDK. Native MCP support, visual workflow design through AgentKit, tracing, and now Promptfoo's eval tooling folded in. If your team is already deep in the OpenAI stack, this is the shortest path from prototype to production.
Claude Agent SDK. The harness that powers Claude Code, exposed as a library in Python and TypeScript. Renamed from the Claude Code SDK in September 2025. Tool-use loops with sandboxed execution and permissioning built in from the start instead of bolted on afterward. Given what happened to OpenClaw this year, safety-first architecture stopped being a nice-to-have and turned into a procurement requirement.
Notable players
Agent Development Kit (ADK). Google's framework, built around A2A. Strongest if your data estate is on GCP.
Microsoft Agent Framework. Agent Framework 1.0 launched December 2025 under an MIT license, unifying Semantic Kernel and AutoGen into a single Python and .NET SDK. Note that it is distinct from Microsoft Foundry, which is the platform. Vendors conflate the two constantly. Do not.
Temporal. The underrated one. Durable execution for long-running workflows, which is precisely what breaks when an agent runs for six hours and the process dies at hour five. Not marketed as AI infrastructure. Is AI infrastructure.
n8n. The default for mid-market teams that need agent workflows without a platform engineering function. Self-hostable, though under a fair-code license and not an OSI one.
Dify. Open-source LLM app platform with a visual builder. A good middle ground between n8n's automation focus and a code framework.
CrewAI. Good for prototyping multi-agent patterns, token-hungry in production at roughly 3x the tokens of an equivalent LangGraph implementation. Know that going in.
LlamaIndex. Retrieval-focused framework, still the strongest option when your agent's core problem is finding the right document instead of orchestrating the right sequence.
Pydantic AI. Type-safe agent framework, gaining ground fast with teams who want their agent code to look like the rest of their Python.
NVIDIA NeMo Agent Toolkit. Enterprise agent components from chip to orchestration, running on any silicon.
OpenClaw. Enormous adoption at roughly 368,000 GitHub stars and 12 million downloads, with NVIDIA and Tencent contributing engineering. Also the most instructive security story of the year, which I have given its own section below.
Amazon Bedrock AgentCore. Launched October 2025, with policy controls reaching GA in March 2026. Memory, gateway, identity, and observability in one framework-agnostic runtime. The SDK passed 2 million downloads within five months of preview.
Foundry Agent Service. GA since April 2026 with more than 10,000 reported customers. The differentiator is Entra Agent ID, which gives every agent deployed through it a first-class identity in the same fabric that governs human employees.
Vapi (voice). Voice agent infrastructure handling telephony, turn-taking, and interruption. Not a finished product. It is something you build on, which is why it sits here and not in the customer experience column.
Standards and guidance
Model Context Protocol (MCP). Under the Linux Foundation's Agentic AI Foundation. Streamable HTTP transport, OAuth 2.1 with Resource Indicators. Treat it as infrastructure, not as a vendor choice.
Agent-to-Agent (A2A). Google's contribution, now the de facto horizontal coordination protocol, with Agent Cards for capability discovery. Adoption outside Google is real but early.
The OpenClaw problem
This gets its own section because it shows the gap between adoption and readiness more clearly than anything else on the chart.
Five formal CVEs landed between January and June 2026, alongside 137 tracked security advisories. CVE-2026-25253 was a CVSS 8.8 one-click remote code execution via WebSocket hijacking, exploitable even against instances bound to localhost, because the victim's own browser initiates the outbound connection. Worse followed. CVE-2026-32922 rates CVSS 9.9 and lets any legitimately paired device escalate to full admin access through a token rotation race condition.
Then the supply chain. A Koi Security audit of all 2,857 skills on ClawHub found 341 malicious entries, 335 of them traceable to a single coordinated operation. By March the count had passed 1,184 malicious packages. Separately, researchers found the Moltbook database publicly accessible without authentication, exposing roughly 35,000 email addresses and 1.5 million agent API tokens in plaintext.
SecurityScorecard counted more than 135,000 publicly exposed instances across 82 countries, roughly 63 percent of them running with no authentication at all. Belgium's national cybersecurity center issued a government advisory.
I still run agents on it. I would not deploy it raw inside a client environment and neither should you. It is a superb personal and prototyping platform with a privilege model designed for a single trusted operator, currently being adopted by organizations that need something else entirely. One of its own maintainers said it plainly in the project Discord: if you cannot understand how to run a command line, this is far too dangerous a project for you to use safely.
If you deploy it, deploy it patched, authenticated, origin-validated, and by someone who has read the issue tracker.
Layer 8: Security, Safety and Governance
In March this was a policy layer with a few startups attached. It is now a consolidating security market with roughly $74B of M&A behind it, and every major cybersecurity incumbent holding a position.
Two things changed. Agents started taking actions, which means a compromised agent is no longer a data leak but an unauthorized transaction. And the regulatory picture got more complicated instead of stricter, which is worse for planning purposes.
Category leaders
Palo Alto Prisma AIRS. Built on the Protect AI acquisition. The most complete AI security posture story available inside a platform enterprises already buy. If you already have Palo Alto, this is a line item instead of a new vendor relationship, and in mid-market procurement that matters more than feature parity.
Check Point AI Agent Security. Built on the Lakera acquisition, which brought Lakera Guard, Lakera Red, and the Gandalf adversarial dataset into Check Point's platform. Runtime protection covering LLM inputs, outputs, RAG data, and MCP servers, with support for more than 100 languages. Lakera Guard was the best managed low-latency guardrail API in the independent market, and it now ships with enterprise distribution behind it.
Microsoft Entra Agent ID. Agent identity in the same fabric that governs human employees. Agent identity is the unsolved problem underneath most agent security incidents, and Microsoft is furthest along at solving it for customers who already live in its ecosystem.
Notable players
Noma Security. Agent governance and posture management, $100M raised. One of the few independents with genuine enterprise proof points.
HiddenLayer. ML model security and adversarial detection. Still independent, still credible, and one of the few that treats the model itself as the attack surface.
CrowdStrike. AI detection and response with agentic telemetry across endpoints, built on the Pangea acquisition.
Cisco AI Defense. Built on the Robust Intelligence acquisition. Model validation and runtime protection inside Cisco's security platform.
SentinelOne. Inline protection for enterprise generative AI through the Prompt Security acquisition, covering prompt injection defense and outbound data leakage.
F5 AI Guardrails. Formerly CalypsoAI. On-prem capable, which matters for organizations that will not route inference through a vendor cloud. It is also the one entry on the graphic whose exact current SKU name I could not verify, which is why it carries a flag.
Oasis Security. Non-human identity governance, $120M Series B around RSAC 2026. The vertical I would watch hardest, because this is the actual root cause under most agent incidents.
Astrix Security. Identity and SaaS-to-SaaS governance extended to agentic access. Adjacent to Oasis and worth evaluating alongside it.
XBOW. Autonomous offensive security, $120M Series C at a billion-plus valuation, 100-plus customers including Moderna. Red teaming that runs continuously instead of once a year.
WitnessAI. $58M in January 2026, timed to enterprise anxiety about non-human decision-making. Better funded than most independents in the governance slice.
Mindgard. AI red teaming and security testing with a research foundation. Good for organizations that want to test their own models instead of buying a filter.
Standards and guidance
These are not products and should not be evaluated as if they were. They are what your auditor and your enterprise customers will benchmark you against.
OWASP Top 10 for Agentic Applications 2026. Built by more than 100 contributors, and now the de facto threat taxonomy. The best free starting point for a mid-market agent threat review, and better than any vendor framework.
NIST AI Risk Management Framework. What most US enterprise governance programs anchor to. When a customer asks how you govern AI, this is the vocabulary they expect back.
ISO/IEC 42001. The certifiable AI management system standard. Larger customers will start asking for it in RFPs, and it takes long enough to obtain that you want to know about it early.
EU AI Act / EU AI Office. The regulation and its supervisory body, not a framework you adopt. Timeline below.
Anthropic Responsible Scaling Policy. Sets norms the rest of the industry gets measured against, whether or not you use Claude. It belongs here as guidance, not in the vendor row.
The regulatory clock, precisely
Parliament endorsed the Digital Omnibus on AI on June 16, 2026 by 423 to 57 with 174 abstentions. The Council gave final approval on June 29.
Stand-alone high-risk obligations under Annex III, covering employment, credit, education, and essential services, moved from August 2, 2026 to December 2, 2027. High-risk AI embedded in regulated products under Annex I moved to August 2, 2028.
August 2, 2026 is not cancelled. Most Article 50 transparency obligations still apply that day: chatbot disclosure, provider-side content marking, deepfake labeling, and notices for emotion recognition and biometric categorization. Penalties run to 15 million euros or 3 percent of worldwide annual turnover.
There is one carve-out inside the carve-out. Article 50(2), the watermarking and machine-readable marking requirement for synthetic content, gets a grace period to December 2, 2026, but only for systems already on the market as of August 2, 2026. Anything you ship after that date complies on day one. The Omnibus also adds a new Article 5 prohibition covering AI-generated non-consensual intimate imagery and CSAM, effective December 2, 2026.
Read practically: you got sixteen extra months on the hard part and zero extra months on the disclosure part. Spend the runway on inventory and classification, because that is the work that does not get easier with time.
Internal governance
Nothing changed here except urgency. AI usage policy, data classification, model approval workflow, agent inventory, audit cadence. The agent inventory is the new one and the one everybody skips. You cannot govern what you cannot enumerate, and by the time a mid-market company notices, there are usually forty agents running that nobody approved.
Band 4: Business applications
Where value accrues, and where your buying should start.
Layer 9: Applications and Agents
Where the money gets made. Six verticals plus a sidecar.
Every layer below this one is a cost. This is the only layer that produces revenue or removes work. Find the application that solves your most painful problem, then work backward down the stack only as far as you have to.
Vertical A: Agent planes (NEW)
The console where an organization governs its agent fleet. This category did not exist in March and now has five serious entrants. It is a different product from the agents themselves, which is why it gets its own column.
OpenAI Frontier. Leader. The only entrant designed to govern agents from multiple vendors in one place: audit logs across all agents, unified permission controls, unified cost visibility, uniform policy enforcement. If you run agents from six vendors, that is the pitch. If you run one vendor's agents, it is table stakes.
Microsoft Agent 365. Leader. Launched November 2025 as the control plane governing the agent estate, including agents built on Salesforce, ServiceNow, Google, and open-source frameworks. Wins by default in Microsoft shops, and the cross-vendor governance claim is more credible than it sounds, given Entra underneath it.
Agentforce 360. Leader. The broadest production footprint of any platform here, with OpenAI and Anthropic models embedded and Claude specifically positioned for financial services, healthcare, and cybersecurity. Pricing runs $125 per user per month for add-ons and $550 for Agentforce 1 editions. The lock-in sits at the orchestration layer. You can swap the model. You cannot swap Atlas.
Gemini Enterprise. Google Cloud's unified agent portfolio, positioned as the successor to Vertex AI, spanning app, platform, and CX with identity, registry, gateway, simulation, observability, and memory.
Glean Agents. The vendor-neutral option, built on a permission-aware knowledge graph instead of a single vendor's identity fabric. More than 100 million agent actions annually.
Vertical B: Productivity
Microsoft 365 Copilot. Leader. AI inside the tools your company already pays for. The open question in this vertical is whether Microsoft bundles the whole category to death, and Copilot is the reason to keep asking it.
Glean. Leader. ARR crossed $300M in late May, up from $200M in December and roughly $100M in early 2025, at a $7.2B valuation. The pivot is the signal: permission-aware enterprise search became an agent platform, and the context graph now sells as a cost-reduction story instead of a capability story. Its win rate against bundled Microsoft is the clearest indicator in enterprise AI right now.
Gemini for Google Workspace. The Google-native equivalent. The correct name matters here, because "Google AI" is not a product you can buy.
Notion AI. Now with agentic capability and enterprise search. The right answer for companies whose knowledge already lives in Notion.
Claude Cowork. Anthropic's agentic knowledge-work product for delegating multi-step tasks. The non-developer counterpart to Claude Code.
Writer. Enterprise agents and AI Studio with governance built in. Strong in regulated content workflows.
Moveworks. AI assistant platform and agent marketplace focused on internal employee support. The IT helpdesk case, done well.
Vertical C: Code and development
The vertical with the biggest structural change this year.
Claude Code. Leader. Run-rate revenue past $2.5B by February and reportedly around $8B by May, with Menlo Ventures putting it at 54 percent of the enterprise AI coding market. Roughly 4 percent of all public GitHub commits, with SemiAnalysis projecting north of 20 percent by year end. It also took 46 percent most-loved in JetBrains' April survey against Cursor at 19 percent and Copilot at 9 percent. Clearest revenue leader in agentic coding, and the best tool available for hard multi-file work.
Cursor. Leader, with an asterisk. Best full-IDE experience: Composer 2, multi-repo workspaces, parallel local and cloud agents. Roughly $4B annualized by early June, more than doubling from $2B in February, with about $2.6B from enterprise and more than 60 percent of the Fortune 500 using it.
The asterisk is ownership. SpaceX exercised its option on June 16 and signed a $60B all-stock merger agreement for parent company Anysphere, expected to close in Q3 pending regulatory approval, at which point Cursor becomes a wholly owned SpaceX subsidiary. Meanwhile, per Ramp data, its share fell from roughly 41 percent in June 2025 to about 26 percent by May 2026 while Anthropic's climbed toward 50. Growing fast and losing share at the same time is what a margin trap looks like from the outside, and it explains the sale. For regulated buyers, the ownership question belongs in procurement now, not in Q4.
GitHub Copilot. Leader. About 40 percent adoption at companies over 5,000 employees, 4.7 million paid subscribers, usage across 90 percent of the Fortune 100, and IP indemnity. It also now runs Claude models inside its multi-model interface, which is Microsoft conceding the quality argument while keeping the distribution.
OpenAI Codex. A coding agent across cloud, CLI, and ChatGPT. The natural pick if your organization standardized on OpenAI.
Cognition Devin. Acquired Windsurf and more than doubled ARR. Autonomous instead of assistive, which puts it in a different product category from the three leaders.
Replit. Replit Agent for building and shipping apps from a browser. Strong for prototyping and for non-engineers building internal tools.
Lovable. AI app and site builder, closer to vibe-coding than to IDE tooling. Included because that category is now large enough to matter.
Base44. AI app, site, and agent builder, acquired by Wix in 2025 with the product preserved.
One number should shape your tooling policy more than any of the above. In-house teams run a median of 3.1 AI coding tools per developer, with roughly 70 percent of engineers using two to four at once. This is not a winner-take-all market. The dominant pattern is Cursor for inline editing plus Claude Code for hard multi-file work, and fighting that with a single-vendor mandate mostly produces shadow spend.
Vertical D: Customer experience
Sierra. Leader. Raised $950M at a reported $15.8B valuation in May, hitting roughly $200M ARR, up from about $130M at the end of 2025. Outcome-based pricing, charged per resolution instead of per seat. That pricing model is the alignment that justifies the multiple, and the multiple is roughly 100x ARR, so say that part out loud when you evaluate it.
Decagon. Leader. Customer AI agent platform with omnichannel deployment. The most credible direct competitor to Sierra on enterprise deals.
Intercom Fin. Documented as Intercom's AI agent, with its role expanding across service, sales, and ecommerce. The easiest path if you already run Intercom.
Zendesk AI. Self-improving AI agents for service operations. Same logic: incumbent advantage in accounts that already have the suite.
Cognigy. Full-stack AI agents for CX and contact centers, strong in European enterprise.
Kore.ai. Enterprise agent platform with deep CX roots that now spans employee and workflow use cases. Worth evaluating beyond the customer service box.
Parloa. AI Agent Management Platform for enterprise CX, with a genuine voice focus.
The growth vector is voice. Roughly 80 percent of customer service interactions still happen by phone, and almost none of them are automated.
Vertical E: Vertical AI
Harvey. Leader. Legal AI with broad law firm and professional services deployment, trading near 58x ARR. The proof case for this entire vertical thesis.
Abridge. Leader. Clinical conversation and documentation with major health system deployments. It solves a problem clinicians actively hate, which is the shortest path to adoption in healthcare.
OpenEvidence. Leader. Clinician-focused medical search and clinical decision support, $250M Series D in January at a $12B valuation on roughly $100M annualized. The multiple is extraordinary and the institutional partnerships are real.
Legora. Collaborative AI for lawyers. The main European challenger to Harvey.
Ambience Healthcare. Clinical documentation and coding, $243M Series C. Coding is the part that touches revenue, which is why it matters.
Hippocratic AI. Healthcare agents that explicitly limit themselves away from diagnosis and prescribing. The scoping is the product decision worth studying.
Rogo. Finance-specific AI for bankers and investors. Domain vocabulary as a moat.
Hebbia. Enterprise research and document workspace, strongest in finance and legal diligence work.
If you take one recommendation from this piece, take this one. I expect the most mid-market growth here over the next twelve to eighteen months. General-purpose AI is commoditizing on schedule. Domain depth is not. Harvey at 58x, Sierra at 100x, and OpenEvidence at $12B look unhinged next to Glean at 24x, but what is being bought is a workflow with regulatory teeth and switching costs attached, not a model wrapper.
Vertical F: Content and media
Kling 3.0. Leader. Leads the text-to-video arena at 1934 Elo. Kling 3.0 Omni does per-character lip sync with two speakers on separate audio tracks, which nothing could do before this year.
Veo 3.1. Leader. Google DeepMind's video model, still the only one owning 48kHz speech generation, and the enterprise-safe pick given Google's indemnification posture.
ElevenLabs. Leader. $11B valuation on a $500M Series D in February, with more than $330M ARR at the end of 2025. The clear leader in voice.
Seedance 2.5. Demoted from leader this edition, deliberately. ByteDance's model is technically excellent, and its July launch brought 30-second native generation without stitching. But Seedance 2.0 drew a Motion Picture Association cease-and-desist describing its infringement as a feature and not a bug, plus individual letters from Disney, Warner Bros., Paramount, Sony, and Netflix, a SAG-AFTRA condemnation, and a Senate demand that it be shut down. ByteDance paused the global rollout in March, and by March 30 had added C2PA provenance watermarking and filters blocking recognizable faces and copyrighted characters. None of the disputes have been resolved in court, so enterprise teams building on it inherit that exposure. Recommending an unindemnified model under active litigation to a mid-market buyer is bad advice no matter how good the output looks.
Runway. Gen-4.5 was number one at launch in late 2025 and has since dropped out of the arena top 10. It still has the best control surface in the category, with motion brush and reference-image consistency, and the deepest film production ecosystem. Rank it on controllability, not on arena score.
Midjourney. Still leads on image aesthetic, and still the answer when the output has to look intentional instead of generated.
FLUX. Black Forest Labs. Leads open-weight photorealism, and the practical choice for self-hosted image generation.
Ideogram. Still owns text rendering inside images, which sounds narrow until you need a legible sign or logo.
Suno. Major AI music product with Studio, paid plans, and real creator workflows.
MiniMax Hailuo. Video generation and agents, competitive on cost.
Alibaba Wan. Open-source video generation, and the self-hostable option in a category where almost everything is an API.
What has not changed: content production costs dropped 80 percent or more, and the models now generate synchronized audio in a single pass. You still need a human directing it and catching the failures.
Sidecar: Systems of record
Already in your contract. The question is how much of their AI you turn on.
These sit in their own strip on the graphic instead of forming a seventh column, because they are not peers to the categories above. They are the substrate agents run inside. The buying motion is completely different, too: you are not selecting a vendor, you are deciding how much of your existing vendor's AI to switch on, and that decision usually gets made by whoever handles the renewal rather than by anyone evaluating AI.
Salesforce. Agentforce embedded in CRM. The largest installed base for agentic AI by default.
ServiceNow. Workflow automation with agents across IT, HR, and customer service. Strongest where process is already codified.
SAP. Business AI across ERP. Slow, deep, and unavoidable if you run SAP.
Workday. HR and finance agents inside the system that already holds your employee data, which makes the permission story simpler than any bolt-on.
HubSpot. The mid-market CRM answer, and often the most realistic starting point for companies in my target range.
Atlassian. Rovo agents across Jira and Confluence, where your engineering context already lives.
Intuit. AI across QuickBooks and adjacent finance products. The small-business substrate.
How executives should read this
Buy Layers 1 through 5. Build your differentiation at Layers 7 through 9. Nobody wins by owning infrastructure.
Standardize on MCP and A2A at the protocol layer and route models at runtime. Four frontier models shipped in three weeks in July, and that pace is not slowing.
Layer 8 is not optional. The unsolved problem there is agent identity: systems built for humans, handing credentials to non-humans.
Inference spend now exceeds training spend, which means Layer 3 decides your gross margin and Layer 5 does not.
Nearly every enterprise now runs agents in production. Far fewer can govern them. The constraint stopped being capability and became control.
Value keeps moving up the stack, and vertical applications hold their multiples better than horizontal ones.
Deals are quoted in gigawatts now, not GPU counts. Google says it cannot build fast enough and is buying third-party capacity.
EU AI Act transparency rules apply August 2, 2026. High-risk rules slipped to December 2027. Start the inventory anyway.
Ten predictions
Each of these is falsifiable, which is the only kind worth publishing.
1. Inference gets commoditized faster than these valuations expect. Fireworks at $17.5B, Together at $8.3B, Baseten at $13B, all with gross margins around 50 percent because GPU cost sits in COGS. Meta just started selling excess capacity and every hyperscaler wants the same revenue. The revenue is real. The margin structure is not software. I expect at least one down round or consolidation in this tier within twelve months.
2. Agent identity becomes the security category of 2027. Every serious agent incident this year traced to the same root cause: systems designed for human actors, handing credentials to non-human ones. Oasis, Astrix, and Entra Agent ID are early. Expect a major platform acquisition by mid-2027, triggered by a publicly disclosed breach where an agent took an unauthorized action with real money attached.
3. Model names stop mattering to mid-market buyers. Four frontier models in three weeks is not sustainable for anyone building on top. By mid-2027 the default enterprise architecture is a gateway with a routing policy, and asking which model a company uses becomes as strange as asking which CDN edge served a page.
4. Open weights take the majority of enterprise token volume while closed frontier keeps the majority of revenue. The economics only point one direction for high-volume, low-ambiguity work. But the hard, high-stakes, liability-bearing 10 percent stays on the frontier, and that is where the money is.
5. The EU deferral backfires on companies that treat it as an excuse. The hard part was never the documentation template. It was enumerating every AI system in the organization and classifying it. That work does not get easier with time, and the systems keep multiplying. Companies starting their inventory in late 2027 will have weeks, not months.
6. At least one more top-five AI coding tool gets acquired by a non-software company. SpaceX buying Cursor established the template: buy the developer distribution channel, feed it into your own model. Every large company with frontier model ambition and no developer relationship now has a playbook and a comp.
7. The agent management plane consolidates to two or three winners, none of them a startup. Governance planes win on identity integration and audit, which are exactly what incumbents already own. Glean is the only independent with a real shot, and its risk is Microsoft bundling.
8. Power, not silicon, becomes the number in every AI infrastructure headline by Q2 2027. Google already says it cannot build fast enough. CoreWeave and Nebius each have gigawatts contracted with most of it not yet online. The unit of measure has already changed and the headlines will catch up.
9. A high-profile agent incident forces the first real agent liability case. OWASP published the taxonomy, the CVEs are landing, most organizations have agents in production, and almost nobody has an agent inventory. I do not think this one requires much imagination.
10. Vertical AI stays expensive and stays worth it. These multiples look unhinged next to horizontal software. But what is being bought is a workflow with regulatory teeth and switching costs. I expect the vertical names to survive a general AI multiple compression better than the horizontal ones do.
Methodology and limits
I ranked by news cycle share, public valuation, and enterprise footprint. Those are proxies for relevance, not for quality. It is the honest version of what every market map does implicitly and rarely admits.
Figures mix audited filings, company statements, and private-company reporting. The public company numbers are reliable. Private valuations moved 2x to 4x inside single quarters this year and should be treated as directional. Verify anything before it goes in a board deck.
Twelve entries carry a name-confidence flag on the graphic, meaning the printed label is not confirmed as the exact current public product name. Product naming in this market changes faster than anyone reprints a poster. F5 AI Guardrails is the one I am least sure of.
Some deliberate omissions. OPEA, because it is a 2024 sandbox-tier project with no demonstrated 2026 enterprise traction, and putting it beside MCP would be false equivalence. Sora, because OpenAI discontinued the app on April 26 and the API goes on September 24. And a long tail of good tools nobody wrote about this quarter.
Last updated July 2026. This is a living document and parts of it will be wrong by October.
Chris Labatt-Simon / FairWinds Strategies / labattsimon.com
Frequently Asked Questions
A "category leader" here means the vendor a mid-market company should start evaluating first if they have limited time to research a layer. It's not a quality judgment, and it's not based on market share. Three of the picks are open-source projects with no sales team, which makes clear the label is about where to begin, not who won.
Fireworks AI, Together AI, and Baseten lead the inference and model serving layer. This layer determines gross margin for mid-market companies because the cost of tokens sits directly in COGS. Pick the wrong vendor and you pay more for inference than you charge for the product. Pick right and open-weight models handle 80 percent of traffic at a fraction of frontier pricing, with hard requests escalated only when necessary. Together reports customers cutting inference costs by up to sixty-fold versus closed models.
OpenClaw had five formal CVEs filed between January and June 2026, including a CVSS 8.8 remote code execution flaw via WebSocket hijacking and a CVSS 9.9 privilege escalation bug. A third-party audit of its skill marketplace found over 1,000 malicious packages, and a database breach exposed 35,000 email addresses and 1.5 million agent API tokens in plaintext, with more than 135,000 publicly exposed instances running no authentication at all. Before deploying it, organizations should ensure it is fully patched, running with authentication enabled, origin-validated, and set up by someone who has read the project's issue tracker — it is a strong prototyping tool, but its privilege model was designed for a single trusted operator, not a multi-user enterprise environment.
The EU AI Act's Article 50 transparency obligations still take effect on August 2, 2026, covering chatbot disclosure, provider-side content marking, deepfake labeling, and notices for emotion recognition and biometric categorization, with penalties up to 15 million euros or 3 percent of worldwide annual turnover. The Digital Omnibus pushed stand-alone high-risk obligations under Annex III (employment, credit, education, essential services) to December 2, 2027, and high-risk AI embedded in regulated products under Annex I to August 2, 2028. There is one narrow exception: the synthetic content watermarking requirement under Article 50(2) gets a grace period to December 2, 2026, but only for systems already on the market before August 2, 2026.
Vertical AI applications like Harvey in legal or Abridge in healthcare come with regulatory requirements, deep workflow integration, and high switching costs built in — making their value harder to replicate and their pricing power more durable. General-purpose AI is commoditizing quickly, but domain depth is not. That's why vertical multiples like Harvey at 58x ARR and OpenEvidence at a $12B valuation hold up against horizontal tools trading at far lower multiples.