Snowflake adds AI routing to curb enterprise spend
Wed, 19th Aug 2026 (Today)
Snowflake has added dynamic model routing to Cortex AI Gateway to help customers manage AI spending across enterprise workloads.
It is also adding the open models DeepSeek-V4-Flash and GLM-5.3 to Snowflake Cortex AI, expanding the range of models available to customers seeking to balance cost and output quality while keeping data inside Snowflake's governed environment.
The updates address a problem many large companies face as they expand generative AI tools and agents across departments. Using the same advanced model for every task can quickly drive up costs, while selecting and maintaining a growing list of models adds operational work for developers and platform teams.
Dynamic model routing is intended to automate that choice. Cortex AI Gateway can now select a model for each request based on factors such as quality, speed, customer preferences and cost. It can send simpler or repetitive tasks to lower-cost models while reserving more advanced models for work that requires deeper reasoning.
The feature is built into Snowflake's own AI products, including Snowflake CoCo and Snowflake CoWork, and is also available to third-party AI agents that use Cortex AI Gateway. Customers can control which models and providers are available to users, which may matter for multinational companies facing regional restrictions or regulated industries with strict compliance requirements.
Cost controls
Alongside routing, Snowflake is expanding administrative controls to give companies more visibility into AI use. Cortex AI Gateway lets administrators track token usage and costs, set spending limits across applications and agents, and manage consumption across teams.
Those controls also extend into Snowflake CoCo through the company's existing role-based access and tagging framework. Administrators can set default models, attribute usage to teams or cost centres, establish per-user quotas and receive alerts as consumption approaches defined limits.
Snowflake is framing the latest updates around what it calls "intelligence efficiency", or how effectively businesses convert computing, models, data and context into commercial results. While the term is Snowflake's, the underlying issue has become more prominent as enterprises move from AI trials to broader deployment and begin measuring whether usage justifies the cost.
Snowflake said internal testing showed that mixing open and proprietary models for different tasks could maintain similar quality while improving token efficiency. In one evaluation, agents using dynamic model routing to build a dbt pipeline achieved up to three times greater token efficiency than an approach that relied only on frontier models, while maintaining the same quality. In another, engineering teams completed the same number of pull requests with 25 per cent greater token efficiency, according to the company.
Model expansion
The addition of DeepSeek-V4-Flash and GLM-5.3 expands a model catalogue that already includes options from providers such as Anthropic, OpenAI, Google and Mistral. The broader selection is intended to help customers choose combinations of models that fit different workloads without rebuilding applications whenever pricing or model performance changes.
Snowflake's AI Research Team also shared internal benchmark results for the new open models on enterprise-focused tasks. DeepSeek-V4-Flash scored 74.4 per cent on data engineering tasks in recent testing and outperformed the leading proprietary models used in that evaluation, Snowflake said. It added that GLM-5.2 scored 62.8 per cent while using fewer tokens than any other model tested.
Because the new models are open source and can be self-hosted, organisations can keep data within their own environments, according to Snowflake. That may appeal to customers in sectors with strict data-handling rules and growing interest in open models amid concerns about cost and vendor dependence.
Snowflake has been building Cortex AI Gateway as a central layer for governance, routing and oversight of AI use. The latest additions suggest the company sees model choice and cost management as part of the infrastructure challenge for enterprise AI, rather than a one-off purchasing decision.
"Enterprises are becoming much more rigorous about the economics of AI. The question is no longer how much AI they are using, but whether that AI is translating into meaningful business value," said Sridhar Ramaswamy, Chief Executive Officer, Snowflake.
"Achieving intelligence efficiency requires the flexibility to use the best model for each task as the landscape evolves. Snowflake's role is to absorb that complexity so customers can focus on outcomes while we optimise model choice underneath," said Ramaswamy.
Industry analyst Sanjeev Mohan said the challenge is increasingly operational rather than theoretical as the number of models grows.
"Enterprises are drowning in model choices, but the real problem isn't which model to pick. It's the operational overhead of picking the right one for every task, at scale. Snowflake's dynamic model routing directly addresses that gap," said Sanjeev Mohan, Principal and Founder, SanjMo.
"By automating intelligent model selection within Cortex AI Gateway, Snowflake is removing a real friction point that has been slowing enterprise AI deployment. The ability to match workload complexity to model cost, without rebuilding your infrastructure every time a new model drops, is exactly the kind of efficiency enterprises need to move from AI experimentation to AI at scale," Mohan said.