Talk with our engineers by joining the new Onehouse community Slack!

TL;DR
Onehouse AI Gateway combines inference routing and MCP tool access without chaining you to the models your data platform picks. Connect agents to your open tables and run the gateway in your VPC, on-prem or on a neocloud. Any model, any table, anywhere, with no added inference tolls.
Future of Data Access is now clearly “Agentic”
AI models have become so proficient at generating code and SQL that natural language queries are quickly becoming the “de-facto” way to ask questions of data. Ad hoc analysis for day-to-day questions increasingly runs through AI coding agents like Claude Code and Codex, and internal “ChatGPT”-like enterprise applications. Such agentic access democratizes enterprise data at machine speeds, letting non-technical users interact with it directly without learning SQL. These lower barriers to entry are predicted to drive 10x more data access to warehouses and data lakes over the next few years. Gartner expects 40% of enterprise apps to ship task-specific AI agents in 2026, up from under 5% in 2025, and inference cost per agentic workflow to grow more than fivefold through 2028. An AI agent application brings two systems together. A model (typically an LLM) processes a natural language request and decides what to do next via natural language output; inference is the service running that model. It’s text in and text out, with the response instructing the caller to call “tools” before the next turn. A tool retrieves data or executes an operation. For a data agent, the tool might inspect a schema or execute SQL against a lakehouse.
The AI Gateway: the newest addition to the core data platform
But agentic access patterns are markedly different from the static dashboards of the traditional BI era. Most importantly, agents “reason” and issue queries of different shapes and sizes; the same natural language input can produce entirely different agent execution and sequences of queries and tool calls. Since the model knows nothing about your tables (their schema, what the columns mean, or what’s inside them), it can take many attempts (“turns”) to get the queries right. This slows down final answers and hammers the lakehouse with a barrage of queries in quick succession.
Earlier this year, we announced Lakegres, our interactive SQL query engine built to keep up with these agentic access patterns and velocity. But that alone is not sufficient. We need a standardized component that streamlines this flow of information across models, tables and data catalogs, while providing essential functionality to collect traces, control tool calls, and switch models dynamically based on cost/availability. The challenge grows as more teams bring their own agents, models and tools: each agent needs to act with the right permissions, keep sensitive data from reaching unapproved models, and stay within budget even when it retries or takes an unexpected path. Implementing these controls separately in every application quickly becomes unmaintainable.
This has prompted a brand new component in front of your data lake or data warehouse engines: the AI Gateway. It provides a common place to enforce these controls, while tracking who initiated each request, what the agent accessed and how much the work cost. Onehouse AI Gateway combines an inference gateway for model requests with an MCP gateway for tool calls, bringing both under shared access controls and observability. In this blog, we introduce the Onehouse AI Gateway and how it stands out against other AI Gateways for data lakehouses and warehouses, upholding a more open approach to the data platform, much like the rest of Onehouse’s services.
The AI Gateway should not chain you to an inference provider
Optionality and composability guide our platform design, so one principle was easy: don’t couple the model (inference) with the tool call (data access). Given the pace of model improvements and the surge in open-weight models, relying on your data vendor to deploy the latest and greatest models should be your last resort.
Snowflake releases its Cortex AI Gateway in July 2026. Given how Snowflake has preached plenty to the community about true openness and the right open table format through its Apache Iceberg™ efforts, we were shocked to find its AI stack only works when both your models and your data live inside Snowflake. Then we remembered Snowflake still defaults to its closed storage format, and it felt less surprising. The Cortex AI Gateway routes exclusively to Snowflake-hosted models: served from Snowflake’s own capacity, authenticated with a Snowflake token. No bring-your-own API key, no way to reach a model you host or pay for elsewhere; the only gateway that exists is the built-in SNOWFLAKE object.
Snowflake lets you choose among models, but controls the catalog, regional rollout and production availability. Inference is metered in tokens but priced in Snowflake AI Credits per million tokens, adding Snowflake’s own credit currency to the bill.
Source: Snowflake documentation, captured September 29, 2026. Selected rows and columns from the cross-region table. Legend: * = public preview; ** = private preview requiring explicit account enablement. Snowflake says both preview categories are unsuitable for production workloads.
Same story on the data side: Snowflake’s VECTOR data type explicitly excludes Apache Iceberg tables. The open format Snowflake spent years evangelizing gets no first-class vector storage; the supported path is Cortex Search, which reads your Iceberg table and builds embeddings into its own managed index.
The pattern completes itself: the moment you touch AI, models run on Snowflake compute and embeddings land in their native structures. AI workloads are becoming the mechanism that pulls open data back into the walled garden. An AI gateway should govern access (identity, budgets, audit) and stay neutral on where models run and where vectors live. Coupling all three is lock-in with an OpenAI-compatible API in front of it.
The AI Gateway should not collect tolls and taxes
Databricks avoids these obviously wrong decisions: Databricks Unity Gateway supports both Databricks-hosted and external provider models (e.g. OpenAI, Anthropic), adds no extra charge on inference requests routed to external models, and even throws in free tracking and cost analysis of inference tokens at the external provider’s list price, though it does limit features like smart routing to Databricks-hosted models. However, our team has learnt over the years that Databricks often hides the tolls in plain sight: pushing Unity Catalog as a requirement for Apache Iceberg on Databricks, or running Delta Lake with Uniform instead of native Iceberg for two-plus years after telling everyone else to converge. So we naturally looked closer.
Databricks Unity Gateway charges $0.10/GB to log inference requests through Inference Tables, and another $0.10/GB with Usage Tracking enabled. For comparison, AWS S3 Tables charges $0.005/GB for Iceberg maintenance, presumably how Databricks stores and manages this data (or is it Delta Lake?). Agent traces already run several TBs/day uncompressed at many enterprises, and at multiple-agents-per-employee maturity this data becomes a significant chunk of your cloud storage. A 100 MB/s logging stream produces 262,800 GB a month (100 MB/s × 3,600 s × 730 hrs). That’s $26,280/month. Even at 10 MB/s, it’s $2,628/month ($31,536/year) and possibly $5,256/month with usage tracking. These are volumes a single r8g.xlarge in your own AWS account can handle for $172.02/month, at on-demand rates.
The obvious rebuttal would be: Inference Tables and Usage Tracking are optional. But turning them off means running blind. Logging every request and response is how you evaluate model quality, debug agent behavior, attribute spend to teams, and satisfy an auditor. So “optional” is doing a lot of heavy lifting here for Databricks. Most production deployments will end keeping it on since this data is one of the most valuable datasets for post-training models.
The AI Gateway should run wherever data is
Spoiler: Data is everywhere. A vast amount of data still sits in datacenters, and many enterprises are leaning back on their existing datacenter contracts, given the GPU supply-chain issues and runaway cloud costs. In a February 2026 survey of 203 enterprise IT leaders, 79% had already moved some AI workloads out of the public cloud and 93% were repatriating or evaluating it, with 40% saying cloud AI spend blew past projections. There are at least 7-8 “Neoclouds,” running AI training and inference for the new crop of AI-native companies and AI labs, producing massive amounts of agent training and trace data. SRG Research expects this Neocloud market to approach $400B by 2031. So it’s more important than ever to ensure your AI data infrastructure can run wherever you manage data. Models are easy to switch and move; data movement is still costly and cumbersome.
Snowflake has always run most services in its own account, its own closed format, its own cloud storage. Per their latest earnings report, a mere 8.6% of accounts use Apache Iceberg-based open formats. If running anywhere is the goal, Snowflake cannot be your primary choice. Databricks historically favored “classic” BYOC (bring-your-own-cloud), but today recommends “serverless”, with compute in Databricks’ cloud accounts. Nowhere is this clearer than the AI stack: Knowledge Assistant, Supervisor Agent, AI Search and Production monitoring all require serverless compute, with no classic option, and OpenTelemetry ingestion is serverless via Zerobus. Knowledge Assistant only accepts search indexes built on Databricks-hosted embedding models. The agent state follows the compute: both agent products store “temporary data transformations, model checkpoints, and internal metadata that power each agent” in default storage, Databricks’ managed storage, not your bucket; anomaly detection writes its scan results there too. So the agents you tune and the operational metadata they learn live on Databricks infrastructure only.
None of these products shipped a version you can run on your own clusters, and the pattern says that’s by design: every new AI workload lands on Databricks-operated compute. It’s all fine, except you cannot take advantage of your own cloud provider discounts, and Databricks compute cannot run in on-prem or Neocloud environments. This leaves you stranded and strapped to a single cloud data vendor for all your AI needs.
The Onehouse AI Gateway: Any model, Any table, Anywhere
Every AI gateway launched this year promises choice. Databricks pitches “open, multi-AI access without lock-in”; Snowflake calls Cortex “the connective layer for all trusted agent activity”. The real test of choice is what you can still change after you adopt. The Onehouse AI Gateway is built so the answer is: the model, the tables, the catalogs and where it runs.
Any model. Onehouse AI Gateway supports any inference provider with your own credentials: Anthropic and OpenAI today, AWS Bedrock and Baseten very soon. Users publish models behind stable aliases, and existing AI agents can be repointed transparently to route through the gateway, picking up higher uptime from dynamic model switching on failover. There are no Onehouse charges on token volume or otherwise, beyond what you pay your inference providers.
Any table, any catalog. The gateway inherits the universal interoperability Onehouse brings to open tables: Apache XTable™ (Incubating) across formats, and OneSync synchronizing permissions and table metadata across catalogs. Tables are queryable from external engines via MCP servers with tool controls (say, a Snowflake MCP added to the gateway), and natively within Onehouse through Lakegres, purpose-built to feed fast, interactive agent loops. With native support for unstructured data on the same open tables, you can bring any type of data to any model, with data governance and model governance in one gateway.
Anywhere. The gateway is Kubernetes-native and deploys in your account, wherever that account is: the cloud where you have negotiated discounts, the datacenter you’re still paying for, or the Neocloud where your training and inference live. It runs next to the infrastructure that already serves your data, on compute you pay for directly, collecting high-scale observability data in open formats: agent traces, inference activity, latency and estimated spend by caller and model. On AWS, that is the same single r8g.xlarge from earlier: the whole logging stream for $172.02/month, 99.35% lower, with every byte in fully portable data formats.
None of this is abstract; the AI gateway runs on top of our own data lake today. Here is how it comes together, from connecting a provider to querying tables from Claude Code/Codex.
How it works
Under the hood, Onehouse adds support for the major model providers in the industry. Today that is OpenAI and Anthropic, with AWS Bedrock coming quickly for AWS customers and open-weight providers, starting with Baseten, right behind. On top of these, you create logical model providers within the platform for fine-grained control: each carries its own way of supplying API keys for the gateway to talk to the underlying provider, and its own retention policies for the objects associated with it.

At the model level, you configure routing, including uptime optimization: if a base model is unavailable, requests fail over to the fallback models you specify, or you can split traffic, routing X% of requests to one model and Y% to another. Access is just as fine-grained, controlling which users and groups are allowed to call a particular model.

The models page is the single place to see every model currently connected, across all providers in the account.

Each model page shows all activity against that model, and is where you change its routing or the permissions on who can access it.

A model is not very useful without context, and context connects over MCP. The gateway ships with a built-in MCP server for Onehouse’s own services, including Lakegres, our interactive query engine that exposes every table in your data lake while enforcing the governance policies already set on those tables. Alongside it, you can register additional MCP servers: for most companies the operational context lives in the lake, while the rest sits in documents, wikis, GitHub and ticketing systems. All configured MCP servers are namespaced and exposed to an agent such as Claude Code or Codex as a single Gateway MCP server. Every tool call to those servers proxies through the gateway, so the agent connects once and the gateway applies the access controls.

Tool access is controlled centrally too: you decide which groups get which tools, with credentials brokered by the gateway so no agent holds a key. And yes, there is a grant-every-tool option for the auto-mode people.

Once set up, the gateway runs securely inside your own VPC, behind the enterprise security you already have, VPN included. Pointing a coding agent like Claude Code at your context and data is as simple as:
# Paste your token from the gateway token page and hit enter.
# The terminal won't echo it, and read -rs keeps it out of shell history.
read -rs ONEHOUSE_GATEWAY_TOKEN && export ONEHOUSE_GATEWAY_TOKEN
export ANTHROPIC_BASE_URL="<gateway url>"
export ANTHROPIC_CUSTOM_HEADERS="X-Onehouse-Authorization: Bearer $ONEHOUSE_GATEWAY_TOKEN"
export ENABLE_TOOL_SEARCH=true
claude
Below, I am running this against Onehouse’s own internal data lake, asking how much agent coding is actually helping the team move faster. Turns out we are moving about 6x faster this year!
Open the demo full-size or use the player’s fullscreen control. Silent demo: Claude queries Onehouse’s internal data lake through the gateway to analyze engineering productivity.
Where we’re headed
Onehouse AI Gateway now matches Databricks and Snowflake on core gateway functionality, including model access, routing and observability, while preserving open data formats, provider choice and the flexibility to run wherever your data lives. We are building on that foundation with the following capabilities.
Semantic layer. We are extending ingestion to unstructured data, the wikis, Jiras and docs where your data is actually described, and automatically building “semantic” tables that guide inference requests and tool calls. In evals with our Quanton AI Agent, this drastically improved answer quality and cut the number of turns an agent takes to reach a final answer.
Self-organizing data infrastructure. The opportunity we find most exciting: we are moving from humans engineering tables for humans to AI agents reverse-engineering the tables agents need. Agents query on a whim, so no pre-designed system can anticipate their behavior. But with query logs and agent traces flowing through a single point, the gateway can drive the platform to optimize itself, building tables and indexes on open formats, and even automating ETL, from actual agent query patterns.
Tool call optimizations. Ultimately, agents execute a ton of SQL against the lakehouse, and that SQL needs optimizing. The model worries about the semantics of a query and how best to answer it given the context descriptions; shaping those queries over time takes a deep understanding of the underlying infrastructure and table layouts. That is our job, and the gateway sits in the right place to be the central query shaper for everything hitting your tables.
Every open model and Neocloud inference service. Onehouse already runs on any Kubernetes cluster, and we plan to bring open-weight and other proprietary foundation model providers in as built-ins fast. By the end of the year, we want inference services from any of the 7-8 Neoclouds, plugged into the gateway as just another provider, serving models next to wherever your data lives and helping you test/switch providers seamlessly, as the AI revolution gets bigger and bigger.
How are you using your data platform with AI today? What emerging use cases are you seeing? We would love to hear. Get in touch if you want to try the Onehouse AI Gateway, chat with us on our Slack community, or email us.
AI gateway and governance FAQs
What is an AI gateway?
An AI gateway is a shared entry point between agents and their models and tools. It handles model routing, authentication, access controls and observability so each agent does not need its own integration with every provider. Onehouse AI Gateway combines inference and MCP gateways, governing both model requests and tool calls while connecting agents to data through MCP servers and Lakegres.
How does an AI gateway support AI governance?
It gives teams one place to control which users and groups can call a model or use a tool, broker provider credentials, and observe inference activity. Onehouse brings model governance and data governance together: Lakegres enforces existing table permissions, while the gateway controls model and MCP tool access and collects traces, latency and estimated spend by caller and model.
Can I use my own inference providers and API keys?
Yes. Onehouse AI Gateway uses your provider credentials and preserves your direct inference provider relationships. OpenAI and Anthropic are supported today; AWS Bedrock and Baseten are coming next. Stable model aliases let agents use routing and fallback policies without changing their integrations every time a model changes.
Can Onehouse AI Gateway run in my VPC, on-premises or a neocloud?
Yes. The gateway is Kubernetes-native and runs in your account, including your cloud VPC, an on-premises datacenter or a neocloud. Agent traces and inference activity are stored in open formats on infrastructure you control.
Does Onehouse add a per-token inference charge?
No. Onehouse does not add a charge on token volume routed through the gateway. You retain your inference provider bill and pay for the infrastructure running the gateway in your account.
Author

Vinoth Chandar
CEO
Onehouse founder/CEO; Original creator and PMC Chair of Apache Hudi. Experience includes Confluent, Uber, Box, LinkedIn, Oracle. Education: Anna University / MIT; UT Austin. Onehouse author and speaker.


































































































