Talk with our engineers by joining the new Onehouse community Slack!

Snowflake, BigQuery, Redshift: If Iceberg Is the Future, Why Is Closed Format Still Your Default?

Anakin and Padmé meme. Snowflake, BigQuery and Redshift in 2024: “The future is open. Use Iceberg.” Data community: “So you’ll retire your closed formats?” The same warehouses in 2026 say nothing, and the data community asks: “You’re still planning to retire them... right?”

TL;DR

Since 2024, Snowflake, BigQuery and Redshift have added real Apache Iceberg™ support, but all three still default to their closed native formats and ship new features such as vector types and vector indexes there first. None has published a timeline to retire their closed formats. One project cannot absorb every vendor's new features, so data interoperability alongside Iceberg as the lowest common denominator remains the practical path to open data.

Last week, Databricks CEO Ali Ghodsi and I exchanged comments on LinkedIn about how Databricks and Onehouse support open table formats. We were debating recurring questions: Have we standardized yet? Is it a standardized implementation of the standard? Halfway through, I realized that only open format vendors are expected to debate their choices this way. I had hoped Databricks, with its larger audience, would raise the broader question for all of us. But I am raising it myself here anyway:

Why do we still accept proprietary native formats as defaults in Snowflake, BigQuery and Redshift?

Since the summer of 2024, we’ve assumed Apache Iceberg™ is the new Apache Hive™: the common open table format all data vendors will support. Today, Onehouse’s engine Quanton natively supports Iceberg, and Apache Hudi™ can produce Iceberg tables alongside its high-performance native format. We have our disagreements, but Databricks has introduced Iceberg support and Delta Lake is trying to converge to Iceberg. Databricks (with Tabular absorbed) and Onehouse have championed the Open Data Lakehouse for years. Both have always stored all analytical data ONLY in open file formats. Onehouse has been busy bringing cutting-edge features like unstructured data and vector search to the open lakehouse.

Meanwhile, the three big cloud warehouses continue to support and innovate on closed file/table formats in plain sight. Of course, I have a commercial stake in this debate and strong technical opinions about what a high-performance open table format should look like. Databricks should answer hard questions about interoperability. Onehouse should too. So should all three warehouses. The engineers operating and eventually migrating these systems deserve the same candor from all of us.

In this blog, I unpack how the big warehouses have walked the talk on openness and the practical limits of controlling open-source, open data innovation by a committee.

Wait a minute meme. Top caption: We're arguing over which open format is more open... Bottom caption: Snowflake, BigQuery and Redshift still default to closed format?

What’s happened since the summer of 2024?

To be clear, the warehouses have done substantial work to support Iceberg. But customers still have to opt into that future, while native formats acquire capabilities the Iceberg paths lack. I can understand supporting closed formats for old workloads. Building new features on them first raises a different question about where the focus still is. The announcements since 2024 show how much the vendors have already committed to Iceberg.

Date What changed
June 3–10, 2024 Snowflake announced Polaris (now Apache Polaris™), Databricks announced its Tabular acquisition, and Snowflake released Iceberg tables to general availability. Polaris, Tabular, Snowflake GA.
October 2–11, 2024 AWS announced the retirement of Lake Formation Governed Tables in favor of open formats. Google announced managed BigQuery tables for Iceberg in preview. AWS notice, Google announcement.
December 3, 2024 SageMaker Lakehouse brought Iceberg-compatible access across S3 and Redshift. Announcement.
May–June 2025 Google announced GA for managed Iceberg tables; the June 3 release notes recorded it. Announcement, release notes.
October 30, 2025 Google’s release notes recorded general availability of the BigLake Iceberg REST catalog. Release notes.
November 17, 2025 Redshift added generally available Iceberg writes for append-only workloads. Announcement.
April–June 2026 Google announced a preview of multi-engine read/write interoperability. Redshift added UPDATE, DELETE and MERGE. Snowflake made configurable Iceberg defaults generally available. Google, AWS, Snowflake.
July 13, 2026 BigQuery’s release notes recorded GA for partitioning, multi-statement transactions and advanced runtime on Iceberg managed tables. Release notes.
August 31, 2026 Redshift added Iceberg v3 reads and writes on supported deployments. Announcement.
October 5, 2026 Redshift added creation and refresh of Iceberg materialized views. Announcement.

So, real progress on reading/writing Iceberg. But let’s remember most already read/wrote Apache Parquet™ and Hive. What I haven’t found is a public timetable for retiring their closed native formats. Without that, we’ll never fully be in a world where data is open. Whatever your favorite open table format, this is the single most important thing for those of us who’ve spent a decade trying to make open formats the ONLY way to store data and end vendor lock-in.

Each vendor exposes a different part of that gap: Snowflake has made the default configurable, BigQuery keeps important AI capabilities tied to standard tables, and Redshift is expanding interoperability while retaining its native storage path.

Snowflake offers an Iceberg default. Still no public retirement date for closed formats.

Snowflake already lets you make Iceberg the default. Yet it still ships with its own closed format as the default. Since June 5, 2026, nearly two years after Iceberg tables reached GA, customers can configure DEFAULT_METADATA_WRITE_FORMAT at account, database or schema scope. With the required catalog configuration, ordinary CREATE TABLE statements can produce Iceberg tables. Snowflake also offers managed Iceberg storage on supported AWS and Azure deployments, with external access through Horizon, so customers need not configure their own bucket. The managed experience can sit on an open format. Standard tables still use Snowflake’s internal compressed columnar format and micro-partitions. And Horizon’s Iceberg REST endpoint, the path external engines use, exposes Iceberg tables; it doesn’t make native Snowflake storage independently readable by Iceberg engines.

In at least half of Onehouse’s sales conversations with Snowflake customers, they want an open format but don’t know what to do with years of locked-in data. Official guidance from Snowflake would help immensely. Changing the default doesn’t convert existing tables. CREATE ICEBERG TABLE ... AS SELECT can populate a new Iceberg table, and SnowConvert supports Iceberg targets for some migrations, including Teradata, whose settings still default to native tables. CONVERT TO MANAGED changes catalog management for a table that’s already Iceberg. These tools help. But where is Snowflake’s plan for customers sitting on years of closed-format data?

Snowflake CoCo answering "What do I use? Snowflake native format or iceberg?" It recommends native tables when you want the full Snowflake feature set and when performance is the top priority, and Iceberg tables for multi-engine access, no lock-in, existing data lakes and external catalogs.
Snowflake CoCo, asked “What do I use? Snowflake native format or iceberg?” on October 7, 2026.

Meanwhile, Snowflake keeps building new features on its closed format. Snowflake made VECTOR generally available in May 2024 and FILE in September 2025. FILE holds references to staged documents and images for multimodal AI. Both VECTOR and FILE remain explicitly unsupported in Iceberg tables. Iceberg still has no vector type. An upstream proposal was opened on September 28, 2026, more than two years after Snowflake made VECTOR generally available.

Customers can still build AI applications on Iceberg by modeling these as logical types themselves: embeddings can be stored as arrays and cast to vectors, and FILE is a reference, not the document itself. Snowflake supports Iceberg inputs alongside UDFs and Cortex functions in dynamic-table queries and added Iceberg search optimization in 2025. So customers choosing open storage still have to work around these gaps. Snowflake, your June 2026 announcement promises agency over data, and your EVP of Product told Summit that Snowflake is “fully committed to interoperability and openness.” Make Iceberg the shipped default for suitable new analytical workloads, publish the remaining blockers, and give existing customers a supported transition. Why should each customer have to opt into the future you’re already promoting?

BigQuery advances AI on native tables. Still no public retirement date for closed formats.

BigQuery has done substantial work: managed Iceberg tables live in customer-owned buckets, with maintenance and SQL mutations. The newer Lakehouse REST catalog path gives open engines read/write access, with BigQuery’s own writes to those tables still in preview, while the BigQuery-managed path accepts no writes from open engines, so customers modify data through BigQuery. Yet a regular load without an existing Iceberg managed target still creates a standard BigQuery table. The open path remains something customers have to choose.

There are useful migration tools too. BigQuery can populate a new Iceberg table from a query, although CREATE OR REPLACE cannot convert a standard table directly. A Data Transfer Service preview lands data from S3, Azure Blob Storage and Cloud Storage in Iceberg managed tables. The catalog-migration preview handles external Hive and Iceberg metadata. These help customers copy data or start open, but still don’t say when Google intends to move away from its closed format.

But BigQuery makes the gap between closed and open formats surprisingly clear: new AI features are still being built on native tables. Google’s own feature matrix says both Iceberg management paths support BigQuery ML and AI read queries, but neither supports search indexes or vector indexes, including automatic embedding generation. The vector-index DDL explicitly limits creation to standard tables. Meanwhile, vector search and IVF indexes reached GA in September 2024, autonomous embedding generation entered preview in December 2025 and reached GA in June 2026, and partitioned TreeAH indexes and online index rebuilding reached GA in April 2026. So customers can run AI queries on Iceberg, but choosing open tables still means giving up these indexing and embedding capabilities. Governance follows the same pattern: the same matrix shows BigQuery row-level security and authorized views are unsupported on both Iceberg paths.

I can understand Google’s argument that its native storage architecture, using Capacitor, lets storage and execution evolve together. It’s the same argument Onehouse, Databricks and ClickHouse make for their own open formats. BigQuery, why is shipping innovations faster a valid engineering argument for your closed format, but anti-standardization when open formats do it? It’s a weird paradox. Your April 2026 announcement called this “Openness without compromises.” When will these indexed AI features work on Iceberg? When will open tables become the default for suitable new analytical workloads? And what’s the plan for customers already sitting on years of native data?

Redshift opens access. Still no public retirement date for closed formats.

With Redshift, what exactly is becoming open? AWS lets other engines access Redshift Managed Storage through Iceberg-compatible interfaces. But the data still sits in RMS, with a managed Redshift workgroup handling compute when those catalogs are queried. Giving another engine access doesn’t mean your warehouse has become a collection of Iceberg files you can use independently.

Redshift has done substantial work on actual Iceberg tables too. Query support reached GA in November 2023. It now supports direct writes and row mutations, and CREATE TABLE ... USING ICEBERG ... AS SELECT can copy selected native-table data into Iceberg, subject to supported types and configuration. Yet ordinary local table creation still takes the native path. The October 2026 Iceberg materialized views have open inputs and outputs, but their sources must already be Iceberg tables, and Iceberg v3 sources aren’t supported, even though Redshift added v3 support in August. They don’t migrate native tables, lack automatic refresh and automatic query rewriting, and run only on Serverless and RG instances. AWS’s own launch post says they don’t replace standard materialized views in RMS, and suggests loading them back into RMS for latency-sensitive dashboards.

Meanwhile, Redshift keeps improving its native storage. Multidimensional Data Layouts reached GA in September 2025, using query filters to organize rows under SORTKEY AUTO. Again, this goes beyond keeping old workloads running. AWS is building new storage optimizations while its Iceberg path still has automation gaps. That doesn’t mean Redshift ML or UDFs generally require proprietary tables; these are specific gaps.

Redshift, tell customers where native storage is still preferred for performance or compatibility, why, and what needs to change. Everywhere else, make open tables the normal choice and help customers move out of closed formats. For customers who already want more independence between storage and compute for large datasets, Iceberg on S3 should be an easier case to make. Arguably, Redshift has even less reason to leave them waiting.

Delivering every new capability through one project to all engines is impractical

At the end of the day, warehouse teams and open-format vendors like Databricks and Onehouse are building new storage and engine features for their own users, commercial or open-source. Those engineering projects cannot all wait on a single project’s process. The Iceberg PMC stewards Iceberg; the broader community still has to debate real engineering questions: keeping indexes correct, representing new types, coordinating writers. That takes time, as it should. Even Iceberg’s implementation-status matrix shows different levels of support across language implementations. A feature can enter the specification before engines implement it, or ship in a product before anyone proposes a common representation.

That brings our view of format unification, and the standardization policing around it, into question. Why should we expect Iceberg to lead every storage innovation? Iceberg v3’s deletion vectors use the same binary encoding as Delta’s. Its row lineage addresses incremental-processing problems that Hudi has long tackled with record-level commit metadata. Snowflake had years of experience with VARIANT before the Iceberg and Parquet communities aligned on an open representation. Proposals for Iceberg v4’s metadata tree borrow heavily from Hudi and Delta Lake’s transactional log designs to avoid metadata bloat. These examples show how ideas can develop in individual systems and become interoperable across projects.

The right way to think about Iceberg is as the “lowest common denominator”: the capabilities you can actually rely on across the engines you use. That isn’t putting Iceberg down. It’s an important role, and the backing from the big five—the three big clouds, Databricks and Snowflake—is a good thing. That common ground will grow. But expecting it to include every vendor’s newest feature the moment it ships is unrealistic.

Some execution techniques and rebuildable indexes don’t even need a table-spec change. Vendors may keep those closed as a competitive moat. We should distinguish those choices from keeping the underlying data in a proprietary format. That is where the consequences for users become very different.

Constraining open-source builders does not bode well for the industry

So, the key difference is where these bleeding-edge features land. Onehouse builds around open formats. Our innovations land in our own Iceberg fork and in Hudi, an independent Apache top-level project that has been around for almost a decade now, with its own storage and transaction model, merge-on-read logs and incremental queries. Databricks continues developing Delta Lake alongside Iceberg support. UniForm is still around, generating Iceberg metadata alongside Delta metadata over the same Parquet files. Delta remains a separate table protocol, but separate metadata doesn’t necessarily mean another copy of the data. ClickHouse’s open-source implementation has its own MergeTree layout, with sorted data parts and indexes. Its cloud architecture has also evolved toward SharedMergeTree over object storage. It doesn’t have to become Iceberg to be open source. An open implementation lets users inspect, operate and keep developing a system outside the vendor. Interoperability can also let them change engines without rewriting data.

New features landing in proprietary formats remain the key concern of this blog. Hudi, Delta Lake and Apache Paimon™ have different metadata designs; that does not mean data must be duplicated. UniForm and Apache XTable™ show how compatible representations can share files. When warehouse features require closed formats, moving that data to Iceberg can mean a rewrite and a second copy during migration. If each new feature sends users back to closed formats until the Iceberg path catches up, they risk a perennial cycle of migrations. An optional Iceberg integration cannot solve that.

The warehouses keep innovating outside their Iceberg paths: BigQuery builds indexed AI on standard tables, Snowflake adds AI data types its Iceberg tables can’t persist, and Redshift improves native layouts while its Iceberg automation catches up. So why should other open systems face a different test? You can’t defend your proprietary format as necessary for innovation, then dismiss someone else’s open format because it isn’t Iceberg yet. You’re asking open-source builders to accept constraints you won’t accept yourselves.

Open Asks

I’m not asking these questions after a month. We’ve given all three warehouses more than two years. If the easy answer is “there’s a lot of data sitting in these closed formats,” that’s still not good enough. I can understand why moving that data takes time. But why does that stop you from publicly deprecating the format and giving users a timeline? Build the migration tools. Tell us which workloads still need closed formats and what needs to change. AWS announced the retirement of Lake Formation Governed Tables in October 2024, effective that December, in favor of Iceberg, Hudi and Delta Lake. A different product, yes, but a public commitment. I still haven’t found one for these warehouses’ native formats.

Users also deserve to know what will actually work when they move: which engines can write, what happens to history and permissions, and which AI features they give up. Keeping a proprietary index over an open format that can be rebuilt is one thing. Keeping the only usable copy of the data in a closed format is another. Of course, Onehouse and Databricks should answer the same questions about our own catalogs and services. We don’t get a free pass because we built an open-source project. Neither should the warehouses get one because they added Iceberg support.

And this is an appeal to users too. Please think more carefully and judiciously about open data. Ask what happens when you create a table, use a new feature or try to leave. If you don’t see through these issues for yourselves, open vendors alone cannot change the status quo. The choices you make matter as much as the formats we build. Without your watchful eyes, we risk slipping back into the 2010s where vendors happily locked in all your data into closed formats and threw the keys away.

Maybe we all bit off more than we could chew with this dream of grand unification. There will be new formats, vendors moving at different speeds, and different levels of functionality even within Iceberg. If its biggest champions can’t stick to Iceberg for their own new features, how do they expect everyone else to follow? We need to be realistic about what one format can achieve. After a decade of working on this, I still come back to the same answer: interoperability is the lasting solution.

Author

Profile Picture of Vinoth Chandar, ‍CEO/Founder

Vinoth Chandar

‍CEO

Onehouse founder/CEO; Original creator and PMC Chair of Apache Hudi. Experience includes Confluent, Uber, Box, LinkedIn, Oracle. Education: Anna University / MIT; UT Austin. Onehouse author and speaker.

Read More:

Subscribe to the Blog

Be the first to hear about news and product updates

We are hiring diverse, world-class talent — join us in building the future