Talk with our engineers by joining the new Onehouse community Slack!

How Apache Gluten works, and what to watch for

Apache Gluten: what to watch for. A Spark plan flows through a Gluten tube into the Velox engine, ending at a magnifying glass marked watch-outs.

TL;DR

Apache Gluten™ speeds up Spark SQL by running the operators it supports on a native engine, usually Velox, while Spark still plans and schedules the job. The gains are real on CPU-bound scans, joins and aggregations, but they hold in production only if three things check out: how much of each job falls back to vanilla Spark, where results differ from Spark's, and whether executors survive off-heap memory pressure. Gluten can only run Spark's plan faster; Quanton, also built on Velox, optimizes the read, the plan and the write as well, and ran about 4.2x faster than tuned Gluten on TPC-DS 10 TB.

Apache Gluten™ is an open-source plugin that makes Apache Spark™ SQL faster by handing the heavy parts of a query to a native C++ engine, most often Velox, the open-source C++ engine that originated at Meta. It needs no query changes, it is licensed under Apache 2.0, and on the workloads it covers the speedup is real.

It also changes how an executor uses memory, how a plan executes, and in some cases what a query returns. Gluten’s value on your cluster depends on how much of each job actually runs natively, and on whether the job stays stable once most of its working set moves off the JVM heap.

What Apache Gluten is

Gluten is a middle layer between Spark and a native execution engine. The name is Latin for “glue,” and the design goal is exactly that: keep Spark’s planner, scheduler, and distributed control flow, and offload the compute-intensive data processing to a faster native backend. The Apache Gluten project site describes it as a way to offload JVM-based SQL engines to native engines.

Intel and Kyligence started the project in 2022, and it became an Apache Top-Level Project in March 2026. Contributors now include Alibaba Cloud, Baidu, Meituan, Microsoft, IBM, and Google, and several commercial Spark platforms build on it. Microsoft Fabric’s native execution engine is based on Gluten and Velox, as is IBM watsonx.data’s accelerated Spark engine.

Conceptually, Gluten sits in the same category as Databricks Photon and Apache DataFusion™ Comet: a native, vectorized execution layer under Spark. Photon is proprietary to the Databricks runtime. Gluten, like Comet, is open source and runs on your own Spark 3.4, 3.5, 4.0, or 4.1 deployment.

How Gluten works inside a Spark job

Gluten intercepts a query after Spark has planned it and decides, operator by operator, what can run natively. Two mechanisms explain most of its behavior: the plan translation that feeds the native engine, and the columnar shuffle, memory, and fallback machinery that keeps Spark in charge of everything else.

Diagram: a Spark plan flows into Gluten. Supported operators go on to the Velox engine; unsupported ones fall back to the Spark JVM.
Spark still plans the query. Gluten sends the operators Velox supports to native code and leaves the rest on Spark.

From Spark physical plan to native operators through Substrait

Spark builds its physical plan exactly as it normally would. Gluten then walks that plan, validates each operator and expression against what the native backend supports, and replaces supported nodes with “transformer” nodes. Gluten’s JVM-side plugin converts those segments into a Substrait plan, a cross-language specification for relational operations, and passes it to the native side through JNI.

On the native side, the backend (Velox by default) converts the Substrait plan into its own operator pipeline and executes it on columnar batches. With the Velox backend, that means vectorized, SIMD-friendly C++ operators for scans, filters, projections, hash aggregations, joins, sorts, and window functions. Results come back to Spark as columnar batches through Spark’s Columnar API, using Apache Arrow™ as the in-memory format.

The effect is that Spark still decides what to run and where, while the per-row work that takes up CPU time moves out of generated Java code. The Gluten README credits this split for most of the gain: Spark’s whole-stage codegen has not made individual operators much faster since Spark 2.0, and native engines have.

Columnar shuffle, memory, and the fallback path

Gluten keeps data columnar across stage boundaries with its own ColumnarShuffleManager, which reuses Spark’s shuffle service but writes and reads Gluten’s columnar format. That avoids converting batches back to rows before every exchange.

Memory moves too. The native engine allocates its hash tables, sort buffers, and vectors in off-heap memory, which Gluten tracks through a unified memory manager that reports back to Spark. This is why Gluten requires spark.memory.offHeap.enabled=true and a large off-heap size, and why garbage collection time drops sharply once most work runs natively.

When Gluten meets an operator, expression, data type, or file format it cannot run, it falls back to vanilla Spark for that part of the plan. The query still completes. The cost is a pair of conversion operators, ColumnarToRow and RowToColumnar, at every boundary between native and JVM execution, plus ordinary Spark speed for the fallen-back segment.

Gluten’s native backends: Velox, ClickHouse, and Bolt

Gluten is built to be backend-agnostic, and the backend choice determines which operators and functions are available natively. Meta’s Velox library is the default and by far the most widely deployed backend. It is a C++ execution library also used in Presto C++ and several other engines, which is why “Gluten + Velox” is the usual shorthand for the project.

The ClickHouse backend, driven largely by Kyligence, uses ClickHouse’s execution engine instead. A third backend, Bolt, was added to the Gluten codebase in 2026. Coverage differs across all three, so the operator and function support you read about for Velox does not carry over automatically to the others.

Most public benchmarks, production write-ups, and vendor integrations (Fabric, IBM, Palantir) use Velox. Unless you have a specific reason to pick another backend, the Velox path has the most documentation and the most production use.

Setting up Gluten on an existing Spark cluster

Gluten ships as a bundle JAR per Spark version and platform, and enabling it is a handful of Spark configs rather than a code change. A minimal setup looks like this:

spark-shell \
  --conf spark.plugins=org.apache.gluten.GlutenPlugin \
  --conf spark.memory.offHeap.enabled=true \
  --conf spark.memory.offHeap.size=20g \
  --conf spark.shuffle.manager=org.apache.spark.shuffle.sort.ColumnarShuffleManager \
  --jars gluten-velox-bundle-<spark-version>-<platform>.jar

The off-heap size is the setting that needs the most thought. Native operators live in that budget, but any segment that falls back to Spark still needs JVM heap. IBM’s watsonx.data defaults 75% of executor memory to off-heap and IBM’s Gluten engine docs recommend dropping below 50% for workloads with heavy fallback, so the right split depends on your native coverage.

Beyond the basics, the Gluten configuration reference lists hundreds of settings for per-task memory, spill, shuffle, and backend behavior. Before any of that, run Gluten’s qualification tool against your Spark event logs. It estimates how much of your existing SQL workload Gluten can offload. That estimate is a better starting point than a generic benchmark.

What Gluten’s benchmarks show

Gluten’s published benchmarks report a 3.02x overall speedup for Gluten with Velox over open-source Spark on TPC-DS and 3.34x on TPC-H at 3 TB, with individual queries up to 13.75x faster on TPC-DS and 23.45x faster on TPC-H. Those runs used Intel Xeon Platinum 8592+ hardware with Spark 3.3.1 as the baseline. Newer Spark releases have improved the JVM baseline since 3.3, so the gap on a current version may be different. The Gluten README reports a lower 2.71x overall speedup on TPC-H at 2 TB, from a June 2023 single-node test against Spark 3.3.2.

An independent AWS Labs benchmark ran Gluten 1.6.0 against Comet 0.16.0 on TPC-DS 3 TB over Apache Parquet™ on Amazon EKS. Gluten finished about 9% faster overall (2,239.93 s vs 2,467.49 s), but Comet finished faster on a majority of queries: 39 of 103 queries ran 20% or more faster on Comet, against 18 on Gluten. Gluten’s largest wins came from choosing shuffled hash joins over sort-merge joins on big fact tables, and the same choice made it 5.95x slower than Comet on q72 (193.5 s vs 32.5 s).

That cluster ran 23 executors with 58 GB each, close to 1.3 TB of executor memory for a 3 TB dataset. A memory-to-data ratio that high lets a native engine keep most of its working set in memory. Production clusters usually run with less memory per terabyte, so the same result may not hold where off-heap pressure is higher.

What to watch for before running Gluten in production

The risks with Gluten are rarely about the happy path. They show up at the edges: the operators it does not cover, the places its results differ from Spark’s, and the way native memory behaves under pressure. The table summarizes the three, and each is covered below.

Watch-out What you see Where to check
Fallbacks to vanilla Spark Speedup far below benchmarks; ColumnarToRow and RowToColumnar nodes in the plan Query plan, Gluten UI tab, qualification tool
Correctness differences Results that differ from vanilla Spark on JSON, dates, regex, or NaN Velox backend limitations doc; diff outputs against Spark
Off-heap memory pressure Executors killed with exit code 137 (OOMKilled) on large joins and aggregations Container metrics, off-heap sizing, shuffle partition count

Behavior reflects Gluten 1.7 with the Velox backend on open-source Spark; limitations move with each release, so confirm against the current project docs.

Fallbacks that quietly erase the speedup

Some fallbacks are whole-plan, not per-operator. Open-source Gluten does not yet support ANSI mode, so with ANSI enabled the entire plan runs on vanilla Spark. Spark 4.0 turns ANSI on by default, so this is the first thing to check on a Spark 4 cluster. Progress is tracked in the ANSI support issue.

File formats are next. Gluten’s Velox backend fully supports Parquet and only partially supports Apache ORC™, and scans of other formats fall back. Partitioned table scans fall back when the partition values are not encoded in the file path, bucketed writes fall back, and so does a Parquet scan that hits a byte-type column.

The long tail is expressions. Anything not in the operator support matrix runs on Spark, and each native-to-JVM boundary adds conversion overhead. A plan with scattered fallbacks can run slower than plain Spark, because it pays for the conversions without keeping enough work native to recover the cost.

Diagram: execution moves from native, through a conversion step, to the Spark JVM, through another conversion, and back to native.
Every switch between native and JVM execution adds a conversion, so scattered fallbacks can cost more than they save.

Correctness differences from vanilla Spark

Gluten aims for Spark-compatible results, and the project documents the cases where it does not get there. Its Velox backend limitations page lists several. Velox’s JSON functions accept only double-quoted strings, so single-quoted JSON produces incorrect results. Regex functions use RE2 rather than java.util.regex, which means no lookaround patterns and slightly different whitespace matching.

Dates and numbers are the other category. Gluten ignores Spark’s Parquet datetime rebase setting on read, so files written with the legacy hybrid Julian-Gregorian calendar can return different values. Velox does not support NaN, and case-sensitive mode can produce incorrect results because Gluten only supports Spark’s default case-insensitive behavior.

None of these affect most ETL. In the jobs they do affect, the query still succeeds and returns a different value. Diffing Gluten output against vanilla Spark on your real pipelines, before cutover, is the only reliable guard.

Off-heap memory pressure and OOM kills

Moving work off-heap removes GC pauses, but it also moves the working set into memory the JVM does not manage. The container limit is enforced on the whole process. Once resident memory crosses it, the executor is killed, even if Gluten’s own accounting shows headroom. Gluten’s documentation notes that spill can still end in an out-of-memory error when shuffle partitions are set high.

Onehouse’s write-up on why native accelerators OOM traces this to two compounding causes. Some allocations escape the engine’s accounting or cannot spill, and the default Linux allocator keeps freed memory resident across many per-thread arenas. The result can be an executor that is well under budget by its own numbers and still OOMKilled.

In Onehouse’s full 99-query TPC-DS 10 TB run on a fixed cluster, Gluten hit OOM failures on q67 and q93 even after tuning, while Quanton completed all 99 on the same hardware and memory budget. Onehouse’s review of the Gluten and Comet public trackers counts roughly 115 verified issues across per-query performance cliffs, correctness mismatches, off-heap OOM kills, and shuffle spill regressions. Gluten’s public issue tracker is the best place to check whether your workload patterns are among them.

How to confirm what actually ran natively

A job can be “on Gluten” while its most expensive stage runs on the JVM, so read the plan rather than assuming coverage. In the Spark UI SQL tab or in df.explain(), native operators end in Transformer (for example ShuffledHashJoinExecTransformer or HashAggregateExecTransformer). ColumnarToRow and RowToColumnar nodes mark the boundaries where execution drops back to Spark.

Gluten also adds a Gluten SQL tab to the Spark UI that counts native versus fallen-back nodes per query, and Gluten’s implicits give you a programmatic summary. In a Scala session, import org.apache.spark.sql.execution.GlutenImplicits._ and then df.fallbackSummary lists each fallback and the reason for it.

Measure the native share on your largest jobs by runtime, not by query count. A pipeline where 90% of queries are fully native can still spend most of its time in the 10% that fall back.

The ceiling on compute-stage acceleration

Gluten accelerates the compute stage, and it does that well. What it does not change is the plan itself. Gluten executes the plan Spark produced, with Spark’s join order, Spark’s shuffle boundaries, and Spark’s algorithm choices, only faster per row.

A separate Onehouse test of a native ROLLUP operator shows what that means in practice. Spark executes GROUP BY ROLLUP with an Expand operator that copies every input row once per rollup level, nine times on TPC-DS q67. Gluten vectorizes Expand, so it runs the same nine-fold row explosion faster. On q67 at 10 TB, that took 4.4 minutes on open-source Spark and 2.7 minutes on tuned Gluten, against 55 seconds with a rollup operator that removes the explosion.

The same limit applies to storage. Gluten does not read Apache Iceberg™ or Apache Hudi™ table metadata to reshape a plan, and it has no index to find the rows a MERGE touches without scanning. For upsert-heavy pipelines, compaction, and MERGE, the most expensive work often sits outside the part of the job a compute-stage accelerator can reach.

Quanton: a Velox-based engine that works past the compute stage

Quanton starts from the same foundation as Gluten, a vectorized C++ engine on Velox, and changes what the engine is allowed to optimize. That affects two things an evaluator weighs against a free accelerator: how much faster the jobs run, and what it costs to try.

Storage-aware planning on an optimized Velox fork

The Quanton engine is a drop-in replacement for Spark’s execution engine, built on an optimized fork of Velox with reimplemented operators. Quanton also optimizes the read and the write. It reads Iceberg and Hudi metadata to reshape plans, runs index-aware joins that locate rows instead of scanning for them, and ships native low-shuffle MERGE and compaction that stay columnar.

On TPC-DS 10 TB (99 queries, 11 × m8gd.4xlarge, identical hardware), Quanton ran about 4.2x faster than tuned Gluten across the full suite (2,034 s vs 8,563 s). Against open-source Spark (12,200 s), Quanton ran 6x faster. Quanton’s memory handling (a purpose-built allocator, native memory bounded under the executor limit, and operators that spill instead of failing) targets the OOM pattern described above. These results are derived from TPC-DS and figures move with each release, so confirm them in the Quanton benchmark results.

Open core, a free tier, and a one-line path back

Quanton is open core, and it reverts to open-source Spark with a single config line. The free tier covers the first 100 GiB processed each month, and past that cap it falls back to open-source Spark with Gluten and Comet pre-installed. The side-by-side against Gluten runs on the same cluster with no separate setup.

Quanton bills per GiB of data processed rather than per compute hour, and it runs in your own cloud account, so your reserved and spot discounts stay yours. Because billing is per GiB, a faster job cuts instance hours without raising the Quanton fee. The current rates and a TCO comparison against Gluten, Comet, and Databricks are on the Quanton pricing page.

Should you run Apache Gluten?

Gluten is worth evaluating. It is free, actively developed, backed by a broad contributor base, and faster on CPU-bound scan, join, and aggregation work over Parquet. Start with the qualification tool, check ANSI mode and your file formats, size off-heap memory for your fallback share, and diff results against vanilla Spark before you cut over.

If the fallback share stays high, if large jobs hit off-heap OOMs, or if your cost lives in MERGE, compaction, and writes, Gluten has reached the limit of what compute-stage acceleration can do. Quanton builds on the same Velox foundation, optimizes the read, the plan, and the write, and reverts to open-source Spark with one config line. You can run it against Gluten on the same jobs on the free tier: create an account and compare the numbers in your own VPC.

Frequently asked questions

Is Apache Gluten free and open source?

Yes. Gluten is an Apache Software Foundation Top-Level Project licensed under Apache 2.0. You download a bundle JAR for your Spark version and enable it with Spark configs, with no vendor license involved.

Does Apache Gluten support Spark 4?

Yes. Current Gluten releases support Spark 3.4, 3.5, 4.0, and 4.1, and support for Spark 3.3 has been removed. On Spark 4, check ANSI mode first: it is on by default in Spark 4.0, and open-source Gluten falls back to vanilla Spark for the whole plan when ANSI is enabled.

What is the difference between Gluten and Velox?

Velox is an open-source C++ execution engine that originated at Meta. It provides vectorized operators, but it does not understand Spark plans on its own. Gluten is the layer that translates Spark’s physical plan into something Velox can execute, manages memory and shuffle between the two, and falls back to Spark where Velox has no implementation.

How does Apache Gluten compare to DataFusion Comet?

Both are Spark plugins that run supported operators on a native engine: Gluten mostly on Velox in C++, Comet on DataFusion in Rust. In the AWS Labs TPC-DS 3 TB benchmark, Gluten finished about 9% faster overall while Comet won more individual queries. The Comet project’s comparison covers the design differences, and benchmarking both on your own jobs is the reliable way to choose.

Is Microsoft Fabric’s native execution engine the same as Gluten?

Fabric’s native execution engine is built on Gluten and Velox, with Microsoft-specific additions such as ANSI support on Fabric Runtime 2.0 and Delta Lake optimizations. Its behavior is close to open-source Gluten but not identical, so Fabric’s documented limitations and open-source Gluten’s limitations should be read separately.

Author

Bhavani Sudha Saktheeswaran

Head of Open Source

Sudha is the Head of Open Source at Onehouse and a PMC member of the Apache Hudi project. She comes with vast experience in real-time and distributed data systems through her work at Moverworks, Uber and Linkedin’s data infra teams. She is a key contributor to the early Presto integrations of Hudi. She is passionate about engaging with and driving the Hudi community.

Read More:

Subscribe to the Blog

Be the first to hear about news and product updates

We are hiring diverse, world-class talent — join us in building the future