Read More:

Onehouse LakeBase is now Lakegres™
Vinoth ChandarJuly 2, 2026
Databricks Iceberg Support Has a Catch. It's Called Unity Catalog.
Kyle WellerJune 11, 2026
Introducing the Quanton Kubernetes Operator for Apache Spark™
Vinoth ChandarMarch 24, 2026
Announcing OneSync™ Permissions: Unified Access Control Across All Your Data Catalogs
Andy Walner, Kyle Weller and Roushan KumarMarch 13, 2026
Bringing Onehouse Cloud to Microsoft Azure
Manohar Kiran, Kyle Weller and Nilesh MahajanMarch 10, 2026
Announcing Onehouse Lakegres™: database speeds finally on the lakehouse
Vinoth ChandarFebruary 17, 2026
Onehouse 2025 Year in Review
Vinoth ChandarJanuary 15, 2026
Choosing Between a Database and a Data Lake
Ahilya Kulkarni and Shiyan XuJanuary 8, 2026
Orchestrating Spark Pipelines on Onehouse with Apache Airflow
Andy Walner and Sagar LakshmipathyDecember 17, 2025
Designing Your Data Lakehouse Tables for Fast Queries
Andy WalnerDecember 10, 2025
Inflated data lakehouse costs and latencies? - Blame S3’s choice of HTTP/1.1
Rajesh MahindraDecember 3, 2025
Onehouse Quanton vs the latest AWS EMR for Apache Spark™ Workloads
Kyle WellerDecember 2, 2025
Introducing Onehouse Notebooks – Interactive PySpark at 4x Price-Performance
Andy Walner, Divik Mittal and Praveen GajulapalliDecember 1, 2025
Apache Iceberg™ on Quanton: 3x Faster Apache Spark™ workloads
Kyle Weller and Rajesh MahindraNovember 12, 2025
Securing Your Data Lakehouse: Best Practices for Data Encryption, Access Control, and Compliance
Amakiri Welekwe and Shiyan XuNovember 6, 2025
Choosing the Right Data Ingestion Method: Batch, Streaming, and Hybrid Approaches
Gaurav Thalpati and Bhavani Sudha SaktheeswaranOctober 28, 2025
Optimizing Performance in Open Source Data Warehouses: Query Tuning, Data Partitioning, and Caching Strategies
Roel Peters and Shiyan XuOctober 23, 2025
Data Lake vs. Warehouse vs. Lakehouse
Alen Kalac and Shiyan XuOctober 16, 2025

Data Architecture Survey Report: The Lakehouse Is Your Data Foundation for AI
Gaetan CasteleinOctober 9, 2025
Hudi
Apache Iceberg™ vs Delta Lake vs Apache Hudi™ - Feature Comparison Deep Dive
Kyle WellerOctober 2, 2025
How to Save Apache Iceberg™ from Chilly Meltdown When Dimensions Too Hot to Handle
Shiyan XuSeptember 26, 2025
Meet Cost Analyzer – a free tool to unearth Apache Spark™ bottlenecks
Chandra Krishnan and Shubham PatelSeptember 25, 2025
Why the Apache Spark™ Default Autoscaler Fails Your Lakehouse (and How We Fixed It)
Rajesh MahindraSeptember 17, 2025
Protecting Your Data Lake from Internal and External Threats: Security Strategies
Thinus Swart & Shiyan XuSeptember 9, 2025
Introducing Onehouse OneFlow: Ingest once, query anywhere
Andy Walner, Jyasveer Gotta and Vishal KumarAugust 1, 2025

AWS S3 Tables : After the 10x Priceberg Plunge
Vinish Reddy Pannala and Kyle WellerJuly 16, 2025

S3 Managed Tables, Unmanaged Costs: The 20x Surprise with AWS S3 Tables
Vinish Reddy Pannala and Kyle WellerJuly 8, 2025
From the trenches: Managing Apache Iceberg metadata for near-real-time workloads
Vinish Reddy PannalaMay 29, 2025
Announcements
Announcing Apache Spark™ and SQL on the Onehouse Compute Runtime with Quanton
Rajesh Mahindra, Andy Walner and Vinoth ChandarMay 20, 2025
Measuring ETL Price-Performance On Cloud Data Platforms
Daniel Lee, Rajesh Mahindra and Vinoth ChandarMay 15, 2025
Towards Open Data - Part 1: Cloud Warehouses Now Love Open Formats
Dipankar Mazumdar and Vinoth ChandarApril 28, 2025
Announcements
Announcing Open Engines™: Flipping defaults to “open” for both data and compute
Vinoth ChandarApril 17, 2025

Product
Apache Flink™ vs Apache Kafka™ Streams vs Apache Spark™ Structured Streaming — Comparing Stream Processing Engines
Sagar LakshmipathyApril 17, 2025

Product
ClickHouse vs StarRocks vs Presto vs Trino vs Apache Spark™ — Comparing Analytics Engines
Chandra KrishnanApril 17, 2025

Product
Ray vs Dask vs Apache Spark™ — Comparing Data Science & Machine Learning Engines
Andy WalnerApril 17, 2025
Data Deduplication Strategies in an Open Lakehouse Architecture
Dipankar Mazumdar and Aditya GoenkaMarch 20, 2025

The Open Table Format War: Merely a Battle on the Path to Engineering a Truly Open Data Platform
Pauline BrownMarch 12, 2025

ACID Transactions in an Open Data Lakehouse
Dipankar MazumdarFebruary 20, 2025

Hudi
Using Apache Hudi™ Data with Apache Iceberg™ and Delta Lake
Bhavani Sudha SaktheeswaranFebruary 13, 2025
Moving Beyond Lambda: The Unified Apache Beam Model for Simplified Data Processing
Ryan GarrettJanuary 30, 2025

What is Clustering in an Open Data Lakehouse?
Dipankar MazumdarJanuary 23, 2025

Introducing Onehouse Compute Runtime to Accelerate Lakehouse Workloads Across All Engines
Kyle Weller and Rajesh MahindraJanuary 16, 2025

Accelerating Lakehouse Table Performance - The Complete Guide
Chandra KrishnanJanuary 9, 2025

Product
Unbundling Your Data Platform: How Open Data Lakehouses are Changing the Game
Ryan GarrettJanuary 2, 2025

Top 5 tips for scaling Apache Spark™
Andy WalnerDecember 19, 2024

Amazon S3 Data Lakes: A Complete Guide
Po HongDecember 17, 2024

Data Infrastructure Transformation at Scale: Conductor's Success with Onehouse
Pauline BrownDecember 12, 2024


Comprehensive Data Catalog Comparison
Kyle WellerDecember 10, 2024

AWS re:Invent Recap 2024: AI & Open Table Formats
Andy WalnerDecember 9, 2024

Open Source Data Summit 2024 Draws Data Engineers and Data Architects
Floyd SmithNovember 26, 2024


Product
Conductor Unlocks the Power of Open Data Architectures with Onehouse
Floyd SmithOctober 31, 2024


Product
Open Table Formats and the Open Data Lakehouse, In Perspective
Dipankar MazumdarOctober 7, 2024


Announcements
The Onehouse Data Integration Datasheet: Your Guide to Seamless Data Connectivity
Pauline BrownSeptember 23, 2024

Case Study
Apna Unlocks AI Job Matching for 50 Million Users With Confluent & Onehouse
Floyd SmithSeptember 19, 2024


Announcements
Announcing: AI Vector Embeddings Generator for the Lakehouse
Chandra Krishnan and Ryan GarrettAugust 22, 2024

Case Study
Solving Data Challenges: Olameter and Onehouse Team Up for Enhanced Predictive Maintenance
Pauline BrownAugust 5, 2024
Hudi
Open Data Foundations with Apache XTable: Hudi, Delta, and Iceberg Interoperability
Floyd SmithJuly 8, 2024
Announcements
Raising our Series B and Our Quest For the Most Open and Interoperable Data Platform
Vinoth ChandarJune 26, 2024



Hudi
The Universal Data Lakehouse: User Journeys From the World's Largest Data Lakehouse Builders
Floyd SmithJune 4, 2024

Running Apache Hudi™ on Databricks
Sagar LakshmipathyMay 22, 2024

Announcements
Announcing OneSync™ - Multi-Catalog Sync with Snowflake, Databricks, BigQuery
Kyle WellerMay 20, 2024

Hudi
Optimizing JobTarget’s Data Lake: Migrating to Hudi, a Serverless Architecture, and Templated Glue Jobs
Floyd SmithMay 13, 2024


Hudi
Dremio Lakehouse Analytics with Hudi and Iceberg using XTable
Dipankar Mazumdar & Alex MercedApril 23, 2024










Product
2023: A Breakthrough Year for Onehouse and the Universal Data Lakehouse
Floyd SmithDecember 22, 2023

Solutions
Onehouse Offers First Turnkey End-to-End CDC Ingestion Lakehouse
Kyle WellerNovember 30, 2023








Product
Building a near real-time data lake with Onehouse and Starburst
Kyle Weller and Matt FullerAugust 10, 2023



Solutions
The Ultimate Data Lakehouse for Streaming Data Using Onehouse + Confluent
Andy WalnerJuly 18, 2023

Hudi
Knowing Your Data Partitioning Vices on the Data Lakehouse
Bhavani Sudha SaktheeswaranJuly 12, 2023

Solutions
Integrating Onehouse with the Amazon Sagemaker Machine Learning Ecosystem
Chandra KrishnanJuly 7, 2023

Solutions
Optimize Costs by Migrating ELT from Cloud Data Warehouses to Data Lakehouses
Po HongJune 30, 2023

Hudi
Exploring New Frontiers: How Apache Flink™, Apache Hudi™ and Presto Power New Insights at Scale
Nadine FarahJune 16, 2023




Solutions
Instantly unlock your CDC PostgreSQL data on a data lakehouse using Onehouse
Po HongMay 17, 2023




Product
Onehouse Product Demo - Building a data lake for Github analytics at scale
Po HongApril 21, 2023



Hudi
Getting Started: Manage your Hudi tables with the admin Hudi-CLI tool
Sivabalan NarayananFebruary 22, 2023





Hudi
Comparing Apache Hudi's™ MOR and COW Tables: Use Cases from Uber and Shopee
Andy WalnerDecember 27, 2022




Hudi
Apache Hudi™ vs Delta Lake - Transparent TPC-DS Data Lakehouse Performance Benchmarks
Alexey KudinkinJune 29, 2022
Hudi
Hudi’s Column Stats Index and Data Skipping feature help speed up queries by an orders of magnitude!
Alexey KudinkinJune 9, 2022


Hudi
Introducing Multi-Modal Index for the Lakehouse in Apache Hudi™
Sivabalan Narayanan and Ethan GuoMay 17, 2022





Hudi
Apache Hudi™ Z-Order and Hilbert Space Filling Curves
Alexey Kudinkin and Tao MengDecember 29, 2021


Hudi
Building an ExaByte-level Data Lake Using Apache Hudi™ at ByteDance
Ziyue Guan (translated to English by Ethan Guo)September 1, 2021



