Incremental view maintenance, democratized

Stop recomputing.
Refresh only what changed.

OpenIVM compiles SQL materialized views into incremental SQL. Inserts, updates and deletes propagate as deltas, so a refresh costs what changed, not what exists.

The problem

ETL costs scale with your data. They should scale with your changes.

LLMs and declarative tools like dbt generate more queries than ever, and most of them recompute from scratch on every run. The cost of goods sold for ETL goes through the roof.

Apache Spark is the primary ETL engine behind many lakehouses, yet there is no open standard, and no competitively performant open-source engine, for incremental view maintenance.

OpenIVM is that open standard: an IVM compiler that integrates transparently with the engines you already run.

What an IVM engine needs

Three things, done properly.

01

Rigor

Best-effort, operator-by-operator matching does not scale. OpenIVM derives delta rules from the algebra of each operator (Z-sets, in the DBSP tradition), so a view is maintained exactly, or classified up front as needing a full refresh. A few misplaced CTEs should not throw it off.

Z-sets · delta rules
02

State

Incremental aggregates need somewhere to live between refreshes. OpenIVM keeps view state and auxiliary state as ordinary tables, so it works for batch tables, not only streams, and survives restarts without a separate state store.

aux state · delta tables
03

Deletes

Real pipelines update and delete. Every change is a signed tuple (+1 or −1), consolidated before refresh, so inserts, updates and deletes are handled by the same algebra instead of special cases.

+1 / −1 multiplicities

LLM-readyA rigorous compiler gives agents a fast feedback loop: take any query, rewrite it to be deterministic, verify the incremental plan in CI, and benchmark the savings automatically.

Engines

One compiler, many engines.

DuckDB

available · source build

The reference implementation. Materialized views, automatic refresh, cascading pipelines, adaptive cost model and DuckLake integration. Community extension planned for end of 2026.

github.com/ila/openivm →

Apache Spark

in development

Incremental refresh for Spark ETL and dbt models, including Spark-dialect view bodies such asVERSION AS OF.

Coming soon

Your engine

open standard

OpenIVM emits SQL, not engine internals. If your engine speaks SQL and can store a table, it can host an incremental view.

Propose an integration →

Benchmark

Open, reproducible numbers.

A live benchmark fed by ivm-bench: incremental refresh versus full recompute across engines, workloads and delta sizes. Results are published as a static Parquet file and rendered right in your browser.

coming soon

Research

OpenIVM: a SQL-to-SQL Compiler for Incremental Computations

SIGMOD Companion 2024 · doi:10.1145/3626246.3654743