2026

Broad compatibility, incomplete convergence

Modeled snapshot · through 8 October 2026

Analytical stack in 2026. Modeled snapshot · through 8 October 2026. Three horizontal layers, each totaling 100 percent. Access is at the top, compute in the middle, and storage at the bottom. Curved bands connect access to compute and compute to storage, covering the full edges of the boxes. Vendor APIs stay separate by engine. Access routes after 1991 are illustrative estimates. Select a box to inspect its percentage and emphasize its connections. Within a layer, use the left and right arrow keys to move between boxes.

Widths show shares of analytical activity. Select a box to trace its connections.

All figures are estimates. Did we get something wrong? Let us know.

All percentages for 2026

Methodology

This chart combines documented product history with estimates of how analytical work was done. It is a way to explore changes in the data stack over time. No consistent public dataset measures all three layers across this period, so the percentages are editorial estimates informed by the available evidence.

What the percentages measure

Think of each year as 100 representative units of analytical work, divided among access interfaces, query engines, and storage options. Each layer totals 100%. A share describes that technology’s place in the year’s analytical activity, rather than its share of revenue, customers, installations, bytes stored, or individual queries.

The scope covers SQL reporting, data warehousing, and batch or interactive analysis of persisted data. It includes analytical work on general-purpose databases, but excludes their day-to-day application transactions. Streaming, document and other nonrelational databases, vector search, and BI presentation or semantic layers are outside this view.

A falling share can still accompany growth in absolute use. A missing name may be included in a broader family or be too small to show separately.

How the estimates were built

Release announcements, technical documentation, research papers, and historical accounts establish when products and capabilities became available. Adoption reports, client defaults, and the breadth of each ecosystem help inform the usage estimates. Customer counts, surveys, and vendor revenue are useful context, but cannot be converted directly into shares of analytical work.

The model sets annual estimates and fills gaps between selected years by interpolation. Figures describe a representative late-year configuration. The historical research ends on 8 October 2026, so that year is a snapshot through that date.

A company’s founding, a research prototype, and a shipping product are different milestones. A new product starts with a small allocation when it enters use. Earlier reporting remains in “Legacy reporting,” so the first relational databases do not appear to own the whole market. Release dates are generally better established than the numerical shares, especially in the earliest years.

Access: count the path the client uses

Access means the interface through which a client receives analytical results. Each path is counted once, even when software is layered: a Python client using an ADBC driver belongs to ADBC; one using ODBC belongs to ODBC. “Python DB-API” covers the remaining Python driver paths.

Direct Arrow access includes clients explicitly receiving Arrow results through an Arrow-returning API or Flight/Flight SQL. Arrow used internally by an engine, or conversion to Arrow after rows have been fetched, does not qualify. An ordinary ODBC or JDBC interface stays in its own category even if its driver uses Arrow underneath. ADBC and direct Arrow are separate boxes, so they can be added without counting the same work twice.

Driver availability establishes a possible route, not how often it is used. The estimates allow for adoption through tools such as Power BI and dbt, including drivers users may never configure themselves. ODBC and JDBC are also weighted differently by engine and year: the Microsoft SQL family leans more toward ODBC, while several Java and lake-query ecosystems lean more toward JDBC.

Compute: count the engine or service once

The compute layer follows the engine or service selected for the work. Spark inside Databricks counts as Databricks, and Spark inside Fabric counts as Fabric LH (Spark). Their work is not counted again under Apache Spark. The Apache Spark group includes Spark on Amazon EMR and Synapse Spark; EMR’s Hadoop/Hive, Presto, and Trino work belongs to those respective groups.

SQL Server, Synapse SQL, and Fabric SQL execution share a compute group because of their related execution engines. Fabric WH includes both Warehouse and Lakehouse SQL endpoints. Fabric’s estimated activity is divided equally between SQL and Spark, an explicit assumption in the absence of comparable usage data. SQL Server and Synapse retain separate native storage underneath the combined compute box.

Broader families and “Other analytics” keep smaller products in the picture without giving every product an unsupported individual share.

Storage: count the representation being queried

Storage describes the durable representation serving the analytical work. An engine’s native tables remain native storage whether their files sit on local disks, shared storage, or cloud object storage.

Iceberg, Delta, Hudi, and DuckLake each include their data files and table metadata. Parquet inside an Iceberg table therefore counts only as Iceberg. The standalone Parquet and ORC boxes cover files outside those table formats. A table exposed through another format’s metadata stays with the format that controls its updates and history; compatibility alone does not turn a Delta table into Iceberg.

If data is copied into a native warehouse, subsequent work on that copy counts as native storage. Temporary caches, spill files, and catalog databases do not create extra shares. Work that reads several storage types is divided among them. Queries into another database’s native tables use a separate federated-storage pool where the source vendors cannot be reliably assigned.

How the connections are estimated

Each engine’s analytical activity is divided among the storage options it uses. For example, if an engine has 10% of compute activity and uses Iceberg for 30% of its work, that connection contributes 3% to the full storage layer. Adding the connections from every engine gives each storage box’s width.

Access connections are balanced to match both the interface totals and the engine totals, using supported paths and the client preferences described above. Vendor-specific access is initially divided in proportion to compute share. All connections are estimates of use, rather than a map of every possible integration.

The two sets of connections are estimated separately. Following a path from access through an engine to storage does not establish which interface was used for a particular storage option.

What “managed” means

For Iceberg and Delta, managed means a vendor service handles ongoing table maintenance, such as compaction, optimization, and cleanup. Customer-owned storage can qualify. A hosted catalog alone, or a maintenance job the customer operates, does not. Unmanaged means customer-managed, not unmaintained.

For ADBC, managed means a vendor-operated access service. Simply bundling an ADBC driver inside a client application does not make it managed under this definition.

These are subdivisions of the parent box. Their chart percentages use the full-layer denominator and add to the parent’s share before rounding. Selection and “All percentages” also show each part’s percentage within its parent. The engine connections remain attached to the full box.

Rounding and animation

Extra decimal places keep the arithmetic consistent; they do not imply that the estimates are known that precisely. Box labels are rounded, so displayed numbers may not add up exactly. Very small management subdivisions stay hidden in the chart while their values remain available in the details.

Playback interpolates between annual estimates to make changes easier to follow. Those in-between frames are not additional observations. Pausing settles the chart on the displayed year’s values.

Created with assistance from GPT-6 Astra.

Sources

The research draws on product histories and technical documentation. These sources support dates, architecture, and availability; the usage percentages remain our estimates.