A lakehouse is a converged infrastructure design environment that combines the semantic flexibility of a data lake with the production optimization and delivery capabilities of a data warehouse. Data lakehouses are considered transformational and can serve as the foundational analytic data store for the organization. They are designed to unify the capabilities of data warehouses and data lakes into a single platform to support comprehensive data management and AI lifecycle.
Data ingestion: Collecting data from sources and transferring to the lakehouse, including batch ingestion, CDC, stream ingestion and file transfer.
Persistent storage: Leverages simple object storage and is expected to be in an open table format (OTF) that may be complemented by other data types.
Data catalog: Ability to identify and discover data objects, data governance, security, lineage and metadata management of information associated with data assets to enhance integration, access and utility across an organization.
Data management: The lakehouse must ensure features like capacity planning, backup and disaster recovery are performed. This can be executed by the lakehouse or delegated to other services.
Unified data management: Multimodal data storage, schema flexibility and processing for a broad range of data types.
Converged design architecture: Unifies the architecture and workloads of a data warehouse and data lake on a single platform.
Data sources: Types of information sources that are utilized as inputs, including unstructured, semistructured, structured and streaming data.
Data science/machine learning: Involves the application of predictive and prescriptive analytics methods to extract insights and build models.
Query engine(s): Execute queries from one or more query engines that share the same metadata and physical assets of the lakehouse.
Workload management: Ability to execute different workloads without conflicts and with acceptable availability and performance.
VMware Tanzu Data Intelligence is Broadcom's data intelligence platform for private cloud, built for enterprises aiming to unlock more value from all of their data - structured, semi-structured or unstructured. It includes the capabilities provided by Tanzu Greenplum, Tanzu Data Lake, Tanzu GemFire, Tanzu RabbitMQ, Tanzu for Postgres, Tanzu for MySQL, Tanzu for Valkey and Tanzu Data Flow. The data lakehouse architecture of Tanzu Data Intelligence unifies fragmented data estates, enables real-time insights, powers AI/ML workloads, and business outcomes by enabling queries of all data types. It also provides governance and multi-platform agility for organization’s most critical workloads and AI-powered apps. Tanzu Data Intelligence is purpose-built enterprises looking for data sovereignty, while accelerating their data-driven strategies. Tanzu Data Intelligence, combined with Tanzu Platform, offers a secure, opinionated, and full-stack solution for agentic and generative AI app delivery.
Starburst is the Enterprise Intelligence Platform that connects, organizes, and activates AI on distributed enterprise data — without moving it. Built on enterprise-hardened Trino, Starburst delivers high-performance federated SQL across cloud lakes, warehouses, and operational systems on Apache Iceberg's open table format.
The Enterprise Context Layer organizes distributed data into shared ground truth — serving governed business context, data products, and semantic definitions at query time so humans and AI work from the same trusted foundation.
AIDA, Starburst's AI Data Assistant, connects models, tools, and agents through a governed, model-agnostic agentic layer grounded in certified enterprise data.
Deployed as Starburst Enterprise Platform (SEP) for on-premises/hybrid environments or Galaxy, a fully managed SaaS offering.
Primary use cases: federated analytics, enterprise data products, AI-readiness for regulated industries, and governed agentic workflows.
Dremio provides an agentic lakehouse platform designed to support AI-driven analytics and automation. It enables AI agents and users to access and analyze data across sources through federated query capabilities, unstructured data processing, and an AI-powered semantic layer that adds business context. The platform automates performance management and query optimization, reducing manual administration and supporting scalable, self-managing data operations. Dremio is built on open standards, including Apache Iceberg, Apache Polaris, and Apache Arrow, and is used by global enterprises across industries to accelerate data access and insights.
Microsoft Fabric is a data analytics software that integrates multiple tools for data integration, data engineering, data warehousing, data science, real-time analytics and business intelligence into a single platform. The software supports connectivity across sources and provides a unified experience for data preparation, transformation and modeling. Microsoft Fabric enables organizations to store, manage and analyze data from various sources by offering access to lakehouse architecture, semantic models and reporting features. The software addresses challenges related to disparate data tools and data silos by offering centralized governance, security and collaboration functions aimed at streamlining analytics processes for business decision-making.
Amazon SageMaker AI is a software that enables developers and data scientists to build, train, and deploy machine learning models at scale. The software offers a managed environment that supports various machine learning frameworks and algorithms, including built-in tools for data labeling, model tuning, and data preparation. It provides infrastructure automation for distributed training, as well as model hosting for real-time and batch inference. Users can take advantage of integrated Jupyter notebooks to perform data exploration and preprocessing. Amazon SageMaker AI supports deployment across cloud and edge environments, helping organizations accelerate and standardize machine learning workflows. The software addresses the challenges of operationalizing machine learning by streamlining development and deployment processes.
Google Cloud’s borderless Lakehouse is an open data foundation that transforms fragmented data estates into an active, real-time system of action. Built on open standards like Apache Iceberg and the REST Catalog API, it eliminates lock-in by enabling bidirectional read/write interoperability across engines like BigQuery and Managed Service for Apache Spark, unlocking advanced analytics and AI capabilities. It solves multi-cloud costs and fragile ETL via Cross-Cloud Interconnect and block-level caching, letting teams run high-performance analytics directly on AWS S3 and Azure data and provides bi-directional catalog federation with AWS Glue, Databricks Unity Catalog and Snowflake Horizon Catalog. Designed for AI, its Knowledge Catalog establishes an enterprise context engine to help ground autonomous agents with high precision. It enforces unified table-level governance across clouds, and accelerates BigQuery and Spark workloads to scale analytics and secure AI grounding.
IBM watsonx.data is a hybrid, open data lakehouse that helps organizations easily access, integrate, and analyze structured and unstructured data across hybrid cloud and on-premises environments. It combines data lakes and data warehouses to support enterprise AI, analytics, and real-time workloads, using open engines such as Apache Spark, Cassandra, and Presto, along with real‑time data services like DataStax optimized for price and performance.
IBM watsonx.data prioritizes data governance and security with end-to-end lineage, consistent access control, and open formats/APIs to prevent vendor lock-in. By unifying data sources, enabling flexible analytics, and providing no-code, low-code, and pro-code interfaces, watsonx.data helps break down silos, streamline workflows, and prepare data for reliable AI and analytics.
Data Lake Analytics is a software provided by Alibaba Cloud that offers a serverless interactive analytics service for enterprises to process and analyze large volumes of structured and unstructured data. The software supports SQL-based analysis and integrates with data sources such as Object Storage Service, enabling users to run queries directly without the need for infrastructure management. Data Lake Analytics provides compatibility with multiple data formats and facilitates real-time or batch processing, supporting data warehousing, business intelligence, and reporting requirements. It addresses the business problem of extracting insights from distributed datasets while optimizing resource usage and operational efficiency in data processing environments.
Amazon SageMaker is the center for all your data, analytics, and AI. Its open data foundation is built on Amazon S3 and Apache Iceberg, supporting multiple query engines such as Amazon Redshift, Amazon Athena, and third-party options through federation. Consistent governance and metadata span every workload with SageMaker Catalog, fine-grained access controls, and full lineage so teams can discover, trust, and secure their data in one place. From this foundation, teams work in a single development environment to build ETL pipelines, query data in SQL, and create analyses in serverless notebooks, all accelerated by the built-in data agent. They can also train and deploy ML and foundation models with SageMaker AI (including HyperPod, JumpStart, and MLOps), and build agentic workflows with Amazon Bedrock and AgentCore in the same platform. SageMaker meets developers where they are with remote IDE connectivity and open protocol support so people and agents can use it programmatically.
Oracle Autonomous Data Warehouse is a cloud-based software designed to automate database management and optimize data analytics workloads. The software utilizes machine learning techniques to handle routine tasks such as patching, upgrading, and tuning without human intervention. It provides scalable storage and compute resources tailored for analytical processing and reporting. Oracle Autonomous Data Warehouse supports integration with various business intelligence tools and data sources, enabling organizations to aggregate, analyze, and visualize large volumes of data efficiently. The software addresses business demands for secure data storage, automated performance optimization, and simplified management, helping organizations focus on extracting insights rather than on database maintenance.
Teradata VantageCloud is an analytics software designed to manage and analyze large-scale data across multiple cloud environments. The software provides data integration, advanced analytics, and artificial intelligence capabilities to help organizations process, store, and extract insights from diverse data sources. VantageCloud enables querying, reporting, and machine learning using a unified interface, supporting both structured and unstructured data. It addresses business challenges related to complex data management and analytical workloads by offering scalable performance, workload management, and governance features. The software is designed to facilitate informed decision-making by enabling users to explore and operationalize data across hybrid or multi-cloud architectures.
ClickHouse Cloud is a software designed for cloud-native managed analytics. It enables users to perform real-time data analysis and scalable processing of large volumes of data through column-oriented database technology. The software provides features such as automatic scaling, high availability, and seamless integration with various data sources. It supports fast queries and concurrent data ingestion, aiming to solve the business problem of analyzing complex datasets efficiently in operational and analytical environments. ClickHouse Cloud is intended for organizations seeking to extract insights from their data without managing underlying infrastructure, offering flexibility and performance for business intelligence and data warehousing applications.
Cloudera Open Data Lakehouse is a software that integrates data warehousing and data lake capabilities, enabling organizations to store, manage, and analyze structured and unstructured data from multiple sources in a unified environment. The software supports cloud-native and on-premises deployments, offering features such as data governance, security, and scalability. It provides a consistent data architecture to facilitate data engineering, business intelligence, and machine learning workloads. Cloudera Open Data Lakehouse addresses business challenges related to data silos, analytics accessibility, and regulatory compliance by enabling seamless data sharing and management across the enterprise.
Databricks Data Intelligence Platform is a software designed to unify data, analytics, and artificial intelligence workloads under a single platform. It enables organizations to store, manage, and analyze structured and unstructured data at scale while supporting collaborative data engineering, machine learning, and business intelligence projects. The software provides tools for data warehousing, data lakehouse integration, automated data workflows, and governance capabilities, facilitating secure sharing and discovery of data assets. By streamlining the creation of analytics solutions, Databricks Data Intelligence Platform aids businesses in deriving insights, building machine learning models, and operationalizing data science processes to address complex analytical tasks and inform decision-making.
Dell Data Lakehouse is a software designed to unify data analytics by combining the features of data lakes and data warehouses. The software provides a single platform to store, manage, and analyze structured, semi-structured, and unstructured data. It supports data integration, governance, and security, enabling organizations to manage large volumes of data from different sources. Dell Data Lakehouse allows users to execute analytics and machine learning workloads directly on their data without the need for migration between environments. The software addresses the business problem of fragmented data by providing a centralized repository and analytics platform, which streamlines data workflows and increases operational efficiency for data-driven decision making.
EverFlex AI Data Hub as a Service is a software designed to support data integration, management and analytics for organizations handling complex data environments. The software aggregates diverse data sources into a unified platform, enabling centralized access and governance. It offers features for data ingestion, transformation, cataloging, and lineage tracking, as well as integration with artificial intelligence and machine learning workflows. The software facilitates automation of data operations and supports compliance through data security and privacy controls. Its architecture is designed to scale with organizational data volume and performance requirements. EverFlex AI Data Hub as a Service addresses challenges in data silos, accessibility, and complexity, helping organizations to streamline data-driven processes and enable insights across different business functions.
HPE Data Fabric Software is a data platform software designed to enable organizations to manage, access, and analyze large-scale data across hybrid and multicloud environments. The software provides features such as data storage, data integration, and real-time data streaming, supporting both structured and unstructured data. It offers unified data access, support for various analytics tools, and capabilities for data governance and security. By facilitating the seamless movement and management of data, the software addresses challenges related to managing diverse data types, enabling organizations to derive insights and build data-driven applications while maintaining data consistency and control across distributed infrastructures.
IOMETE is a software that delivers a unified data platform designed to streamline analytics infrastructure and data engineering workflows. It enables organizations to manage, process, and analyze large datasets efficiently by integrating data lake, data warehouse, and ETL capabilities within a single environment. Its features include centralized data management, real-time data processing, support for scalable storage, and compatibility with various data formats and BI tools. The software addresses the business problem of fragmented data systems by providing a cohesive solution for data ingestion, transformation, and access, facilitating collaboration among data teams.
MatrixOne Intelligence is a software designed for data management and analytics, offering features for real-time data processing and analysis. It supports SQL-based querying and enables integration with multiple data sources. The software facilitates transactional and analytical workloads within a single unified platform, helping organizations manage data storage, retrieval, and complex analytical tasks. It aims to assist businesses in deriving insights from structured and unstructured data while ensuring data consistency and concurrency. The software addresses challenges related to big data handling and seeks to streamline information workflows, supporting decision-making processes through efficient data operations.
MongoDB Atlas is a software that provides a managed cloud database service, built on the MongoDB database platform. It offers features such as automated backups, scalability, security controls, and real-time performance monitoring. The software enables users to deploy, operate, and scale databases across major cloud providers, including AWS, Azure, and Google Cloud. MongoDB Atlas integrates with various development frameworks and supports global data distribution, high availability, and data privacy options. The software addresses business requirements for reliable database management, operational efficiency, and uninterrupted data access, serving as a solution for organizations looking to handle structured and unstructured data at scale while reducing infrastructure management overhead.