Why OceanBase Built Lakebase

On June 29, OceanBase launched Lakebase as the core engine of its AI database family. Product GM Yanran (Han Fusheng) explains why enterprises need a unified data lake and database foundation for multimodal data, hybrid search, and agent-ready access.

Animated banner introducing why OceanBase built Lakebase for AI-era data

Editor’s note

On June 29, the OceanBase Hours online launch introduced OceanBase’s AI database for integrated data lake and database workloads, marking the debut of the OceanBase AI product family.

In that lineup, OceanBase Lakebase is the core engine of the OceanBase AI database. It combines data lake, database, and multimodal capabilities so structured, unstructured, and vector data can be managed, processed, retrieved, and accessed within one architecture.

Why build OceanBase Lakebase? What is the underlying technical logic? Which business use cases does it support? This article explores those questions.

Author | OceanBase Product GM Yanran (Han Fusheng)

Portrait of OceanBase product GM Han Fusheng speaking about Lakebase

OceanBase Product GM Han Fusheng

AI is rewriting the rules of enterprise data systems.

For decades, a database’s core job was to manage structured data. Transactions, orders, accounts, and finance were organized and queried as tables, and they powered the most critical business systems.

As AI advanced, text, images, audio, video, and other multimodal data entered those systems. They stopped being mere attachments and became assets that can be understood, analyzed, and used to support new business use cases.

Enterprises are not short of data. Data lakes store raw data, databases support transactions, and warehouses serve analytics. In AI use cases, that fragmented architecture struggles to process and understand multimodal data in a unified way. For AI applications to understand the business better, they need a new foundation that connects structured and multimodal data and applies AI capabilities to process and extract value from both.

That is why we released OceanBase Lakebase.

Diagram of OceanBase Lakebase unifying structured and multimodal AI data

A unified data lake and database foundation for AI

We define OceanBase Lakebase as a unified data lake and database foundation for AI business use cases.

It is not a new data lake, and it is not a lateral expansion of database features. It is an attempt, in the AI era, to rethink how enterprise data should be stored, managed, computed, and searched.

The core logic of OceanBase Lakebase is straightforward: manage multimodal data with the same rigor as structured data.

In the past, unstructured data could be stored, but it was hard to put to real use. Documents, images, video, and audio were scattered across systems, without unified metadata, unified indexes, unified compute, or unified search. Many enterprises already have large volumes of high-value data, yet business users and AI applications cannot use it efficiently.

That is the problem OceanBase Lakebase is meant to solve.

Architecture showing Lakebase ingest, compute, search, and agent access

First, text, images, audio, and video can be ingested and processed in one place. For AI applications, that means more previously dormant data can be brought back into use.

To get there, we chose an integrated data lake and database architecture. The openness of the lake matters because AI workloads need massive volumes of diverse data in open formats. The management capabilities of the database matter just as much because enterprise applications need stability, governance, access control, and reliable data services. We want those two sets of capabilities truly combined—not merely stitched together at the interface layer.

That architecture also enables more open and diverse ways to use data. AI workloads vary widely and cannot depend on a single computing model. OceanBase Lakebase therefore supports SQL, Spark, Daft, and other computing paradigms, allowing data engineering, algorithm development, and business analysis teams to work with data in the way that suits them best.

Search also needs to converge. Users need keyword search, vector search, and precise filters on structured fields. We want those capabilities unified so people can find what they need by semantics, by keywords, and by business conditions at the same time.

Looking ahead, people will not be the only consumers of data. More and more agents will access it continuously. In the agent era, data is not only queried by humans; it is also consumed by agents. Agents need more than a knowledge base: they need real-time context, long-term memory, business state, action records, and isolated data environments that support rollback.

What OceanBase Lakebase aims to do is turn these data capabilities into infrastructure that AI applications can call reliably—with standardized interfaces and tools that make it easier for AI agents to understand and use enterprise data.

Diagram of Lakebase standalone and overlay deployment models

Two deployment models

As an enterprise data foundation, we know many customers already run long-lived data systems that hold large historical assets. OceanBase Lakebase is therefore not designed to force a rip-and-replace, and it does not require every dataset to be migrated in before it can be used.

We designed two deployment models for OceanBase Lakebase:

  • Standalone mode, for net-new business use cases. When users launch a new AI application, they can deploy a complete end-to-end stack. OceanBase Lakebase can bring the system online quickly with a relatively small initial footprint and provide all the capabilities the new use case needs, including storage and compute.
  • Intelligent overlay mode, for use cases that need to reuse existing storage and data assets. If a customer has already accumulated a large data lake, OceanBase Lakebase can run alongside the current systems, connect existing data with newly managed data, and present a consistent access layer to applications.

In short, new applications can be built quickly, and existing systems can be enhanced smoothly. Users can choose the model that best fits their use case.

Illustration of Lakebase overlaying existing enterprise data systems

Intelligent driving: find the valuable clips

Intelligent-driving companies collect large volumes of video, images, sensor data, and GPS from engineering vehicles and test vehicles every day. The real problem is not whether the data can be stored, but how to find the valuable fragments in that flood.

Extreme conditions, collision risk, abnormal roads, and severe weather matter a great deal for model training. If the work depends on manual screening, the pipeline becomes very long and inefficient.

In this scenario, the core value of OceanBase Lakebase is to make data storable, computable, and usable at a cost teams can afford.

What it does here is turn video and multimodal data into assets that can be processed and searched. The system supports video splitting, event slicing, keyframe extraction, scene recognition, and feature vectorization, then combines vector search, structured query, and multimodal search so teams can quickly find the samples they need in massive driving data.

For intelligent-driving companies, OceanBase Lakebase is not only a storage system. It is a data foundation for continuous model iteration. It helps customers turn large volumes of driving data into training and test samples, lower data-preparation cost, and speed up model iteration.

Intelligent-driving pipeline turning video into searchable training samples

Securities: unify structured and unstructured research data

Securities firms are not short of data—quite the opposite. They have abundant structured data such as market quotes, trades, financials, and customer records, as well as unstructured data such as research reports, announcements, policy documents, news, and public sentiment. The real challenge is the variety of data types, the difficulty of processing them, and inefficient integration.

In this scenario, OceanBase Lakebase can serve as a processing and serving hub for many data types. It can connect heterogeneous sources in a unified way, then parse, semantically understand, and extract content from research reports, announcements, and policy documents, and build indexes.

In intelligent report parsing, for example, the system can automatically parse research notes, industry reports, and company studies, extracting titles, abstracts, tags, industries, security identifiers, and research institutions so investment materials are easier to retain and reuse. In policy-document management, it can classify, identify, summarize, and tag regulatory rules, compliance files, and anti-money-laundering documents to support subsequent policy retrieval and compliance Q&A.

Securities workflow parsing research reports and compliance documents

From an integrated database to an AI data lake and database

A transactional database alone cannot cover multimodal processing. A data lake alone struggles to meet enterprise-grade management, search, and service requirements.

The AI era needs a new foundation that unifies the data lake and database: one that can manage structured data, process multimodal data, and provide analytical capabilities for people as well as programmable data services for agents.

OceanBase already has a complete database product line that keeps many customers’ core systems running. Building on that foundation, we developed an AI database with OceanBase Lakebase at its core to address emerging AI-era data-processing needs.

As we evolve from an integrated database to an AI data lake and database, we want to continue supporting the stable operation of core enterprise systems while accelerating the deployment of new AI applications.

Lakebase is a step OceanBase is taking in that direction. Through Lakebase, we hope to help enterprises consolidate, manage, process, and search scattered data, ultimately supporting AI applications and business innovation.

Many fellow practitioners will travel this path, and many engineering problems still need to be solved. The direction is clear: future enterprise data infrastructure must understand both structured logic and unstructured semantics.

OceanBase wants to strengthen that foundation alongside its customers and turn these use cases into production realities.

Closing visual of OceanBase Lakebase as the AI data foundation

Related-article card recommending earlier OceanBase community content

Related-article card highlighting another OceanBase technical post

Related-article card pointing to a prior OceanBase product story

Related-content banner repeated from the WeChat article footer

Event recommendation

Event poster recommending an OceanBase community activity

Learn more

Welcome to join the open-source community Discord.

Welcome to join the open-source community Discord