AI & TechnologyAgentic

Artificial Intelligence Meets the Database in a Rapidly Converging Market

By Joe McKendrick

At the recent OceanBase Hours conference in Singapore, executives from OceanBase made the case for employing a converged, transaction-first database platform to support burgeoning artificial intelligence workloads in enterprises. They reiterated OceanBase’s technology strategy, offering capabilities that parallel those offered by global market leader Databricks, but with a different vision on fusing data with AI and agentic AI.  

The need for a converged data and AI architectural approach is pressing. For today’s businesses, AI and agentic AI are no longer niceties, they are necessities. The challenge is assembling the right data resources and infrastructure that can process AI workloads and deliver information in a fast-moving digital economy.  

Many organizations’ data infrastructures simply aren’t ready for the AI wave. More than one-third of data managers responding to a recent survey of 1,125 technology professionals and executives published by Cockroach Labs believe their data infrastructures cannot handle the workload demands of AI. A majority, 83%, agree that AI demand will exceed the capacity of their data infrastructures in the coming year.  

Driven by this demand, the data market – databases, analytics platforms, and associated tools — has been undergoing the most significant shift since its inception many decades ago. No longer are data environments relegated to the backend of enterprises, with passive storage and limited accessibility. Databases and data environments need to serve as the intelligent layers of AI-driven enterprises, handling streaming data, vector retrieval, retrieval augmented generation (RAG) interfaces, and multiple AI agents.  

Data environments are also required to deliver real time capabilities, the operational mode of the AI era. For decades, analytics was largely delivered via batch mode, with data processed and reports delivered on a schedule. Now, with AI and agentic AI, the data feeding analytic and AI systems, along with the information and insights delivered, need to move at light speed. Data latency, inconsistent versions, or fragmented access controls can lead directly to incorrect AI decisions and business actions.  

That’s why today’s generation of data platforms is integrating AI into every aspect of their pipelines and systems. The goal is to build a single modern architecture that supports mission-critical transactions, real-time analytics, data-lake capabilities, as well as multi-model and multimodal formats. 

Two of data platform providers in this space – OceanBase and Databricks –represent two different paths toward the future AI data platform. Databricks is evolving from a data lake foundation, while OceanBase is evolving from a database foundation, with both approaches contributing to the development of next-generation AI data platforms. 

The symmetry is not accidental. Both platforms are converging on the same destination — a unified data platform for the AI era — but from opposite starting points. Databricks is bringing transactional guarantees to a system built for analytics. OceanBase is bringing analytical and AI capabilities to a system built for transactions. The question is which starting point gives you more to trust when an AI agent is making decisions in real time. 

Databricks has historically been structured around a data lakehouse architecture that leverages data that is captured, stored, and maintained for current and future applications. Transactional capabilities were eventually built into the architecture, via its Lakebase offering, as well as LTAP for real-time, unified data serving. More recently, the company has announced an agentic AI coworker, called Genie One, as well as a developer toolset for building agents called Agent Bricks. The vendor has also rolled out a runtime governance layer for models, agents, tools, and MCPs (Model Context Protocols) called Unity AI Gateway.  

OceanBase’s starting point is not the data lakehouse, but with mission-critical transactions and live operational data. The platform is built on an open-source, distributed native relational database management system architecture designed to provide high availability to production systems. It supports AI and agentic initiatives via heavy transactional and analytical workloads (HTAP) concurrently, vector searches, and high availability. Recently, OceanBase announced that it is expanding from mission-critical transaction processing to include real-time analytics, multimodal data, and AI workloads, aiming for a unified AI data platform.  

Currently, OceanBase holds the largest distributed database market share in China. Its strategic move into AI infrastructure opens it up to the global marketplace for unified, real-time data platforms for enterprise AI. 

OceanBase and Databricks are responding to a number of shifts that AI is creating in today’s data and analytics markets: 

Demand for real-time responsiveness  

Today’s data managers and professionals are actively seeking real-time capabilities built within their tools and platforms, versus being offered as standalone or bolted-on features. Tellingly, interest in standalone real-time analytics has dropped from 50% to 32% over the past two years, a survey of 259 data managers conducted by Unisphere Research finds. The drop in real-time analytics projects “suggests that this capability is increasingly assessed as part of broader data architectural strategies rather than as a standalone research priority,” according to the study’s author, John O’Brien, principal advisor with Radiant Advisors.   

This has far-reaching implications for the development of real-time data platforms. “A copy that is three seconds old, for example, can already be wrong,” according to Indranil Bandyopadhyay, principal analyst with Forrester. “An agent acting on the wrong copy may become a liability rather than a productivity gain. The flow is therefore reversing. Retrieval, inference, and even governance are moving down to the data layer itself because that is the only place where freshness, consistency, and authorization can all be maintained at once.”  

Agents as users  

AI is also spurring a shift in “data gravity,” in which data environments were designed to keep data close to compute resources. “What has changed is the consumer. It is no longer only a human consumer,” Bandyopadhyay said. “We now have AI systems, especially agents, acting in the moment on live data.”  

Importantly, agents are becoming the primary consumers of applications arising from data platforms. “Many of today’s data platforms were designed with a human at the other end,” Bandyopadhyay pointed out. “Humans and programmatic systems queried databases with some tolerance for latency and working hours. Agents have none of that. They query continuously, at machine speed, around the clock. They need structured facts, unstructured context, and relationships in a single authorized step.” 

Data constellations as business entities  

AI models and agents need to understand live business data if they are to deliver accurate results within their proper context. Business entities or events may be drawn from a constellation of data of varying formats or categories. Such an event may be a customer-service interaction, for example, with data related to customer profiles, order records, chat transcripts, call recordings, uploaded images, product manuals, and vector representations. Or AI may call upon a logistics workflow that includes order fields, route events, driver notes, proof-of-delivery images, and semantic address information.  

The leading data platforms handle these data constellations in different ways. With Databricks, it means extending its data lakehouse toward databases and online transactions. Data and AI assets are managed through Delta tables, Unity Catalog, file storage, and Mosaic AI Vector Search. Indexes are built from Delta tables and combined with vector similarity, keyword search, filters, and hybrid retrieval within the platform’s governance framework.  

 OceanBase enables live operational data to serve analytics and AI directly. In doing so, it is extending its foundation in mission-critical transactions and real-time analytics toward data lakes, multimodal data, and AI workloads. This involves a table-centric approach through multimodal tables, which manage structured data, text, images, audio, video, vector, model-generated outputs, and all other forms of data or content. The underlying data does not have to use the same physical storage format, but it can share metadata, access controls, lifecycle policies, and query interfaces. Othe result is all data is contained within the same business entity – such as order status, customer profile, or uploaded image, rather than as separately maintained copies.  

The convergence between databases and AI platforms has been accelerating as AI adoption has gained speed, paving the way toward a single, more manageable architecture. “Databases are absorbing capabilities that used to sit above them, such as native vector retrieval and embedding integrations,” said Bandyopadhyay. “At the same time, context-assembly and AI platforms are reaching down to operational data. Both camps are converging on a similar product surface because they have learned the same lesson: AI is only as good as its access to live, governed, and trustworthy data.” 

By Joe McKendrick

Related Articles

Back to top button