MatrixOne Introduction

MatrixOne is a unified data infrastructure for AI agents and intelligent applications. In one database, it connects business tables, JSON, documents, and external files; combines transaction processing, real-time analytics, full-text search, and vector search; and uses Git for Data to provide snapshots, branches, diffs, selective application, merges, and recovery for data changes.

AI agents do more than read data. They continuously retrieve context, generate results, transform data, and may perform operations that affect business state. Traditional data architectures distribute business data, object files, keyword indexes, and vector indexes across separate systems. This creates repeated data movement, complex permission boundaries, and changes made by agents that are difficult to review or reverse. MatrixOne brings these capabilities into a MySQL-compatible SQL data layer, allowing agents to access different forms of data through a unified interface and work in controlled data workspaces.

A Unified Data Layer for AI

MatrixOne organizes different forms of data and their context in one data system:

  • Structured data: Use transactions, SQL queries, and real-time analytics for business data such as orders, users, devices, and metrics.

  • Semi-structured data: Use types such as JSON for events, model output, tool-call results, and dynamic attributes.

  • Unstructured data: Use Stage and DATALINK to connect documents, images, audio, video, datasets, and model files in object storage, remote file systems, or local file systems, while managing their references and metadata in tables.

  • Vector data: Store embeddings produced from text, images, or other content by external models and associate them with source objects, business attributes, and permission information.

MatrixOne stores and queries structured, semi-structured, and vector data, while managing references, metadata, and SQL access for external files. The external file content remains in the underlying object store or file system. Parsing content such as images, audio, and video and generating embeddings are typically performed by external models or AI services. Their results can be written back to MatrixOne for retrieval and business computation.

Unified Retrieval Context for Agents

An agent’s context often combines exact facts, keyword-relevant content, and semantically similar content. MatrixOne supports these retrieval methods in the same SQL data layer:

  • Structured queries: Obtain accurate and current business facts through filters, joins, aggregations, and transactions.

  • Full-text search: Use full-text indexes to retrieve keywords from documents and text columns.

  • Vector search: Use vector types, distance functions, and IVF indexes for semantic similarity search. HNSW indexes are experimental and require experimental_hnsw_index to be enabled. See Vector Search.

  • Hybrid search: Combine structured predicates, full-text matching, and vector similarity to provide richer context for RAG, enterprise search, recommendations, and intelligent question answering.

Applications can access these capabilities through SQL, the MySQL-compatible protocol, and familiar development tools. Model inference and agent orchestration can remain in existing frameworks, while MatrixOne acts as the unified data layer for facts, retrieval context, and persistent execution results.

Git for Data: Safe Data Workspaces for Agents

Git manages code changes; Git for Data manages data changes produced by people and agents. An agent can clean, enrich, classify, or transform data in an isolated branch. Application policies or human reviewers can then inspect the differences and decide which results to accept.

Table branches support Diff, Pick, and Merge directly. A database branch creates an isolated workspace containing multiple tables, but its relevant tables must be diffed, picked, or merged one by one. MatrixOne currently has no database-level DATA BRANCH DIFF, DATA BRANCH PICK, or DATA BRANCH MERGE statement.

Git concept

MatrixOne capability

Role in an agent workflow

tag

Snapshot

Save a known data baseline before an agent begins a task

branch

Data Branch

Create isolated data workspaces for different agents, tasks, or experiments

diff

Data Branch Diff

Review rows inserted, deleted, or updated by an agent

cherry-pick

Data Branch Pick

For tables with an explicit primary key, accept only selected, verified results

merge

Data Branch Merge

Apply approved changes to the target table using a conflict strategy

restore

Snapshot / PITR

Return to a trusted state when an operation does not produce the expected result

A typical workflow is:

  1. Create a snapshot, or ensure that a PITR policy covering the relevant object and time range is configured before the task begins.

  2. Use DATA BRANCH CREATE to create a table-level or database-level workspace for the agent.

  3. Let the agent read data, call models, and write results in the branch without affecting the source data.

  4. For a table branch, use DATA BRANCH DIFF to review row-level changes. For a database branch, inspect each relevant table separately.

  5. For a table with an explicit primary key, use DATA BRANCH PICK to accept selected results. Tables without an explicit primary key do not support Pick. Alternatively, use DATA BRANCH MERGE to merge changes from a table branch into the target table.

  6. Delete branches that are no longer needed. If merged results are not as expected, recover with a snapshot or a PITR policy that was configured before the task and still covers the recovery time.

This pattern gives agent-driven data operations isolation, review, approval, and recovery controls. For the complete concepts, permissions, and SQL operations, see Git for Data and the Git for Data tutorial.

A Foundation for Production AI Workloads

MatrixOne’s AI and agent-oriented capabilities run on a distributed database foundation:

  • Unified transactions and analytics: Use the same data for real-time writes, transactions, complex analytics, and retrieval, reducing cross-system synchronization.

  • MySQL compatibility: Connect applications through familiar SQL, protocols, and ecosystem tools.

  • Disaggregated storage and compute: Scale resources for retrieval, writes, and analytical workloads independently.

  • Multi-account isolation: Separate the data and access boundaries of teams, applications, or agents within one cluster.

  • High availability and recovery: Protect continuously running intelligent applications with distributed logging, snapshots, backups, and PITR.

  • Open deployment choices: Run in local, private-cloud, and public-cloud environments to build AI applications close to their data.

Typical Use Cases

  • Enterprise knowledge bases and RAG: Connect business tables and external documents, then combine full-text, vector, and structured retrieval to provide current, traceable model context.

  • Agent-driven data processing: Let agents clean, classify, enrich, or transform data in isolated branches, then apply the results after diff review and approval.

  • Training data and evaluation-set management: Snapshot and branch datasets, compare versions, and reproduce the exact data state used by an experiment.

  • Multimodal applications: Manage references, metadata, and embeddings for external images, audio, and video to support content search and recommendations.

  • Intelligent business applications: Combine transactional data, real-time analytics, and AI retrieval in one system, reducing the delay caused by moving data into separate databases and search systems.