DataCebo Launches SDV 2.0 To Generate Enterprise Databases On Demand

MIT spinoff DataCebo launches SDV 2.0, a generative relational model platform that lets enterprises build synthetic copies of Oracle, SQL Server, BigQuery and Spanner databases for AI training and analytics.

DataCebo Launches SDV 2.0 To Generate Enterprise Databases On Demand

Boston-based MIT spinoff DataCebo has released SDV 2.0, a commercial platform that lets enterprises turn any relational database into a generative model, producing synthetic copies of production data for AI training, analytics and agent evaluation without exposing the original records.

Generative Relational Models, Not Just Row Generators

SDV 2.0 goes beyond earlier synthetic tabular tools by auto-detecting schema, primary and foreign key relationships, formats, constraints and business rules across an entire database, then learning a joint generative model over the graph. Once trained, it emits realistic synthetic data that preserves referential integrity and statistical structure. Native connectors ship for Oracle, SQL Server, BigQuery, Spanner and AlloyDB, and pricing starts at $500 a month with unlimited tables under a consumption model.

"A generative relational model lets them capture that intelligence once and put it to work across the business," CEO Kalyan Veeramachaneni, an MIT Data-to-AI Lab principal research scientist, said in the launch.

Server rack showing enterprise database infrastructure

The Bet: Synthetic Data As An AI Governance Layer

SDV started as an open-source MIT project and now counts more than 18 million downloads, 5,000-plus research paper citations and 30,000-plus data-science users. DataCebo, founded by Veeramachaneni and CPO Neha Patki with backing from Link Ventures, Uncorrelated Ventures and Zetta Venture Partners, positions SDV 2.0 as a governance layer for AI: enterprises can spin up realistic non-production data to train agents and models without regulatory or privacy exposure.

Enterprise AI Turns To Data Foundations

The launch lands the same week OpenText and Cohere unveiled a sovereign agentic AI partnership, Factory closed $200 million for enterprise coding Droids and Cisco extended Splunk AI into private clouds. Each highlights the same shift: agentic AI programs stall not on model quality but on whether teams can safely put enterprise data underneath them.

Reporting based on coverage from AIwire, Solutions Review, NewsBreak and DataCebo's announcement.

Category: Machine Learning

Tags: Machine Learning synthetic biology AI synthetic training data Enterprise Software

Related Articles