Search across 704 pages

Try a tool name, category, or "lifetime deal"

Moscone Center, San Francisco, California, United States Past event

Databricks Data + AI Summit 2026 Recap: What Actually Shipped

June 15-18, 2026

TL;DR: Databricks Data + AI Summit 2026 ran June 15-18, 2026, at Moscone Center (North, West, and South) in San Francisco - not “June 14-17” as the ACF date field on this page previously showed, and the venue itself...

TL;DR: Databricks Data + AI Summit 2026 ran June 15-18, 2026, at Moscone Center (North, West, and South) in San Francisco - not “June 14-17” as the ACF date field on this page previously showed, and the venue itself was never named at all before this rewrite. More than 30,000 people attended in person from 150-plus countries, with tens of thousands more watching the keynotes online, on a schedule Databricks itself calls the world’s largest data and AI conference. CEO Ali Ghodsi opened Day 1 with one argument repeated across both keynote days: AI does not have an intelligence problem, it has a context problem. Everything else - Genie One going generally available, a Postgres-on-the-lake product called Lakebase, a genuinely new architecture called LTAP, and a real-time query engine called Reyden - traced back to that single line.

I check AI event pages the same way I check AI SaaS tools before recommending them: primary sources first, never a directory site’s guess made months in advance. Other pages in this batch on my AI events hub had date fields off by a day or two, and at least one pointed to an event that had quietly changed format. This page’s problem was smaller but still real: the date was wrong by a day on each end, and the original content never named the venue at all - it just said “the largest conference dedicated to data, analytics, and AI infrastructure” and left it there. I confirmed the corrected dates and venue against Databricks’ own event site, the Moscone Center’s own venue listing, and independent recaps published after the event closed.

This is a recap, not a “should you register” guide, because Data + AI Summit 2026 already happened over a month ago. Here is the corrected date and venue, what Ghodsi and his co-founders actually announced, what independent data engineers who watched the stream thought of it, and an honest read on whether the summit was worth it for the people it is built for. Every claim below traces to Databricks’ own newsroom and blog posts, the official summit site, or independent recaps from Flexera, Atlan, and Data Engineering Central - and where a number could not be independently confirmed, I say so instead of rounding it into something cleaner than it is.

Key Takeaways

  • The real dates were June 15-18, 2026, not June 14-17. Databricks’ own event site, the Moscone Center’s venue listing, and its travel page all confirm the same four days at Moscone North, West, and South, San Francisco.
  • More than 30,000 people attended in person from 150-plus countries, per Databricks’ own framing and corroborated separately by Flexera’s and Atlan’s independent recaps - with tens of thousands more watching virtually.
  • Ali Ghodsi’s keynote centered on one line: “AI does not have an intelligence problem, it has a context problem.” He organized two days of launches around four challenges: context, cost, control, and choice.
  • Genie One went generally available, alongside a new context layer called Genie Ontology, an expanded Agent Bricks platform, and an open source “meta-harness” called Omnigent, introduced by CTO Matei Zaharia.
  • The two announcements data engineers actually cared about were Lakebase and LTAP - serverless Postgres on open lake storage, and an architecture that lets transactional and analytical workloads read the same copy of data without CDC pipelines.

What Is Databricks Data + AI Summit?

Data + AI Summit is Databricks’ flagship annual conference, the company’s answer to Salesforce’s Dreamforce or AWS’s re:Invent. It is not an academic, peer-reviewed venue like CVPR - there is no submitted-paper track. It is a product and platform event built around two packed keynote days, 800-plus breakout sessions, an expo floor, hands-on labs, and certification training, aimed at data engineers, data scientists, platform architects, ML engineers, and the executives who fund them.

The event grows every year, and 2026 marked at least its second consecutive year at Moscone Center, per independent recaps. Apache Spark, the open source engine Databricks commercializes, now sees more than three billion downloads a year per Ghodsi’s own keynote framing, and that installed base is a big part of why the summit draws data platform teams from PepsiCo, Mastercard, AstraZeneca, Atlassian, and dozens of other named enterprises to talk publicly about production deployments rather than pilots.

If you build or evaluate AI infrastructure - data pipelines, vector search, agent orchestration, governance tooling - Data + AI Summit is where a meaningful share of the year’s roadmap gets announced in public, well before it shows up as a line item on a vendor comparison page like my best AI tools list.

What Happened at Databricks Data + AI Summit 2026: Dates, Venue, and Scale

Data + AI Summit 2026 ran June 15-18, 2026, at Moscone Center - specifically Moscone North, West, and South, 747 Howard Street, San Francisco, CA 94103 - per Databricks’ own event site and the Moscone Center’s own venue listing, which describes the summit as “the world’s leading conference for data engineering, analytics, machine learning, and artificial intelligence.” That corrects the “June 14-17, 2026” date this page’s ACF field carried, a one-day shift on both ends, and adds a venue the original thin content never named.

Quick facts:

DetailWhat I could verify
EventDatabricks Data + AI Summit 2026
DatesJune 15-18, 2026
VenueMoscone Center (North, West, South), 747 Howard Street, San Francisco, CA
OrganizerDatabricks
Attendees30,000+ in person, from 150+ countries; tens of thousands more virtual
Sessions800+ breakout sessions across four days
FormatTwo keynote days (June 15-16) plus two days of breakouts, labs, and certification training (June 17-18)
PricingGeneral admission $1,895 standard; keynote and expo-only pass $195; early-bird discounts of roughly 50% before April 30
Headline keynote lineAli Ghodsi: “AI does not have an intelligence problem. The problem is that AGI is not really permeating our organizations.”
Major launchesGenie One (GA), Genie Ontology, Lakebase updates, LTAP, Lakehouse//RT, Unity AI Gateway, Omnigent (open source)
Sponsors and partners240+, including Microsoft, AWS, Google, OpenAI, Accenture, Deloitte, EY, Infosys, Fivetran, and Salesforce
Official sourceDatabricks’ own Data + AI Summit 2026 site

Those figures are corroborated across sources with no reason to coordinate: Databricks’ own event and travel pages, the Moscone Center’s independent venue listing, and two third-party recaps - Flexera’s FinOps-focused blog and data governance vendor Atlan - both published after the summit closed and both landing on the same “30,000-plus attendees from 150-plus countries” figure. When three parties with different incentives converge on one number, that is about as solid as conference attendance data gets without an audited headcount.

Keynote Recap: Ali Ghodsi’s Context, Cost, Control, and Choice

Ghodsi opened Day 1 with a keynote that ran well over an hour, returning after nearly every announcement to one argument: AI does not have an intelligence problem. It has a context problem.

He made the case with an audience poll. He asked how many people believed AGI had already arrived - about 90% said no. Ghodsi disagreed: he put a genuinely obscure graduate-level math problem on screen (the reduced 12th-dimensional Spin Bordism of the classifying space of the Lie group G2), noted only one person could raise a hand for it, and pointed out that “all the frontier agents and AIs today can solve this.” His conclusion: the models are already smart enough. What is missing is context - the ability for those models to reliably reach the right internal knowledge, at the right permission level, fast enough to be useful inside a real company.

From there, Ghodsi organized the two-day keynote around four challenges: context, cost, control, and choice - and was candid Databricks has not fully solved any of them yet. “I’m not going to say we’ve completely solved it. But at least we’re making some big leaps towards solving these four things,” he said, a more measured framing than most vendor keynotes offer.

Beyond Ghodsi, the stage carried real range. Ryan Blue, creator of Apache Iceberg and co-founder of Tabular (acquired by Databricks), confirmed Iceberg v3 support is now generally available, unifying the physical storage layer between Delta Lake and Iceberg. Reynold Xin, co-founder and chief architect, introduced the Reyden query engine and LTAP. Matei Zaharia, co-founder and CTO, gave the full introduction to Omnigent. Two fireside chats brought in outside voices: OpenAI president Greg Brockman argued that “connecting \[AI\] models to the world is actually becoming this most important critical challenge,” while a pre-recorded conversation between Ghodsi and Microsoft CEO Satya Nadella covered moving past “frontier model worship” toward what Nadella called a “frontier ecosystem” built on proprietary enterprise data and “token capital.”

Enterprise customers got real stage time too. Magesh Bagavathi, PepsiCo’s global chief data and AI officer, described consolidating 60-plus data lakes into one lakehouse over roughly six years at a company with 320,000-plus employees - PepsiCo’s Genie deployment inside its procurement platform logged nearly 30,000 interactions in its first few weeks. Federico Cohen Freue, Mastercard’s EVP of AI and data operations, described consolidating around 80 services onto Lakebase to support a “Virtual C-Suite” initiative, built as an MVP in seven weeks.

Major Announcements: Genie One, Lakebase, LTAP, and Omnigent

Databricks did not make one headline announcement the way NVIDIA’s GTC tends to - it made roughly two dozen across four days. Here is what actually matters, grouped by the problem each one solves.

Genie One and Genie Ontology (agents). Genie One, now generally available on web, iOS, and Android, is Databricks’ rebuilt agentic assistant for business teams. “This is not that Genie. This is Genie One,” Ghodsi said. It connects to 50-plus enterprise apps at launch - Google Drive, Salesforce, Jira, Slack, Confluence - and goes beyond answering questions to produce documents, run scheduled tasks, and take action through MCP integrations. Underneath it sits Genie Ontology, a self-improving context layer that builds a knowledge graph of organizational context using a ranking algorithm called OntoRank, modeled on Google’s PageRank. Ken Wong, Databricks’ senior director of product for Genie, shared internal test data claiming Genie Ontology improved answer accuracy to 84.5% while cutting runtime roughly in half versus leading general-purpose coding agents - a specific, checkable number, though it comes from Databricks’ own benchmark rather than an independent evaluation.

Lakebase and LTAP (data infrastructure). This is the pair the data engineering crowd actually cared about, judging by independent post-event commentary. Lakebase is serverless Postgres built on open lake storage, letting a database scale to zero when idle and branch - a git-style, copy-on-write snapshot - in roughly 500 milliseconds. Databricks says it already handles 12 million database launches a day in production and, new at this summit, added cross-cloud disaster recovery, which the company calls the first fully managed cross-cloud DR for serverless Postgres. LTAP (Lake Transactional/Analytical Processing) goes further: it automatically converts Postgres-native transactional data written into Lakebase directly into columnar Delta Lake and Iceberg formats at write time, so analytical engines query the same copy of data without CDC pipelines or replicas. Xin put it bluntly: “CDC doesn’t stand for change data capture. It really stands for continuous data corruption.” Forbes covered the announcement as Databricks claiming to have cracked a 40-year-old database problem - a bold framing that still has to survive messy production workloads before anyone calls it settled.

Lakehouse//RT and Reyden (real-time analytics). Xin called this “probably the largest single innovation we have done since our introduction of lakehouse.” Lakehouse//RT is a new SQL warehouse type built on a new engine called Reyden, trained by collecting traces from trillions of real production queries rather than academic benchmarks. In preview customer testing, Databricks reported response times as low as 10 milliseconds, roughly 12,000 queries per second of throughput, and up to 16 times better performance than dedicated real-time serving stacks like ClickHouse or Druid - running directly against existing Delta or Iceberg tables with no data copies. It entered beta on June 16.

Unity AI Gateway and Omnigent (governance). Unity AI Gateway is a runtime governance layer for AI spend and agent activity: one entry point for every model and agent request, with spend caps, smart routing between model providers, cross-provider failover, and guardrails against prompt injection, open sourced through Unity Catalog and MLflow. Omnigent, introduced by Zaharia, is a separate Apache 2.0 open source “meta-harness” sitting above individual agent frameworks like Claude Code, Codex, LangGraph, and CrewAI rather than replacing them, letting organizations compose multiple frameworks under centralized governance. A managed version runs on Databricks in beta.

Everything else, briefly. LakeFlow crossed 100-plus connectors with three capabilities reaching GA at once, including a Kafka-compatible ingest API and a real-time streaming mode replacing what used to need separate Flink deployments. Databricks Free Edition - 500,000-plus users - added five previously paid products at no cost, including Genie Code and serverless GPUs. On security, Databricks announced Lakewatch, an agentic SIEM, alongside its intent to acquire Panther, an AI security operations platform founded by Jack Naglieri - its third security acquisition.

Highlights Recap: The Partner Ecosystem and What Practitioners Actually Cared About

Days 3 and 4 (June 17-18) shifted from announcements to implementation: 800-plus breakout sessions, hands-on labs, and 25-plus certification courses covering Lakebase, Genie, and agent development. More than 240 sponsors and partners had a visible presence, including Microsoft, NVIDIA, OpenAI, Anthropic, LangChain, LlamaIndex, Accenture, and Deloitte - a genuinely broad spread rather than a single-vendor trade show.

What’s most useful for judging real substance over marketing volume is how independent data engineers reacted afterward. Daniel Beach, who writes the Data Engineering Central newsletter and watched remotely rather than attending in person, singled out Lakehouse//RT and LTAP as genuinely significant, while giving “a big ho-hum” to the broader wave of agent launches like Genie Ontology and Unity AI Gateway: “the world of Agents and AI is still in flux, no clear winners, everyone is throwing stuff at the wall to see what sticks.” That is a meaningfully different read than the vendor’s own framing - infrastructure landed as more durable with a skeptical practitioner audience than agent tooling did, at least in the first weeks after the summit.

Flexera’s own recap flagged a similar pattern from hallway conversations rather than the keynote stage: attendees repeatedly raised rising compute costs from agentic workloads, agents still sitting in pilots rather than production, and the growing operational complexity of a platform now spanning data engineering, ML, agents, apps, and security under one roof. None of that contradicts the announcements above - it is the honest gap between what got announced on stage and what attendees can actually ship next quarter.

Attendance and Scale: How Big Was Data + AI Summit 2026, Really?

Databricks’ own materials describe the summit as bringing together practitioners from more than 160 countries, and both the company’s framing and the Flexera and Atlan recaps put in-person attendance at more than 30,000, with tens of thousands more watching the keynote livestreams. Ghodsi called it “the largest data and AI conference in the world,” an organizer’s claim rather than an independently audited figure, but consistent with the summit’s scale in prior years and with Apache Spark’s own reported three-billion-plus annual download figure.

I did not find an independent attendance audit confirming the exact 30,000 figure down to the person - that precision essentially never exists at this scale, and any recap claiming otherwise should be read skeptically. What I can say with more confidence: the figure is consistent across Databricks’ own framing and two outside recaps published by parties with no reason to inflate it on the company’s behalf.

One scale detail worth flagging: Databricks reported that roughly 80% of databases created on Lakebase are now built by AI agents rather than human developers - a specific claim from Xin’s keynote that, if accurate, says as much about where enterprise infrastructure demand is coming from in 2026 as any attendee headcount does.

Was Databricks Data + AI Summit 2026 Worth It? My Honest Take

For the audience this event is built for - data engineers, platform architects, and ML teams running production workloads on or adjacent to Databricks - Data + AI Summit 2026 delivered real, checkable substance. Lakebase’s cross-cloud disaster recovery and sub-second branching, Iceberg v3 unifying storage with Delta Lake, and Lakehouse//RT’s benchmarks against ClickHouse-class serving layers are specific, falsifiable claims tied to a shipping or beta product, not vague roadmap language. That is a higher bar than most vendor keynotes clear, and it is why an independent, sometimes-skeptical voice like Data Engineering Central still walked away calling Lakehouse//RT and LTAP genuinely important.

The honest caveat: this is Databricks’ own conference, on Databricks’ own stage, with Databricks’ own internal benchmarks doing a lot of the heavy lifting behind claims like Genie Ontology’s 84.5% accuracy figure or LTAP’s “40-year-old problem solved” framing. None of those numbers are independently audited, and the practitioner reaction on the ground - rising agentic compute costs, agents still stuck in pilots, and a platform spanning data engineering, ML, agents, apps, and security under one increasingly complex umbrella - is a legitimate counterweight to keynote-stage confidence. If you already run workloads on Databricks, this summit’s announcements are worth a serious evaluation queue. If you are shopping agent tooling broadly across vendors, my MCP server leaderboard tracks the vendor-neutral side of that comparison.

Frequently Asked Questions

When did Databricks Data + AI Summit 2026 actually take place?

The summit ran June 15-18, 2026, at Moscone Center (North, West, and South) in San Francisco. This corrects a “June 14-17, 2026” date that previously appeared in this page’s event data and adds the specific venue, which the original content never named.

Who gave the keynotes at Data + AI Summit 2026?

CEO Ali Ghodsi delivered the central keynote across both days, organized around “context, cost, control, and choice.” CTO Matei Zaharia introduced Omnigent, and chief architect Reynold Xin introduced Lakehouse//RT, Reyden, and LTAP. Apache Iceberg creator Ryan Blue joined Ghodsi on stage, and fireside conversations featured OpenAI’s Greg Brockman and Microsoft’s Satya Nadella.

How many people attended Data + AI Summit 2026?

Databricks’ own framing and two independent recaps - Flexera and Atlan - put in-person attendance at more than 30,000 from 150-plus countries, with tens of thousands more watching virtually. No independently audited third-party headcount exists at that level of precision, which is typical for conferences this size.

What did Databricks actually announce at the 2026 summit?

The headline launches were Genie One (GA) and its Genie Ontology context layer, an expanded Agent Bricks platform, the open source Omnigent meta-harness, Unity AI Gateway, and data infrastructure launches - Lakebase updates, the new LTAP architecture, and Lakehouse//RT powered by the Reyden query engine.

Is Data + AI Summit only relevant if you already use Databricks?

Not entirely - the OpenAI and Microsoft conversations, the Omnigent open source project, and the broader argument about agents needing context rather than more raw model intelligence apply to anyone building agentic systems. But the deepest, most checkable announcements are Databricks-ecosystem-specific, so practical value scales with how committed your organization already is to that platform.

The Bottom Line

Databricks Data + AI Summit 2026 ran June 15-18, 2026, at Moscone Center in San Francisco - not June 14-17 as this page previously stated, with a specific venue the original content never named - and drew more than 30,000 in-person attendees, making it, by every source I could independently check, genuinely the scale event Databricks claims it to be. Ghodsi’s “context problem, not an intelligence problem” framing held together roughly two dozen product announcements, and the two that data engineers outside Databricks’ own marketing machine flagged as genuinely significant - Lakehouse//RT and LTAP - tackle a real, decades-old architecture problem rather than repackaging an existing feature with new branding.

The broader lesson from this batch of event-page corrections keeps repeating: a page can carry the wrong date, or no venue at all, for months before anyone checks it against a primary source - small enough to look trivial until a reader tries to plan a trip around it. Every date, attendance figure, and product claim above traces to Databricks’ own event site and newsroom, the Moscone Center’s venue listing, or independent recaps from Flexera, Atlan, and Data Engineering Central.

If lakehouse-native infrastructure, agent governance, or the Databricks-Microsoft-OpenAI partnership are relevant to your stack, the Lakebase, LTAP, and Lakehouse//RT product pages are worth a look now that they are shipping in beta or GA, not roadmap slides. For the rest of the AI events calendar, including which recurring conferences are still real and which quietly changed format, check my AI events hub. Want a heads-up the moment new AI event coverage and deal alerts go live? Subscribe here.

← Back to all AI events