Q&A: How Clinical Data Hub from ZS and Databricks will power the next generation of cures

Key takeaways

From genomics and imaging to wearables and real-world evidence, clinical trials generate more data than ever—most of it originating outside traditional electronic data capture (EDC) systems. While this data could be used to accelerate today’s trials and generate novel clinical insights, much of it remains siloed, hard to use and disconnected from the systems that need it.

ZS and Databricks have partnered to create the Clinical Data Hub, a unified, cloud-based platform that brings together structured and unstructured trial data to enable streamlined regulatory submission, exploratory analytics, data reuse and AI-ready workflows across trials and programs. ZS’s Jeffrey Rieske and Databricks’ Christina Busmalis discuss how the Clinical Data Hub can be used to inspire the next generation of therapies.

ZS: Let’s start at the top. What problem in the market did the Clinical Data Hub set out to solve?

Jeffrey Rieske: As long as we’ve had electronic clinical data, we’ve had repositories and places to store it. Historically, though, those have been very siloed. For years, the industry has been trying to figure out how to break down those silos—how to stop making so many copies of data, how to stop sending people to multiple places and how to make it easier to track lineage and prove provenance to regulators. The core idea behind a clinical data hub is about breaking down fragmentation and creating one place where clinical data can live and be used more broadly. On top of this foundational technology, sponsors can deliver on the promise of agentic AI by establishing a new operating model for clinical data—one where teams spend far less time preparing and reconciling data and far more time using it to make decisions.

Christina Busmalis: If you look at how study teams work today, they spend an enormous amount of time just preparing data for use. Within a single trial, data comes from multiple systems, in multiple formats, and a lot of effort goes into just getting it into a usable state. Then, if you zoom out and look across trials, teams are often repeating the same preparation work again and again. That repetition is incredibly inefficient, and it’s been a persistent pain point across the industry.

ZS: This is something the industry has been talking about for decades. What’s changed now that makes it possible to solve this problem?

CB: The cloud is a huge enabler, and that simply wasn’t available in the same way in the past. But it’s not just the cloud. The world has also moved on from having documents and tables that don’t talk to each other. With lakehouse capabilities, you can bring structured and unstructured data together in a single environment. That changes what’s possible.

At the same time, there’s been a movement toward standardization. It’s not perfect—everyone technically has SDTM [study data tabulation model], but then there’s SDTM plus, plus‑plus, and so on. There’s enough alignment now that we can start bringing data together more seamlessly and actually do something meaningful with it, rather than just storing it.

JR: A really important shift is moving from files into data tables. Traditionally, SDTM/ADaM [analysis data model] and even TLFs [tables, listings and figures] are file‑based. When you move those into databases, you can manage access without creating copies. You can allow blinded and unblinded users to look at the same underlying data table, but with different views, and to do that securely and compliantly and in a way that protects trial integrity. That’s a fundamentally different model from what most sponsors are used to, and it unlocks a lot—both for submission and for exploratory use.

ZS: Before we talk about Databricks specifically, I want to stay on that point. What actually changes, and for whom, when you move from files to tables and you no longer have to make copies to enable use in different places?

CB: One of the biggest things is data lineage. When files are scattered across different locations, it’s very hard to know what the master is. You end up with multiple versions of the truth, and that creates risk. Once you move into tables, you can actually track where the data came from, how it’s been transformed and how it’s being used. That alone simplifies the overall data architecture significantly.

It also becomes much easier to integrate data. Things that used to be very manual—pulling data from different files and stitching it together—can be handled much more seamlessly. You can start to use agentic AI or other advanced capabilities to make integration smoother. And ultimately, you get data consistency: Everyone is working from the same underlying data, instead of their own local copies. This eliminates repeated reconciliation cycles, manual handoffs and duplicate validation.

JR: When files move into data tables, you’re able to work faster because you can bring AI to bear more effectively. You’re also reducing risk because you’re not managing multiple uncontrolled copies of the same data. That combination of speed and risk reduction is really powerful.

ZS: Could you give a brief overview of how the shift changes the day-to-day work across a couple of different roles?

JR: Sure. For statisticians, it means moving from file-based analysis to working from governed tables that provide views across users, without the need for copies of data. For data managers, it’s about stewarding a single source of truth rather than coordinating handoffs. And for safety teams, it means gaining near-real-time access to integrated trial data, enabling continuous signal detection.

ZS: You’ve both mentioned AI. Why does AI change the stakes here?

JR: AI effectively becomes a new kind of user, which creates a new imperative: How do I get my data AI‑ready—and, increasingly, agentic AI‑ready? A lot of the work Christina described—data preparation, integration and reconciliation—is still highly manual. There’s real opportunity to automate that work with AI, but only if the data is in a state where AI can actually use it. That means it has to be governed, structured and accessible.

Fragmentation and siloing become even bigger problems once you start thinking about AI at scale. If your data is scattered and inconsistent, you can’t safely or effectively apply AI to it.

author-image-bottom
Fragmentation and siloing become even bigger problems once you start thinking about AI at scale. If your data is scattered and inconsistent, you can’t safely or effectively apply AI to it.
Jeffrey Rieske
Associate Principal, ZS
Testimonial CTA
#
true

CB: The distinction that matters is whether AI is working from your data or from generic models. What sponsors actually need is AI grounded in their trial history, their patient populations, their protocol decisions over time. When data is fragmented across archives, file systems and local copies, AI has nothing coherent to work from. You’re pointing a sophisticated tool at a mess. But when it’s governed, structured and unified, AI can surface patterns across studies that no individual team would have the bandwidth to find on their own. That’s when the answers become genuinely novel. And governance isn’t separate from that—it’s what makes it trustworthy enough to act on.

ZS: And what does AI enable that wasn’t possible before?

JR: AI-assisted SDTM/ADaM generation with traceability, cross-study harmonization and continuous safety signal detection pipelines are now possible. Work that once took months can be completed in weeks.

ZS: With that context, let’s talk about Databricks. What does Databricks bring to the table that’s unique in this environment?

JR: One differentiator is how Databricks applies its lakehouse and cloning technology, which allow us to avoid duplicating data while still supporting different use cases and access needs. Another big differentiator is its Unity Catalog. The out‑of‑the‑box governance it provides is a huge accelerator for clients that want to adopt a clinical data hub because it simultaneously protects the data, enables tighter governance and allows a broader set of users to access data in a controlled way.

CB: I’d also point to the medallion architecture—bronze, silver and gold. Other companies do something similar, but Databricks does it really well with an open architecture. You bring in raw, uncleaned data at the bronze level, you format and standardize it at silver and, by the time you get to gold, it’s actually usable for analytics and decision‑making. It turns something that used to be very technical and abstract into an understandable business concept.

Imagine the ability now for all of your teams to have a conversation with your data—being able to ask a question in plain language and get an answer from your organization’s own trial data—that’s now something we can layer on top of a well-governed foundation. It doesn’t replace the technical workflows, but it puts insight in the hands of people who never had direct access before. And for or those teams using workflows for SAS, R, Python or something else, they don’t have to give up the tools they’re already comfortable with. That significantly lowers the barrier to adoption across the organization.

JR: I’d add also the way Databricks builds contextual metadata around tables, columns and how data is structured. Beyond just asking questions in natural language, the metadata enables richer answers because the system understands the context of the data. That contextual layer is what’s going to power the next generation of agentic AI to power specific use cases across the clinical submission pipeline and beyond.

ZS: Let’s shift to the partnership. From your perspective, what does ZS bring to the table alongside Databricks?

JR: ZS has deep expertise in the clinical domain, not just in technology but in how clinical organizations actually operate. There’s also strong architecture expertise and the ability to bring together large, transformative programs and execute them in a way that clients actually realize value.

Another thing: When I joined ZS, I was surprised by the depth of our adviser network. We have former FDA personnel, CIOs, former heads of biostatistics and experts across the clinical landscape who give us very direct, very raw feedback on our thinking. Adviser input actively shapes the product.

CB: ZS differentiates by owning the transformation end-to-end—from business strategy and operating-model design to platform delivery and regulatory execution. We feel confident in their ability to bridge business, technical and regulatory requirements because we have a lot of experience scaling programs in these highly regulated environments. I’d also emphasize delivery. As AI makes it easier and cheaper to generate ideas and information, delivery is where humans still differentiate. Change management matters. User experience matters. People need to feel heard, and they need to understand the value of changing how they work. The experience built into the Clinical Data Hub has to be intuitive. Users shouldn’t need 200‑page manuals to figure out how to use it.

ZS: Speaking of delivery, how are you driving adoption? What early signals or wins are you seeing?

JR: Safety is emerging as a really strong early use case. Regulators are increasingly interested in earlier access to trial data, and safety data in particular lends itself to near-real‑time analytics. There have been recent examples of regulators working with sponsors to get real‑time access to phase 1 and phase 2 trial data so they can identify safety signals sooner. Databricks technology enables that kind of use case, and it’s driving real adoption.

Beyond safety, we’re seeing a lot of interest in exploratory reuse. There’s huge potential for translation, reverse translation and biomarker identification through cross-study reuse and AI enablement. Why would you run a trial and then let the data sit in an archive forever? Preparing that data to work beyond the initial trial submission—by making it findable, accessible, interoperable and reusable (FAIR)—is going to help create the next generation of treatments and cures.

CB: This is where the distinction between a “hub” and a “repository” matters. A traditional clinical data repository was designed primarily for submission: You lock the database, do the statistical analysis and submit. A hub is about much more than a single trial. It’s about systematically reusing data across trials, learning from past experience and using those insights to design and execute better trials going forward.

JR: And over time, that hub can expand. Some clients are already thinking about bringing in real‑world data, or even synthetic data. It becomes a place where data lives together on a modern platform and where AI can be applied in a much more seamless. Where AI becomes powerful is in accelerating things that are very hard today, like harmonizing data across studies at different points in a program’s life cycle. That kind of reuse is possible now, but it’s slow and manual. The goal of the Clinical Data Hub is to create a foundation that makes that work faster and more systematic over time.

ZS: We’ve talked a lot about early adoption for safety and reuse. Are there other stakeholders you think it’s important to highlight? And what’s the value for them specifically?

JR: First and foremost, clinical data exists to support regulatory submissions. So, it’s all the people responsible for transforming the data and conducting the analyses that support regulatory submissions: Data managers, statisticians, medical monitors and so forth.

Even though the Clinical Data Hub is designed for them, they often become users later in the journey—not because of technology limitations but because submission is a highly regulated, GxP process. And so changing how those teams work requires significant change management, process modernization and, most of all, time. Getting adoption right in that environment is as much a business transformation as it is a technical one.

CB: But once that foundation is in place, the opportunity really does expand across the organization, from clinical into adjacent areas like research—where insights from trials can feed back into earlier discovery and new indications over time. That’s one of the most important shifts we’re seeing. A medical director or clinical ops lead shouldn’t have to open a ticket and wait days to get a question answered about their own trial data. With conversational AI, they can ask in plain language—“How does our dropout rate in this study compare to the last three programs?”—and get an answer drawn from their organization’s actual data. Not a generic benchmark. Not a curated slide. Their data, on demand. When access is that easy, people ask more questions. And when people ask more questions, you catch things earlier.

ZS: To close, what does success ultimately look like?

CB: Success is when data isn’t just something you lock for submission and then forget about. It’s when data is reused systematically across trials and programs to reduce amendments, improve trial design, execution and deliver medicines to the market faster to drive better patient outcomes.

JR: Ultimately, it’s about unlocking the next generation of medicines by putting data to work—safely, securely and at scale.

Add insights to your inbox

We’ll send you content you’ll want to read—and put to use.
Sign me up
/content/zs/en/forms/subscription-preferences
default

Meet our experts

left
white
Eyebrow Text
Button CTA Text
#
primary
default
default
tagList
/content/zs/en/insights

/content/zs/en/insights

zs:topic/data-digital-and-technology