Employee story

The Great Data Connector: Architecting the flow of healthcare data

#BigData #DataInteroperability #DataMesh #RocheCareers #Tech4Life
Phenomblogheader4
Data is only as powerful as our ability to use it. At Roche, we have a wealth of information, but the real challenge is making sure it can be accessed and understood both within a function and across the entire organization. Within Roche Digital and Technology (RDT) - Data Organization, Frank and his team act as the "Great Data Connectors." They are building the digital nervous system that allows scattered data to flow together, ensuring that siloed information becomes the foundation for bringing new medicines to patients faster and more reliably.

In the current healthcare landscape, data is everywhere - hospitals, labs, research platforms, and various functional silos. We see each silo as a different country speaking a different language. My team provides the technology and knowledge to act as the 'Great Data Connector,' translating these languages so data can be brought together faster and in a way that truly makes sense for our scientists, business teams, and AI agents.

“Roche has extraordinary data and extraordinary people. Our role is to connect them through a trusted federated data ecosystem - so insights can move faster across domains, AI agents can act with context, and decisions can translate more quickly into value for patients.”



PhenomQuotes
Today, we are focused on making data AI-enabled and scalable across data platforms. We are moving into a space where insights are generated based on many different data modalities, like structured data, unstructured data, images, etc. We want to ensure that a researcher can access and leverage all these data modalities for insights and decision-making, leveraging AI.

To build this 'digital nervous system,' our day-to-day work involves architecting a future where data moves as fast as our scientists' ideas. Some of the core areas we are evolving to optimize this global data flow include: 
 
  • Converting unstructured data into AI-readable data: We build and expand a 'semantic layer' that provides crucial context to raw data, ensuring it is instantly understood and actionable by both humans and AI.
  • AI-Powered Metadata: We are utilizing AI at scale to automatically create metadata - the descriptions that make complex data easily discoverable and usable for machine learning workflows.
  • Automatic Lineage and Observability: We’ve built systems for automatic lineage to trace the exact origin and journey of our data. Combined with automatic observability, we can monitor and fix issues in our data pipelines in real time. 

We maintain a heavy focus on large-scale data integration, which ensures that external data securely enters our ecosystem while internal data streams in real time to fuel automation. Today, these data assets are discoverable in a centralized marketplace where people and AI agents can find what they need using natural language. 

For me, the success isn't in the complexity of the code, but in the efficiency it creates for the whole company. By using AI to automate our programming, data pipelines, data asset creation and data management, we are scaling faster than ever. We aren't just managing data; we are building a foundation where the path from a raw data point to a life-saving insight is shorter and clearer than it’s ever been.

Frank Rydbirk 
Head of Data Platforms
seperator
Enabling data at this scale is a significant engineering challenge in healthcare, which supports getting medicine to patients in need faster. If you want to build the technology and infrastructure that turn massive datasets and multi-modal data into medical breakthroughs, you belong here. Here, your skills are on the front lines of enabling faster medicine to patients in a trusted way.