Healthcare
Rhino FCP

What Does the World’s Largest Decentralized Health Data Science Network Need?

By
The Rhino Team
Daniel Feller, PhD
, 
Director of Data Solutions
October 8, 2026

Ahead of the OHDSI Global Symposium in New Brunswick, N.J., later this month, I wanted a clear picture of what slows down collaborative health data science today. As Director of Data Solutions at Rhino Federated Computing, I work on that question every day, so I went to where practitioners talk about it: the public OHDSI Forums. I analyzed two years of threads and coded each for the technical problem behind it. The result is a ranked map of 29 distinct needs.

Harmonization dominates the map. Two problems recur most: mapping local data to the OMOP Common Data Model without losing meaning, and then proving that two sites actually mean the same thing once they've both mapped. These are the problems the Rhino Data Harmonization Engine (DHE), a feature of our Rhino Data Activation solution, was built to solve.

‍

Key Findings

  • Harmonization accounts for half of all pain points, and vocabulary mapping and ETL conversion alone produce 43% of need statements.
  • Vocabulary gaps and defects form the single largest need at 55 threads.
  • Infrastructure carries the highest severity, with a third of its threads rated as blockers.
  • Federated studies almost always execute properly: only one thread in two years reports one failing. Multi-site threads ask a harder question, which is whether results from different sites can be trusted as equivalent.

‍

Why Analyze an Online Health Data Science Forum? 

When I was a PhD student in 2019, federated learning, or training machine learning models on distributed health data, was an exciting idea that rarely left academic research. Seven years later, hospitals run it in production, and Rhino Federated Computing is the global leader in those deployments, with more than 100 since 2021. The next challenge is making data from different institutions mean the same thing so results across sites can be trusted.

The Observational Health Data Sciences and Informatics (OHDSI) collaborative is the preeminent network of researchers who generate evidence from distributed health data. This community has tremendous insight into the challenges that create friction in the deployment of federated computing across healthcare. The OHDSI community has discussed more than 8,000 topics at forums.ohdsi.org since 2014. These forums capture the difficulty of doing cross-site analyses. So, I analyzed two years of those threads to answer a practical question: what technology does collaborative health data science still need, and where does today's tooling fall short?

‍

1. What Are the Key Challenges?

Each of the 429 substantive threads analyzed from forums.ohdsi.org was tagged with the stage of the OHDSI lifecycle it concerned, from converting source data through vocabulary mapping and data quality to cohort building and analysis. Figure 1 plots each stage by how many threads it attracted and how severe they were. Severity indicates whether a project could proceed. Infrastructure was rated as high severity because a failed installation halts all progress, while challenges with vocabulary mapping only emerge later during analyses. 

Figure 1. Thirteen lifecycle stages are positioned by the number of need threads tagged with the stage (horizontal; a thread may carry several tags) and mean coded severity (vertical; 0 none, 1 minor, 2 major, 3 blocker). Bubble radius encodes the share of the stage’s threads rated blocking.

Two regions matter. On the far right sit vocabulary mapping (127 need threads) and ETL/CDM conversion (114 need threads). Together, they account for 43% of all need statements. Their severity is mostly “major” rather than “blocker” because the work continues around the gap, resulting in lost information rather than a halted pipeline. These two stages read like a problem statement for the Rhino Data Harmonization Engine, which pairs LLM-recommended mappings with expert review. The threads show exactly what reviewers need to catch.

In the upper middle sit infrastructure and deployment (64 need threads) and data quality (60 need threads). Their threads cover installing Atlas and WebAPI, authenticating against LDAP, and running Achilles or the Data Quality Dashboard on large databases. A third of infrastructure threads were blockers, the highest share of any stage. The per-tool breakdown and platform need clusters are in the repository. For anyone building tools for this community, the threads set a clear bar. Users measure every new tool against an Atlas install that can take weeks, so deployment has to be easy before anything else gets a hearing.

Notice what is small. Network study execution, the activity OHDSI is best known for, appears in only 13 threads, and Strategus, the framework that runs network studies, is named merely seven times in two years. People struggle far more to trust results across sites than to run the study.

‍

2. A Taxonomy of 29 Needs

Topic models, which group text automatically by shared vocabulary, are easy to run on raw forum posts but produce results that are hard to interpret. When I tried topic modeling on the forums, the analysis returned topics like "symposium" and "community calls," which say nothing about what people need. So, I took a different route. 

I used a language-model coder to write a one-sentence pain point for each thread that had one, 392 in all, and I clustered those sentences. I then labeled each cluster, wrote a need statement based on the threads in it, and scored it on an Unmet Need Index. The index weighs volume, unresolved share, severity, trend, and unresolved volume, and I published every component so readers can re-weight it.

Figure 2. 29 bottom-up need clusters ranked by Unmet Need Index. Bar color shows the capability area each need maps to after discovery: data harmonization (12 clusters, 194 pain points), federated analytics (1 cluster, 18 pain points), or none, meaning tool bugs, installation, authentication, performance, and community process (16 clusters, 177 pain points). Right-hand labels give thread count and the share not fully resolved on the forum.

Half of all pain points (194 of 389 clustered) fall into 12 harmonization clusters, and three of the top four needs are about getting data into OMOP. Merging the two vocabulary clusters, as the labeling review suggested, makes “vocabulary content gaps and defects” the single largest need at 55 threads. The sections below summarize the technical substance of each group, drawn from the need statements and the threads behind them, along with what each one means for harmonization tooling.

‍

Vocabulary Gaps and Defects (55 threads)

The labeling review merged two vocabulary clusters into this group, the largest in the analysis. Many source codes have no standard concept to map to, including ICD-O-3 site, histology and behavior combinations from tumor registries, ATC vaccine classes, rare-disease codes, DRGs, and compounded drugs. Where concepts do exist, some are wrong. Members report duplicate TNM concepts, deprecations without a replacement, and truncated names. Usagi, the community mapping tool, cannot load a custom vocabulary for fuzzy matching, so a site that builds its own extension gets no tooling help.

Several ICD-O-3 site/histology/behavior code combinations present in our source data do not exist.

Tumor registry conversion, 2025.
‍

How Rhino Helps: Usagi leaves a site on its own once no standard concept exists. The Rhino DHE treats that case as part of the workflow. It proposes candidate concepts for a domain expert to accept, override, or extend, and when no standard concept fits, the expert can map the code to a custom vocabulary. The accepted mapping becomes a reusable artifact, so the next site with the same codes starts from a recorded decision.

‍

Complex Source Data and Silent Patient Loss (33 threads)

This group combines two clusters: mapping complex source data (26 threads) and mappings that silently lose patients (7). The complex cases include combination drugs where strength is known for the product but not per ingredient, clinical-trial instruments, compound survey concepts, family history, and radiotherapy. Guidance for these is unclear or absent, and the mappings that do exist lose granularity.

I don’t have information on how the total 55 mg dose is split between the two ingredients.

Mapping a combination product to RxNorm ingredients discards strength, 2026.

‍

How Rhino Helps: These two threads make the strongest case on the forum for validating a mapping after it runs. A mapping can look correct and still drop 21,667 patients from a cohort. The Rhino DHE closes its workflow with data quality checks and monitoring on the harmonized output, which gives teams a place to catch that kind of loss before an analysis depends on it.

‍

CDM Conventions and Domain Gaps (106 threads)

Three clusters make up this group: CDM conventions (39 threads), oncology and genomics (32), and labs, drugs, notes and pipelines (35). Conventions threads ask where to put derived age, pregnancy episodes, missed visits, patient-reported drugs without dates, dispensed versus prescribed records, and visit-level costs, and the answers members receive conflict. Oncology disease status, imaging findings, biopsies, screening phases, and regimens do not fit the Episode table as specified.

Genomics has no production pattern at all. Members ask how to represent NGS results, germline versus somatic variants, and specimen-level entities, and how to link variants to phenotypes at scale. In the remaining threads, the boundary between Procedure, Measurement, and Observation is ambiguous for tests without results, ATC-to-RxNorm relationships are inconsistent, and common sources lack reusable open-source ETL pipelines.

Most CDMs don’t yet have genomic data.

‍On the feasibility of Mendelian randomization across OMOP sites, 2026.

‍

How Rhino Helps: Mechanically, a conventions question asks which transformation to apply to which column, and a syntactic mapping encodes exactly that. A mapping built once with the Rhino DHE’s no/low-code ETL design can be reused and refreshed across sites, which turns a convention from a forum answer into an executable, shared decision. Oncology and genomics are where I’ve spent time recently. Earlier this year, I used the Rhino DHE to harmonize the public AACR Project GENIE breast cancer release, 18 source files from three cancer centers, into a canonical genomic data model. The same questions the forum asks about variants, specimens, and regimens came up at every step.

‍

3. The Multi-Site Slice

Forty-six substantive threads (11%) involved more than one institution, a network study, or data that could not leave a site. I flagged these conservatively as a thread had to describe multiple sites or partners, not merely use the word “network.”

Figure 3. Each dot is one substantive thread involving more than one site, placed in the quarter it was posted. Filled dots were not fully resolved on the forum; hollow dots were. Multi-site questions arrived steadily across the two years.

Read together, the 35 multi-site threads with a stated pain point make one argument. Nobody asks how to run a study across sites. They ask how to trust the result when they do.

When two sites report a cohort of heart failure patients, they may not mean the same thing.

A call for collaborators on detecting cross-site semantic divergence, 2026.

‍

Any differences could reflect differences in how each dataset was mapped, undermining conclusions.

On validating two independently converted OMOP datasets against each other, 2026.

‍

I want to share the prediction results with external collaborators, I’m not allowed to share patient level tables.

On publishing a prediction model’s outputs without the underlying data, 2025.

‍

The technical tasks fall into three groups.

  • Comparability: Members want to detect when sites encode the same concept differently, validate one site's conversion against another's, stratify results by site, and derive doses consistently when sites populate quantity differently.
  • Sharing without moving patient data: Members want model outputs and aggregate results that collaborators can use without patient-level access, plus phenotype definitions that travel with their implementation.
  • Provenance: Members want to tag which registry each record came from, so data stewards can verify what they hold.

How Rhino Helps: Comparability splits into two problems. The first is transformation drift, and the Rhino Data Harmonization Engine prevents it structurally: every site runs the same mapping, as one shared artifact, behind its own firewall. The second is semantic divergence. Sites that share a mapping can still differ in what their source data meant, which is the problem behind the forum's 2026 call for collaborators. Detecting that divergence without moving patient-level data is the clearest open problem in this analysis.

‍

What I Am Taking to the Symposium

For the OHDSI community, the data argues for two things that no vendor can supply alone: an opinionated conventions reference maintained alongside the CDM and mapping validation that reports hierarchy-unreachable codes before they drop patients from a cohort. The relevant workgroups already know about the genomic and oncology gaps. The forum shows how many implementers are waiting on them.

For Rhino, the forum confirms the real-world use case for the Rhino Data Harmonization Engine. The DHE pairs LLM-recommended mappings with expert review, runs one shared mapping at every site behind its own firewall, and checks the harmonized output for data quality. Our work with cancer research consortiums puts us squarely in the oncology and genomics gaps the forum describes. The hardest problem the forum surfaced, semantic divergence across sites, is still open.

Rhino has a stake in these findings, which is why every table behind these figures, the coding schema and prompts, the analysis code, and the pseudonymized coded dataset are public at github.com/danieljfeller/ohdsi-community-signals. Rerun it, re-weight the Unmet Need Index, or challenge a cluster label, and open an issue where the analysis gets it wrong.

If your team is working on cross-site comparability, mapping validation, or sharing results without moving patient-level data, find me in New Brunswick, October 20–22, or start a conversation with the Rhino team.

‍

How I Did This: Methods and Limitations

‍

Running the Study in Claude Code

Reading 776 threads by hand would take a week, and coding them against a twelve-field schema would take several more. I ran the whole study in Claude Code over a weekend, and I want to be specific about how, because the method is reusable.

  • Collection and cleaning were ordinary code that Claude wrote and I reviewed: a paced crawler against the forum's public API, followed by pseudonymization (salted hashes for handles, with emails and @-mentions scrubbed).
  • Coding was done by Claude Haiku 4.5 against a fixed schema I wrote: post type, lifecycle stage, tools, pain point, job-to-be-done, severity, resolution, multi-site and harmonization context, a verbatim quote under 25 words, and confidence. Claude Code ran it as 18 parallel batches of 45 thread digests.
  • Reliability was checked the way any content analysis checks a human coder. A second, independent coder (Claude Sonnet 5.5) re-coded an 80-thread sample without seeing the first coder’s output.
  • The taxonomy came from clustering the coded pain-point sentences. Claude Opus 5.5 labeled each cluster and mapped it to a capability area with a written rationale, which I reviewed and, in a few cases, merged.
  • The writing, including the figures, was drafted by Claude Fable 5.1 from the resulting tables and edited by me and my colleagues.
Figure 4. Agreement between the primary coder and an independent re-coder on 80 threads, as Cohen’s κ with percent agreement. Shaded bands follow the Landis and Koch convention.

Agreement was substantial or better on most fields this post depends on: whether a thread had a pain point (κ 0.90), whether it concerned harmonization (0.77), post type (0.70), and severity (0.60). It was weak on resolution status (0.36), so I replaced that field with a metadata definition instead of trusting either coder. A thread with no reply from anyone but its author counts as unanswered. A reliability check that only confirms what you hoped isn't doing its job. This one told me which number not to use.

Total model cost at API list prices was under $5 for the coding, clustering, and labeling. The expensive part was the back-and-forth between me and the orchestrating model.

‍

Methods

  • Corpus: I collected all public topics on forums.ohdsi.org created between Oct. 1, 2024, and Oct. 1, 2026, on Oct. 4, 2026. The corpus includes 776 topics, 2,164 posts, and 334 authors, with 429 substantive threads.
  • Exclusions: I also collected GitHub issues but excluded them after a pilot showed that 73% were maintainers' own task tracking.
  • Clustering: I embedded the pain-point sentences (bge-small-en-v1.5) and clustered them with Ward linkage into 32 clusters, then dropped three singletons.
  • Scoring: The Unmet Need Index weights rank-normalized volume at 0.35, unresolved share at 0.25, mean severity at 0.20, positive trend at 0.10, and unresolved volume at 0.10.

‍

Limitations

  • Teams migration: The forum under-represents study coordinators and methodologists who now work in Microsoft Teams, so platform and harmonization needs are probably over-weighted relative to analysis and execution.
  • Survivorship: People who abandoned OHDSI tooling without posting leave no trace.
  • Small counts: Many need clusters have fewer than 10 threads, so readers should treat the lower half of Figure 2 as rough groups instead of a strict ranking.

Full methods, the reliability table, and the per-tool and platform findings are in the repository. All quotes are verbatim, under 25 words, and unattributed. If one is yours and you'd like it removed, contact me.

Build your Federated Network