Rhino FCP

Can We Use This Data? A Plain-English Guide to Secure Data Sharing

By
The Rhino Team
Tom Heys
, 
Director, Product Marketing
September 30, 2026

For the leaders who need to approve, fund, or explain a data collaboration — not build one.

If you work with sensitive data, you’ve probably had a version of this conversation:

  • A hospital has decades of patient records and a research team capable of building something genuinely useful with them, but the model would be far better if it learned from 20 hospitals instead of one, and those records can’t leave any single hospital.
  • A bank can see fraud patterns in its own transactions, but the patterns that matter most cross institutions, and sharing customer data with a competitor is not on the table.
  • A pharmaceutical company needs more experimental data than it can access to answer the question in front of it, and every organization holding the rest is either a rival or a partner bound by its own obligations.

The projects with the highest payoff tend to face the most barriers to approval.

Your IT team has not failed to solve this. For most of the history of data analysis, two options existed: get permission to bring the data together, or accept that the work could not be done. Governance, contracts, and careful legal structuring were the only levers available, and in some situations none of them were enough.

Of course, technologists have not been sitting on their hands. A range of technical approaches, known collectively as privacy-enhancing technologies (PETs), now sits between those two poles. Some run a computation without the data ever becoming readable to whoever is running it. Others create controlled conditions under which data can move. They are not interchangeable, and which ones are open to you depends on facts about your situation rather than on preference.

This guide covers that landscape in plain terms: why data gets stuck, what your existing protections actually accomplish, what the available approaches do, who needs to be part of the decision, and what to ask the vendors selling these approaches.

The Four Reasons Data Can't Move

Almost every stalled data project is stuck for one of four reasons, and the reasons look similar from the outside. Misreading which one you have is how a project ends up in the wrong conversation for months.

  1. It isn't lawful. Data residency rules (such as China's PIPL or India's payment data localization mandate), cross-border transfer restrictions (such as the GDPR's limits on moving personal data outside the EU), banking secrecy statutes (as in Switzerland), and sector-specific regulation (such as HIPAA in US healthcare) restrict data movement. In some situations, no lawful route exists at all, and no contract or safeguard changes that.
  2. The custodian won't permit it. The transfer may be perfectly lawful, and yet the organization holding the data still declines because its governance committee, research office, or risk function has a policy that says no. The constraint is institutional, which makes it sometimes negotiable and sometimes not.
  3. The data is the asset. You could move it and are probably allowed to, but the data creates enormous value for your organization. Handing it to a partner who may also be a competitor gives away core intellectual property.
  4. Pooling it creates a new risk. Even a fully lawful, well-governed centralization concentrates exposure that did not exist before. One large repository is a more attractive target than 20 separate ones, and the consequences of a single breach or failure are larger.

The practical value of separating these is that each has a different owner. The first needs a lawyer. The second needs a governance conversation with the custodian, early. The third needs a commercial structure. The fourth needs a security assessment.

Many data collaborations spend months arguing about the first reason while actually being blocked by the second or third.

The Basics: Data Encryption and Anonymization

Your organization almost certainly uses at least one of the techniques below. All of them answer a core question: "How do we protect the data itself?” Understanding these baseline data protection practices forms the foundation for business stakeholders who need to join the technical discussion to deliver on their goals.

Encryption, at rest and in transit, is a locked filing cabinet and an armored van respectively. It protects data while it is sitting still and while it is being moved. But to use data — to run an analysis, train a model, or execute a query — it must be unlocked. At that moment, the system doing the work has the real data in front of it. Security teams call this third state data in use, and most conventional protections leave it exposed.

Anonymization transforms data so that no one can link it back to an individual by any reasonably available means. Data that clears that bar is no longer personal data, so it falls outside GDPR regulation entirely. Regulators set the bar deliberately high; rich, high-dimensional data rarely clears it: a genomic sequence, a medical image, or a long individual history is not meaningfully anonymous just because the name field was removed.

Tokenization replaces a sensitive value with a meaningless stand-in (a token) and keeps the mapping in a separate vault. It works well at shrinking the risk footprint of systems that touch regulated data. The catch is what happens when you want to analyze it: either you work with the meaningless token, which is useless for many modeling and analytics tasks, or you unlock it, at which point the data is exposed. Healthcare and life sciences have a specific variant, linkage tokens, that let two organizations recognize the same individual in both datasets without exchanging names. However, the consistency that makes matching possible also makes re-identification more possible for anyone holding the key.

Pseudonymization removes or replaces direct identifiers. The GDPR defines it explicitly as processing personal data so it can no longer be attributed to a specific person without additional information kept separately. Business stakeholders often don’t realize that pseudonymized data generally remains personal data under the GDPR. You have lowered your risk, but you have not sidestepped regulatory requirements.

De-identification is the term you will hear often in a healthcare context. In the United States, the HIPAA Privacy Rule sets out two routes. Safe Harbor means removing eighteen specified categories of personal identifiers. It’s mechanical and inexpensive, but destructive of analytical value, since it takes out precise dates, most geographic detail, and exact ages over 89. Expert Determination means a qualified expert in statistical methods certifies that the risk of re-identification is very small, which preserves far more of the data's usefulness. The trade-off is that Expert Determination is a professional judgment. A determination does not transfer cleanly between organizations, and each participant's expert in a multi-party collaboration may reach a different conclusion about largely concordant datasets.

These techniques all leave the previously mentioned third blocker untouched: the data as a competitive asset. A de-identified claims file or imaging cohort is just as useful to a competitor as the original. Importantly, these techniques cannot protect the model you built from the data or the question you asked of it, which in a sensitive data collaboration are  the value drivers.

One pattern runs through everything above: each technique changes the data itself. Each one reduces the chance that a record can be traced to a person, and none of them changes where the computation happens. Somewhere, some system still processes the data in readable form. Where that work happens is the other lever you can pull, and one question decides how.

The One Question That Determines Everything

Can your data leave its source?

If the answer is yes, under controlled conditions, you can bring the data into a governed environment and protect it there. If the answer is no, the computation has to travel to the data instead.

Three things are worth knowing about this question.

  1. The answer is not a technology preference. It reflects a fact about your situation, and the people who settle it are your legal team and the organization holding the data rather than IT or a vendor.
  2. The answer can differ across datasets in the same organization, and it often does. Ask per dataset rather than once.
  3. "No" is frequently "no for now." Regulations change, and so do institutional policies. However, designing on the assumption that the answer will soften is how projects end up waiting indefinitely for a policy change that may never arrive.

The four blockers from earlier map onto this question. A legal or custodian blocker usually defers to the answer of "no," which points to approaches that keep data in place. Concentration risk points the same way, since leaving data at each source avoids building one large target. When the data itself is the competitive asset, the answer may be "yes," but you still need an approach where partners only ever see summary information derived from the data, and where your model and your queries stay protected from other parties too. Your answer to the big question determines which of the next two sections applies to you.

When the Data Can Move Under Controls

Two approaches apply when transfer is lawful and the data owner permits it, but you still want controls on what happens once the data moves.

Trusted research environments (TREs). TREs work like a secure reading room. Data is brought in, approved researchers work with it inside the environment, and nothing leaves without review. Excellent when one organization holds the data and wants to give many researchers controlled access to it, which is the model behind a great deal of academic and national health research infrastructure. Honest limitations: The data had to move, which rules this out where transfer is unlawful or where several organizations each need to remain custodian of their own records; and what researchers work with is a copy, which goes out of date.

Data clean rooms. Two or more parties contribute data to a neutral environment, run queries against the combined set, and receive only aggregate results. Familiar to work with, fast, and low effort to set up. Honest limitations: In many clean rooms, the data has moved into someone else's environment, and the workloads supported are mostly queries and reporting rather than model development.

When the Data Cannot Move

Four approaches keep data where it lives and allow it to be processed in a highly secure manner.

Federated computing runs data processing where the data already lives, rather than moving the data to a new location for computation. Only summaries, such as statistics or model updates, exit — never the raw records. Federated computing is the broadest of the options in this group since it supports training models, running analytics, making predictions, and preparing or standardizing data across sites. Federated computing is also the most flexible because it is an architectural approach that companies can use together with the additional protections below, as well as the techniques above that transform data to protect it. Setup follows the same pattern at every participating site: software runs locally alongside the data, and each site's data is mapped to a common format so joint analysis can run across sites. Platforms built for federated work provide tooling for both steps.

Confidential computing, which relies on trusted execution environments (TEEs), acts like a sealed room inside the computer's processor. Work happens on real data at near-native speed, but nothing outside the sealed room can see into it, including the operating system and the cloud provider running the hardware. You can also verify cryptographically which program runs inside the enclave before you send anything to it. A TEE can run on your own servers, in which case the data never leaves. If the enclave sits in a cloud or partner environment instead, the data has to travel to it, which matters when that environment is in another jurisdiction. Either way, the cryptographic guarantee depends on trusting the chip manufacturer's implementation.

Multi-party computation (MPC) allows several organizations jointly compute an answer without any of them seeing the others' input data. This technique is best suited to specific, bounded questions, such as totals, counts, thresholds, or "does this pattern appear in both of our datasets." It is substantially slower than ordinary computing, which makes it a good fit for targeted questions and a poor one for training large models.

Homomorphic encryption (HE) involves computing on data that stays encrypted the entire time using advanced math. It provides one of the strongest guarantee available because the data is never readable by anyone but its owner. That said, it has historically been far too slow to be practical for most purposes. Dedicated hardware is closing that gap. In early 2026, Intel demonstrated a prototype of a purpose-built chip running certain encrypted operations roughly 1,000 to 5,500 times faster than a standard server processor. However, it remains best suited to narrow, well-defined tasks rather than general-purpose work.

Two Additional Privacy Techniques

Differential privacy. Deliberately adding a small, carefully measured amount of statistical noise to results that limits how much any one person's record can affect the result. It comes with a cost, however: more privacy means less precision, and less precision means less useful answers. Someone in your organization has to evaluate how much precision is required to solve the problem to a satisfactory degree, and that choice is a business decision as much as a technical one.

Synthetic data. An artificial dataset built to behave statistically like the real one, designed not to contain real records. It is best for testing, development, and getting work moving with lighter governance overhead. Its greatest weakness results from how it is generated since it learns what is typical and under-represents what is rare. Under-representation is a problem precisely when the rare cases are the point, as in fraud detection, rare disease research, and many other business outputs.

Who Needs to Participate

Data collaborations fail more often on process than on technology. The most common failure is running the evaluation as an IT purchase and failing to bring in legal or the data custodian until the end, where they say no.

The people below need to be involved early, and each of them owns a question:

  • The data owner or custodian: Will they permit it, regardless of what the law allows? They are frequently the real blocker, and the last person invited.
  • Legal and privacy: Can this data lawfully move, on what basis, and into which jurisdictions? They own the answer to the question at the center of this guide.
  • Information security: What specifically are we defending against, and does this approach address that threat rather than a different one?
  • Data science and analytics: What computation do we actually need? "Train a model," "run a query," and "let researchers explore" lead to genuinely different data protection needs.
  • The commercial owner: If the constraint is competitive rather than legal, this is the only person who can restructure the arrangement.
  • IT and infrastructure: What has to be installed, where, and who maintains it once the project is live?

If you take one thing from this section, let it be this. Talk to the data custodian first, rather than waiting to loop them in.

What to Ask a Vendor

These questions will tell you more than any product demonstration, and they help you contribute to solution design alongside your technical colleagues. A software and technology provider with established practices and an experienced team should be able to answer each question in plain language. If they can't, take note: a vendor who can't explain their approach to you is unlikely to be a good partner in reaching your business goals, even if your technical team is on board.

Each question below also carries a concern from someone on the team you just assembled. Asking them yourself brings every partner's priority into the conversation in terms a vendor cannot sidestep, and you hear the answers firsthand.

"At what point in this process is our data readable in the clear, and by which systems?"

Your information security team's first question.

A good answer is specific, concrete, and names where, when, and under whose control. A red flag is "it's encrypted end to end" with no further detail. Almost every product encrypts data at rest and in transit. It’s table stakes and not what you asked.

"Does anything leave our environment, including model updates, logs, metadata, or a temporary copy in an enclave? Where does it go?"

For your legal and privacy team, since jurisdiction follows every outbound flow.

A good answer is an inventory of every outbound flow, what each contains, and which jurisdiction it lands in. A red flag is an answer about raw records only. Model updates, error logs, and schema information can carry personal information too, and they cross borders just as easily.

"Is the privacy guarantee enforced by the technology, or by policy and contract?"

For legal and security together, since each owns one half of the answer.

Contracts, data use agreements, and institutional policies determine what people are permitted to do. Technology determines what is possible. Both are legitimate, and plenty of good systems rely on governance, but an organization can have airtight agreements in place and still expose raw data to every system in a compute environment the moment processing begins.

A good answer shows the vendor knows which one they are selling you and keeps the two distinct. A red flag is presenting governance controls as though they were technical ones.

"Which protections can you combine on a single project, and can the combination differ by dataset?"

For your data science team, who will weigh what each added layer costs in precision and speed.

A good answer explains how the layers stack: keeping data at its source, running computation inside a sealed enclave, adding measured noise to results, and de-identifying inputs where required. The vendor should be able to match layers to each dataset's constraints, since the answer to "can this data leave?" often differs across datasets in the same organization. A red flag is one technique presented as the whole answer. A vendor with a single tool tends to describe every problem as the one that tool solves.

"Could someone reconstruct individual records from the results or the trained model?”

For your privacy team, who carry the risk if results reveal more than intended.

A good answer names specific safeguards: minimum group sizes below which results are suppressed, review of outputs before release, and differential privacy where precision allows. A red flag is "only results leave," as though results were automatically safe. Results about small groups, or models trained on few records, can reveal more than intended.

"What does each organization holding data need to run, and what access does it grant?"

For IT and infrastructure, and for the security team at every site that holds data.

A good answer is a list a hospital or bank security team would recognize: what software runs on site, whether it needs only outbound network connections, what administrative rights it requires, and who maintains it. A red flag is "lightweight deployment" with no detail. Each data holder's security team will ask these questions, and a vendor who has answered them before will have the answers ready.

"Who approves each analysis before it runs against our data, and can we see a record afterward?"

For the data custodian, who needs control over every analysis as well as the agreement.

A good answer: the data holder has the ability to review code or queries before execution, can revoke access at any time, and receives a log of everything that ran. A red flag: approval happens once, at contract signing. After that, a partner or the vendor can run new work against your data without your seeing it first.

"Which organizations holding data like ours have approved this through their own security and privacy review, and may we speak to them?"

For the data custodian, who will want to hear from peers who faced the same review.

A good answer is named references from the side that holds the data, such as a hospital, a bank, or a research institution, running in production. A red flag is references only from organizations buying analysis, never from ones supplying data, or a pilot described in language that makes it sound like production. Buyers experience the demo; data holders experience the scrutiny.

"Who owns the trained model and the results, and who else can see our work?”

For the commercial owner, who needs to know the arrangement protects the asset.

A good answer names who owns each output, spells out what other participants and the vendor itself can see, and explains how models and code are protected in use. A red flag is deferring the whole question to "whatever your contract says." Contracts settle ownership, but visibility is a technical question.

One general principle: every approach in this guide has conditions where it works best, and a strong vendor can tell you what those conditions are for theirs and how they handle everything else. A vendor who cannot do so does not understand their approach well, or does not expect you to ask.

Glossary

Anonymization: Transforming data so it is no longer personal data at all, and therefore outside privacy regulation. A high bar, and hard to clear for rich data.

Clean room: A neutral environment where parties contribute data, run queries against the combination, and receive only aggregate results.

Confidential computing: Protecting data while it is being processed, typically using hardware-based isolation. See trusted execution environment.

Data in use: Data that is actively being processed, as distinct from data at rest (stored) or in transit (moving). The gap that most conventional protections leave open.

De-identification: Removing or obscuring identifiers so data no longer identifies individuals. Under HIPAA, achieved via Safe Harbor or Expert Determination.

Differential privacy: A mathematical method for adding measured noise to results so no individual record can be isolated from the output.

Federated computing: Sending computation to where data lives rather than moving data to the computation. Covers model training, analytics, inference, and data preparation.

Federated learning: Federated computing applied specifically to training machine learning models.

Hash: A fixed-length code generated from a value by a one-way function. The same input always produces the same hash, but the original can't be recovered from it.

Homomorphic encryption (HE): Performing computation directly on encrypted data, without decrypting it at any point.

Multi-party computation (MPC): A cryptographic method letting several parties compute a shared result without revealing their individual inputs.

Privacy-enhancing technologies (PETs): The umbrella term for this whole category. You will also see "privacy-enhancing computation."

Pseudonymization: Replacing direct identifiers with substitutes. Recognized under the GDPR as a safeguard; the resulting data remains personal data.

Synthetic data: An artificial dataset that reproduces the statistical shape of real data without containing real records.

Tokenization: Substituting a sensitive value with a meaningless token, with the mapping held separately in a vault.

Trusted execution environment (TEE): A hardware-isolated region of a processor where code and data are protected from everything outside it, including the operating system and cloud provider.

Trusted research environment (TRE): A controlled environment holding centralized data, where approved researchers work under governance controls and outputs are reviewed before release. Not to be confused with a TEE.

Where Rhino Fits

The Rhino Federated Computing Platform (Rhino FCP) is built around the federated approach described above and supports many of the other approaches covered in this guide, including MPC, HE, TEEs, differential privacy, de-identification, and encryption at rest and in transit. The Rhino FCP runs analytics and AI where each dataset already lives, whether that's across partner organizations, multiple clouds, or your own servers, and adds other protections from this guide when a project calls for them.

Secure data collaboration is fast becoming standard practice, and the leaders whose goals depend on it should have a direct voice in how it happens. We wrote this guide so you can weigh your options, ask vendors the right questions, and take a full part in those decisions.

Want to go deeper into how federated computing works in practice, and what it has delivered for organizations running it at scale? Check out our white paper on moving beyond data silos.

Download the white paper →

Build your Federated Network