Boyce Data Science
Data Sharing in a Box
For registry teams and advocacy groups preparing to share data with outside researchers. Check your readiness, then fill in each stage to write the pages a researcher expects to find: how to request data, the terms of use, how requests are reviewed, what the registry holds, and how you work with industry and repositories.
Start here
A registry becomes useful to outside researchers when they can see what it holds, how access is decided, and what they must agree to. Each stage below writes one part of that documentation from your answers, and the last stage assembles everything as Markdown, a web page you can publish, or a Word file.
Your registry
These details fill every document. Anything left empty appears in the documents as a bracketed blank, such as [Registry name], so you can see what still needs an answer.
The group that maintains the registry, for example Your Rare Disease Foundation.
Leave empty to use the same address.
Optional.
How your entries are kept
Nothing you type leaves your browser, and nothing is stored when you close the page. Save progress file, at the top of the page, writes your entries to a file on your computer that Load progress file reads again later, on any device.
What each stage writes
- Readiness check-up: a gap list for your own team, covering documentation, governance, infrastructure, and data standards. It is a check of your program, and no participant data are entered.
- How researchers request data: the researcher information page, an intake form, and frequently asked questions.
- Terms of use: data use terms, suggested acknowledgment and data availability wording, and a page for federally funded projects.
- Review committee: operating procedures and a review rubric.
- Describe your data: the ethics oversight statement, cohort overview, data dictionary, and derived fields.
- Storage and access: a record of how data will be delivered or analyzed.
- Industry and repositories: a policy for sharing with trial sponsors and a decision record for outside repositories.
Sample language, for review by your own experts. Everything this tool writes is a starting draft. Have your attorney, your Institutional Review Board (IRB) or ethics committee, and your governing body review each document before it is published or signed. This tool is educational and is not legal or regulatory advice.
Related tools. The Governance Register covers governance bodies, consent, and regulation in depth and drafts a governance packet for your attorney and IRB. Use it to decide your policies; use this tool to write the pages researchers will read. Data Asset Intake catalogs each data set you hold, which makes the readiness check-up quicker to answer.
Readiness check-up
Check what is currently true of your program. The check-up works in layers: whether the data can be shared responsibly, whether others could understand and reuse them, and, if you plan to join a research network, how close the data are to the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM). Many registries begin sharing with the first layer in place and work through the others over time.
This is a check of your program and its documentation. It is especially useful for registries that combine sources such as health records, surveys, biosamples, imaging, and genomics. Unchecked items become the gap list in the last stage.
Layer one: foundational readiness
Registry content inventory
Governance and documentation
Technical infrastructure
Team roles and capacity
Privacy and ethics
Layer two: standardization readiness
Well-governed data can still be hard to reuse when variables, codes, and meanings differ between sources. This layer checks whether another team could interpret your data without asking you.
Variable definitions and semantics
Standard vocabulary use
Data quality and provenance
Layer three: OMOP readiness
Complete this layer if you plan to take part in federated networks, run analyses across registries, or support studies built on shared cohort definitions. It is optional for everyone else.
Structural fit
Clinical domains present in your data
Practicality of an extract, transform, and load (ETL) project
Operational use
When you reach the point of mapping, The OMOP ETL Worksheet follows the project from source profiling through data quality approval, and the OMOP ETL toolkit holds prompts and a source-to-target mapping record.
Advice for registries with several data sources
- Tag each variable with its source, and keep naming conventions consistent across sources.
- Track provenance: what came from where, and when.
- State which sources are required for a participant to count as enrolled and which are optional.
- Consider sharing a core data set by default and restricting sensitive or high-dimensional data to a separate review.
- Draw a diagram of how data move between systems, and publish a data inventory table that lists each source, its access level, and the status of its documentation.
How researchers request data
The researcher information page is the first thing an investigator reads. It states what to submit, who reviews it, how long review takes, and what happens after approval. Set the details here, and open the draft at the bottom of the stage to read the result.
Review time and steps
Write it the way you want researchers to read it, for example 6 to 10 weeks. The same wording is used on every page, so the pages cannot disagree with each other.
A role, such as registry staff or the committee secretariat.
What a complete request includes
What the committee evaluates
Intake form sections
A short questionnaire sent with the request helps you plan data cuts and support. Choose the sections to include. Consult your IRB or governance committee before using the form with real requestors.
Draft: Information for researchers
Draft: Data request intake form
Draft: Frequently asked questions
Terms of use
The data use agreement (DUA) is the contract a researcher's institution signs before receiving data. This stage writes a plain-language summary of its terms for your public pages, which your attorney can then turn into the agreement itself. Check the terms you intend to require and set the numbers.
Terms and conditions
Review periods and project length
Documents to be signed
Projects with federal funding
The National Institutes of Health (NIH) Data Management and Sharing (DMS) Policy has applied to NIH-funded research that generates scientific data since January 25, 2023. Investigators who plan to use your registry in an NIH project will describe your data in their DMS Plan, and that plan cannot promise more than your consent forms and policies allow.
The plan format changed in 2026. NIH notice NOT-OD-26-046 introduced a shorter format built around yes or no questions, required for applications due on or after May 25, 2026, and notice NOT-OD-26-100 asks existing awards to move to it and to report plan changes in the annual progress report from October 1, 2026. The policy itself did not change. These details were read in September 2026; open the notices in the last stage to confirm what applies now.
Why registry data usually cannot be posted openly
Deidentified registry data can still point to a person when the condition is rare, when the community is small, or when a record can be traced through a clinic or a region. Consent forms for registries commonly permit research use under review and do not permit public posting of record-level data. NIH recognizes ethical, legal, and technical reasons for limiting sharing, and controlled access through your committee is a recognized way to meet the policy.
For the same reason, the draft asks investigators not to place registry data in public repositories or manuscript supplements, including code-hosting and data-hosting sites, unless the registry and the relevant IRB have agreed in writing.
Draft: Data use terms
Draft: Projects with federal funding
Review committee
A Data Access Committee (DAC) reviews each request against your policies and records its decision. Written standard operating procedures (SOP) and a shared rubric are what make decisions consistent from one request to the next and defensible when someone is turned down.
Committee structure
Members. Check the perspectives your committee includes.
For example quarterly, or as requests arrive.
Rubric sections
Each section becomes a table of criteria with a rating column. Requests involving more sensitive data call for closer review, so criteria do not have to count equally.
To work through a live request with the committee, The Governance Register includes a request scoring page. The rubric written here is the paper version for your procedures manual.
Draft: Committee operating procedures
Draft: Review rubric
Describe your data
Researchers judge whether a request is feasible from a description of who is in the registry, which fields exist, and how calculated fields were made. This stage writes those pages, along with a statement of the ethics oversight the registry operates under.
Ethics oversight
A date, or a phrase such as continuing review not required.
One or two sentences in plain language, for example that participants agreed to the use of their deidentified data for research reviewed by the registry. If several approvals cover different parts of the data, describe each.
Cohort overview
Mean or median, with the range.
One per line: the most common diagnoses or subtypes, whether treatment history is recorded, and the clinical variables researchers ask about most.
One per line.
Small numbers. In a rare disease registry, a published table can identify a person when a cell holds very few people. Many programs suppress or combine cells below a threshold that they set with their IRB, and report ranges in place of exact values for small sites or regions. Apply your own rule before publishing this page.
Data dictionary
List the fields a researcher can request. Give each a clear description, its type or units (for example Date, YYYY-MM-DD, or Categorical: 0 = No, 1 = Yes), and any note on how it was collected. Import a CSV reads a REDCap data dictionary as exported, including one from The Registry Mockup, or any CSV with columns named field_name, description, type_or_units, and notes. Calculated fields in a REDCap dictionary are added to the derived fields below. If your registry has several modules, a dictionary per module is easier to read than one long table.
| Field name | Description | Type or units | Notes | Remove |
|---|
Derived fields
A derived field is calculated from other fields and was not collected directly. Describe each one in enough detail that another researcher could recreate it from the source data, and cite any outside algorithm it depends on.
| Derived field | How it is calculated | Remove |
|---|
Draft: Ethics oversight
Draft: Cohort overview
Draft: Data dictionary and derived fields
Storage and access
Where the data are stored and how an approved researcher reaches them shape your security obligations, your costs, and who can realistically use the data. Record your choices here; they appear on the researcher information page and in your internal records.
Your approach
Analysis tools your requestors use. The intake form asks each requestor the same question.
Questions to settle with your host
Which privacy law applies. The Health Insurance Portability and Accountability Act (HIPAA) governs health plans, most health care providers, and the businesses that work for them. A registry run by an advocacy group that collects data directly from participants is often outside HIPAA, and may instead fall under state health privacy laws, consumer protection rules, or the General Data Protection Regulation (GDPR) for participants in Europe. Data received from a clinic or health system usually arrive with HIPAA conditions attached. Ask counsel which laws apply to each source before relying on HIPAA terms such as deidentified or limited data set in your documents.
Hosting options compared
Commercial cloud. The large cloud providers offer services that can be configured for health data under a business associate agreement, along with storage, databases, and hosted notebooks for R, Python, and SQL. The registry controls the account and is responsible for configuring it correctly. Product names and compliance programs change often, so confirm current offerings with the provider.
Institutional servers. A university or hospital partner maintains the environment. Some data governance agreements require this. The registry gains direct oversight by a known party and avoids a third party, and gives up some flexibility: scaling, security patching, and accounts for outside collaborators are often slower.
Vendor platform. The registry vendor stores the data and may offer researcher access. Ask how data can be exported in full, in what format, and what happens to access if the contract ends.
Secure enclaves as an alternative to file transfer
In an enclave, also called a trusted research environment, researchers log in to a controlled workspace and analyze the data there. Record-level data do not leave, activity is logged, and the software environment can be kept consistent, which helps reproducibility. The trade-offs are that researchers must adapt to notebook or remote desktop workflows, that people who rely on point-and-click statistics packages may need support, and that someone has to onboard users and review results before release.
Working Inside a Trusted Research Environment explains enclaves for registry holders and for the researchers who will use them, and Enclave Ready translates Stata, SPSS, and SAS code for researchers moving into one.
Matching the access method to the need
| Need | Common fit |
|---|---|
| Sharing statistical files with a small number of approved teams | Encrypted transfer to the recipient institution under a DUA |
| SQL analysis and routine reporting | A hosted relational database with named accounts |
| Collaborative analysis in notebooks | A hosted notebook environment with project workspaces |
| Protected health information or a limited data set | An enclave, with IRB approval and a DUA in place |
When unsure, a small pilot with one or two trusted research teams shows where users struggle before access is opened more widely.
Draft: Storage and access record
Industry and repositories
Requests from companies running clinical trials, and invitations to contribute data to an outside repository, raise questions an academic request does not. This stage writes a policy for the first and a decision record for the second.
Sharing with clinical trial sponsors
When a sponsor is running or planning a trial in your community, data from your registry or symptom tracker can affect the trial even if the sponsor never submits those data to a regulator. Current symptom data could hint at how a blinded trial is going, and access alone can raise questions about whether a trial was influenced. The reading of the guidance below is mine, made in September 2026, and a sponsor's regulatory team may read it differently.
Regulatory background
Food and Drug Administration (FDA), real world data (RWD) and real world evidence (RWE), August 2023. Considerations for the Use of Real World Data and Real World Evidence to Support Regulatory Decision-Making for Drug and Biological Products expects a protocol and statistical analysis plan (SAP) to be defined in advance for a study that will support a marketing application, and asks sponsors to share drafts with the agency for comment. The agency wants confidence that data sources and analyses were not chosen to favor a result.
FDA, assessing registries, December 2023. Real World Data: Assessing Registries to Support Regulatory Decision-Making for Drug and Biological Products describes a registry as enrolling a defined population and collecting prespecified data, and places responsibility on the sponsor to show that registry data are relevant and reliable for the intended use.
European Medicines Agency (EMA), 2023. The Guideline on computerised systems and electronic data in clinical trials sets expectations for validation, user management, security, and the data life cycle, so that trial data are traceable and auditable. Bringing outside application data into a trial without a plan fits poorly with those expectations.
Read together, these documents favor sharing that is planned in advance, governed by an agreement, and kept away from the teams running a trial.
Conditions under which you will share
A delay keeps shared data from reflecting an ongoing trial. Many policies use three to six months.
Situations in which you will not share
Sponsor firewall
Teams that may receive data.
Teams that may not receive data or results drawn from them.
Records you will keep
Contributing to an outside repository
Joining a data repository or aggregator can extend the reach of a small registry, and it changes who controls the data. Work through the considerations with your governing body and write a short note against each. Checked items with their notes become the decision record.
Federated participation as an alternative
In a federated network, each participant keeps its data in place, converts them to a shared structure such as the OMOP CDM, runs the same analysis locally, and returns only aggregate results. Privacy and local control are easier to maintain, and the cost is the work of standardizing the data and running compatible software. This model is often used across jurisdictions and for highly sensitive data.
Draft: Policy on sharing data with industry sponsors
Draft: Outside repository decision record
Assemble your documents
Choose the documents to include, read them through, and download the set. Public pages are written for researchers. Internal records are for your team, your governing body, and your advisers.
Download the selected documents
The Word file is for circulating drafts to reviewers and opens in Word, Google Docs, and Pages. Save as PDF opens your browser's print window with only the selected documents; choose Save as PDF as the destination. Markdown suits GitHub Pages and MkDocs sites. The web page is a single file with a menu that you can publish as it is or paste into your site.
Save progress file, at the top of the page, keeps your answers so that you can return and regenerate the documents after review.
Glossary and sources
Terms used in this tool
- Controlled access
- Data are released only after a request is reviewed and an agreement is signed, as opposed to being posted for anyone to download.
- Data Access Committee (DAC)
- The group that reviews requests for data against the registry's policies and records a decision.
- Data Management and Sharing (DMS) Plan
- The plan NIH requires from funded investigators describing how scientific data will be managed and shared.
- Data use agreement (DUA)
- A contract between the data holder and the recipient institution stating what data may be used, for what purpose, under what protections, and what happens at the end. Some institutions call it a data transfer agreement.
- Deidentified data
- Data from which identifiers have been removed under a recognized standard, such as the Health Insurance Portability and Accountability Act (HIPAA) Privacy Rule in the United States. Deidentified data about a rare condition can still pose a risk of reidentification.
- Derived field
- A field calculated from other fields, such as age at enrollment.
- Enclave, or trusted research environment
- A controlled workspace in which approved researchers analyze data without removing them.
- Extract, transform, and load (ETL)
- The process of moving data from a source system into a target structure such as the OMOP CDM.
- Federated network
- A group of data holders that run the same analysis locally and share only aggregate results.
- General Data Protection Regulation (GDPR)
- The European Union regulation governing personal data, which applies to many registries with European participants.
- Institutional Review Board (IRB)
- The committee that reviews research involving human participants to protect their rights and welfare. Outside the United States it is often called a research ethics committee or research ethics board.
- Limited data set
- Under HIPAA, data that exclude direct identifiers but may include dates and some geographic detail, shared under a DUA.
- OMOP Common Data Model (CDM)
- A shared structure and vocabulary for observational health data, maintained by the Observational Health Data Sciences and Informatics (OHDSI) community.
- Principal investigator (PI)
- The researcher responsible for a project and for compliance with the terms of data use.
- Protected health information (PHI)
- Individually identifiable health information held by an organization covered by HIPAA.
- Provenance
- The record of where a value came from and how it was changed.
- Standard operating procedure (SOP)
- A written description of how a recurring task is done.
- Standard vocabularies
- Shared coding systems, including SNOMED CT for clinical terms, Orphanet (ORPHA) and Mondo for rare diseases, the Human Phenotype Ontology (HPO) for phenotypes, LOINC for laboratory tests with UCUM for units, and RxNorm or the Anatomical Therapeutic Chemical (ATC) classification for medications.
Origin of the templates
- The data use terms are adapted with permission from the Cystic Fibrosis Foundation: CFFPR Data Application and Confidentiality Agreement.
- The committee operating procedures are adapted from ERN EURO-NMD, Standard Operating Procedures, Data Access Committee (DAC), Version 1.0.
- The templates first appeared in the open Data Sharing in a Box template site, github.com/BoyceLab/DataSharingInABox, which remains available for teams who prefer to publish with MkDocs.
Sources
Policy details were read in September 2026. Open the source before relying on a date or a requirement.
- NIH Data Management and Sharing Policy, and its page on sharing and preservation, including acceptable reasons for limiting sharing
- NOT-OD-26-100, Implementation Update: NIH Data Management and Sharing Plan Requirements, which also links to NOT-OD-26-046 on the updated plan elements
- NIH Extramural Nexus, 2026 Pilot Data Management and Sharing Plan Format Available
- NIH sharing policies
- FDA guidance on real world data and real world evidence, August 2023
- FDA guidance on assessing registries, December 2023
- FDA real world evidence program
- EMA guideline on computerised systems and electronic data in clinical trials, 2023
- EMA real world evidence
- Global Alliance for Genomics and Health (GA4GH), for data access governance and data use conditions
- ELIXIR, for governance models and data stewardship in the life sciences
- Guide to Real World Data for Clinical Research, Danielle Boyce and Pavel Goriacko
This tool is educational and is not legal or regulatory advice. Consult legal, regulatory, and ethics experts before putting a data sharing practice in place.