Guides · Funding

Writing the data management plan for a DBT or SERB proposal

Most data management plans are written in the last afternoon, from a template someone forwarded. This is the version that takes about ninety minutes, says true things, and survives contact with a reviewer. With paragraphs you can adapt.

Reading time
11 min
Licence
Free to copy into your own lab documentation. No signup.

A data management plan is the one part of a proposal that nobody enjoys writing and almost nobody reads carefully. That combination produces a particular kind of document: two pages of confident sentences about repositories and backups that describe a lab nobody actually runs.

It matters more than its length suggests. Not because a weak plan sinks a proposal — it rarely does — but because the plan is a promise you make about the next five years, and the person who has to keep it is usually a student who has not joined yet.

Who this is for

Anyone writing a DBT, SERB, ICMR or CSIR proposal in India who has reached the data management section and would rather not invent it. The structure works for most funders; the specifics here lean Indian. Nothing in it requires particular software.

What the section is actually asking

Strip away the phrasing and every funder is asking five questions. Answer these five and you have written the plan.

  1. What data will this project produce? Kinds, formats, rough volume.
  2. Where will it live while the project runs? And who can reach it.
  3. How will it be described so it means something to a stranger?
  4. What is shared, when, and what is not? With the reason.
  5. What happens when the project ends? Retention, and who is responsible.

Everything else is decoration. If a draft is long but leaves one of those five unanswered, it is a weak plan wearing a good suit.

The mistake nearly everyone makes

The commonest failure is not vagueness. It is describing infrastructure you do not have.

Plans routinely promise an institutional repository that does not exist, an automated nightly backup nobody has configured, or deposition in a public archive that has no category for the data in question. None of it is dishonest — it is copied from a template written for a different kind of lab, usually in a different country.

A reviewer who runs a lab can tell the difference between a plan describing a real workflow and a plan describing an ideal one. The real one is shorter and more specific.

It also creates a problem three years later, when a student is told the data must go somewhere the lab has never used, in a format nobody has produced, because a sentence in a funded proposal says so.

Write it from what you already do

Before drafting anything, spend twenty minutes answering these on paper. This is the whole job; the writing afterwards is transcription.

Before you write a word 0 / 8
List every instrument this project will touchAnd what each one actually writes — the vendor file, and whether it can export something open.
Estimate volume per experiment, then multiplyImaging and sequencing dominate. Everything else is rounding.
Name where files sit today, honestlyInstrument PC, a shared drive, a lab external disk, somebody's laptop. Write down what is true.
Check what your institution actually providesAsk the computer centre directly. Many Indian institutions offer storage nobody in the lab knows about.
Find the real repository for your data typeSequence data has one. Most measurement data does not — say so rather than inventing one.
Decide who is responsible by role, not by namePeople leave. "The project's senior research fellow" outlives "Ms Sharma".
Note anything that cannot be shared, and whyHuman subjects, a collaborator's material, a filing in progress. Reviewers respect a stated limit.
Pick a retention period you can actually honourFunder minimum, institutional policy, whichever is longer. Then check you have the disk.

Paragraphs you can adapt

These are deliberately modest. Replace the bracketed parts. If a sentence is not true of your lab, delete it rather than soften it.

Data to be generated

This project will generate [confocal image stacks, plate-reader measurements and analysis outputs]. Primary data will be produced by [instrument or facility], which writes [format], from which [open format] is exported at acquisition. We estimate approximately [N] GB per [experiment/month] and about [N] TB over the funded period. Analysis outputs will be [tabular files, scripts and figures] held alongside the primary data they derive from.

Storage during the project

Working data will be held on [institutional storage / a dedicated lab server / managed cloud storage], with a second copy on [independent medium] refreshed [frequency]. Copies are held in [two] physically separate locations. Access is limited to project members; the [PI and the senior research fellow] can restore from backup. Data are not held solely on instrument computers or personal devices, which are treated as transient.

Documentation and metadata

Each dataset is recorded with the date, the operator, the instrument and settings used, the sample or condition, and a reference to the protocol version followed. Files follow a documented naming convention [see below or attach one page]. Each stated result can be traced to the run and the original instrument file that produced it. Analysis steps, including any manual step, are recorded so a figure can be reproduced from primary data.

Sharing

Data underlying publications will be released at the time of publication [or within N months], deposited in [named repository] where a suitable one exists for the data type. Where no discipline repository exists, data will be deposited in a [general-purpose archive] with a persistent identifier and cited in the publication. Data that cannot be shared — [reason: human participant data, third-party material, protection pending] — are identified here, and processed or aggregated forms will be released instead.

Retention and responsibility

Primary data and the documentation needed to interpret them will be retained for [N] years from the end of the project, in line with [funder / institutional] policy. The Principal Investigator is responsible for the plan; day-to-day custody sits with [role, not a person]. On any team member's departure, their data and documentation are transferred to the shared store before their final month, and the transfer is checked by [role].

Why we care about this

The documentation paragraph is the one most labs cannot honestly write, because the link between a result, the run, and the original file usually lives in somebody's memory. That is the problem woodle.cloud was built for. But the plan above is worth writing whatever you use — and most of it works on a shared drive and a spreadsheet.

See how it works →

Three things reviewers notice

1. A volume estimate that was clearly calculated

“Approximately 40 GB per imaging session, roughly 4 TB over three years” reads as a lab that has thought about it. “Large volumes of data” does not.

2. A named limit

Stating clearly that something will not be shared, and why, reads as more credible than promising everything. Blanket openness is the claim reviewers most often disbelieve.

3. Responsibility that survives turnover

Roles, not names. Every plan written around a specific student expires when they submit.

The budget line nobody claims

Data storage and curation are allowable costs on most schemes, and are among the least contested lines in a budget. A modest, specific request — disks, an archive deposit fee, a share of a storage subscription — is usually approved without argument, and it makes the plan self-consistent: you have described work, and you have asked for the means to do it.

A plan that promises careful curation with no budget line is a plan that quietly relies on somebody doing it in the evenings.

Take it with you

These paragraphs are free to copy, edit and paste into your proposal. No attribution needed and no signup. If a sentence is not true of your lab, cut it — a shorter honest plan beats a long aspirational one.