A collaboration usually begins with a WhatsApp message and a link to a shared folder. Eighteen months later there are four copies of the dataset, two of which have been edited, and an honest disagreement about which one the figure came from.
Nobody behaved badly. The problem is structural: the moment data leaves your lab it stops having a single authoritative version, unless somebody decided in advance that it would.
Anyone about to send data to another lab, or already several months into a collaboration that has become confusing. Works with any storage. The first section costs one email.
Four things to agree before the first file moves
One short email, at the start, answering four questions. It feels excessive and it prevents nearly every problem in this guide.
- Who holds the authoritative copy? One lab. Everyone else holds copies, and knows they are copies.
- What may be done with it? Analysis for this project only, or anything? Can it go into a grant application, a talk, a thesis?
- What happens when it is published? Who deposits it, under whose name, and who is named in the data availability statement.
- What happens if the project stops? Does the other lab delete it, keep it, or keep it for a fixed period?
Write the answers in an email to everyone involved. Not a contract — an email. Its whole value is that it exists in writing, dated, before anyone has an incentive to remember it differently.
Send derived data, keep the originals
The strong default: send what they need to do their work; keep the raw files with you.
This is not possessiveness. It is that raw instrument files are large, often proprietary, and meaningless without acquisition context — and once a second lab holds a copy, two versions of the truth exist. Send the processed, documented, open-format version, and keep the original where it was made.
Where a collaborator genuinely needs the raw files — reanalysis, a method paper, reprocessing with their own pipeline — send them, and record what was sent and when. The mistake is not sharing. It is sharing without a record.
What travels with the data
Data sent without context generates a fortnight of email. A short accompanying note prevents almost all of it.
What this is — one line, in plain words.
When and who — dates of acquisition, and which lab produced it.
How it was made — instrument, key settings, protocol version.
What has already been done to it — normalisation, exclusions, filtering. The single most important line.
What the columns mean — including units. Every time.
What is known to be odd — the plate that ran warm, the replicate that failed.
Who to ask — a role and a person.
Seven lines. Send it as a plain text file inside the folder, not in the covering email — emails get separated from data almost immediately, and the file stays with it.
Keeping track once it has moved
The hard part is not moving files — it is that a copy arrives at the other end stripped of everything that made it interpretable. woodle.cloud keeps a result attached to the run and the file behind it, so what you share carries its provenance with it. The seven-line note does the same job by hand, and costs nothing.
Where your files live →When it has already gone wrong
Several copies exist, at least one has been edited, and nobody is certain which produced the figure. This is recoverable, and the sequence matters.
- Stop editing anything. Tell everyone to work from copies until this is settled.
- Collect every version into one place, each in a folder named for where it came from and when.
- Compare them. Row counts, column names, checksums if you can. Usually two are identical and the differences are small and explainable.
- Nominate one as authoritative — in writing, to everyone — and say why.
- Keep the others, marked superseded. Do not delete them; a published figure may have come from one.
Do this before the manuscript, not during revision. It is a bad week either way, and a far worse one under a deadline.
The awkward one: authorship and credit
Data-sharing disagreements are rarely about data. They are about credit, arriving late and in the wrong conversation.
Two sentences in that first email prevent most of it: what contribution is expected to warrant authorship, and who decides if it becomes unclear. It is uncomfortable to write when everyone is enthusiastic and nothing has been found. It is considerably more uncomfortable eighteen months later, with a result on the table.
The four questions, the seven-line note and the checklist are free to copy into your own practice. No attribution needed and no signup. If you take one thing, send the seven-line note with the next dataset that leaves your lab.