Guides · Running a lab

Sharing data with a collaborator without losing track of it

What to send, what to keep, and the four things to agree in writing before the first file moves. Written for collaborations that will outlast the people who started them.

Reading time
9 min
Licence
Free to copy into your own lab documentation. No signup.

A collaboration usually begins with a WhatsApp message and a link to a shared folder. Eighteen months later there are four copies of the dataset, two of which have been edited, and an honest disagreement about which one the figure came from.

Nobody behaved badly. The problem is structural: the moment data leaves your lab it stops having a single authoritative version, unless somebody decided in advance that it would.

Who this is for

Anyone about to send data to another lab, or already several months into a collaboration that has become confusing. Works with any storage. The first section costs one email.

Four things to agree before the first file moves

One short email, at the start, answering four questions. It feels excessive and it prevents nearly every problem in this guide.

  1. Who holds the authoritative copy? One lab. Everyone else holds copies, and knows they are copies.
  2. What may be done with it? Analysis for this project only, or anything? Can it go into a grant application, a talk, a thesis?
  3. What happens when it is published? Who deposits it, under whose name, and who is named in the data availability statement.
  4. What happens if the project stops? Does the other lab delete it, keep it, or keep it for a fixed period?

Write the answers in an email to everyone involved. Not a contract — an email. Its whole value is that it exists in writing, dated, before anyone has an incentive to remember it differently.

Every difficult conversation in a collaboration is easy while nothing has been discovered yet. Have it then.

Send derived data, keep the originals

The strong default: send what they need to do their work; keep the raw files with you.

This is not possessiveness. It is that raw instrument files are large, often proprietary, and meaningless without acquisition context — and once a second lab holds a copy, two versions of the truth exist. Send the processed, documented, open-format version, and keep the original where it was made.

Where a collaborator genuinely needs the raw files — reanalysis, a method paper, reprocessing with their own pipeline — send them, and record what was sent and when. The mistake is not sharing. It is sharing without a record.

What travels with the data

Data sent without context generates a fortnight of email. A short accompanying note prevents almost all of it.

The note that travels with every dataset

What this is — one line, in plain words.
When and who — dates of acquisition, and which lab produced it.
How it was made — instrument, key settings, protocol version.
What has already been done to it — normalisation, exclusions, filtering. The single most important line.
What the columns mean — including units. Every time.
What is known to be odd — the plate that ran warm, the replicate that failed.
Who to ask — a role and a person.

Seven lines. Send it as a plain text file inside the folder, not in the covering email — emails get separated from data almost immediately, and the file stays with it.

Keeping track once it has moved

Before and after you send 0 / 7
Keep a one-line log of every dataset sentDate, what, to whom, which version. A spreadsheet. This is the whole system.
Send a dated, versioned folder, never a live syncA shared folder that keeps changing means nobody can say what they analysed.
Include the seven-line note inside the folderInside, so it cannot be separated from the data.
Say plainly whether this is raw or processedThe commonest misunderstanding in any collaboration, and the easiest to prevent.
Agree how corrections will be sentThey will be needed. Decide now, or the fix will arrive as an attachment nobody files.
Check whether anything needs approval before it leavesHuman participant data, a third party's material, anything under an agreement.
Note who at the other end actually has itPeople move. "The Kumar lab" is not a custodian; a named role is.
Why we care about this

The hard part is not moving files — it is that a copy arrives at the other end stripped of everything that made it interpretable. woodle.cloud keeps a result attached to the run and the file behind it, so what you share carries its provenance with it. The seven-line note does the same job by hand, and costs nothing.

Where your files live →

When it has already gone wrong

Several copies exist, at least one has been edited, and nobody is certain which produced the figure. This is recoverable, and the sequence matters.

  1. Stop editing anything. Tell everyone to work from copies until this is settled.
  2. Collect every version into one place, each in a folder named for where it came from and when.
  3. Compare them. Row counts, column names, checksums if you can. Usually two are identical and the differences are small and explainable.
  4. Nominate one as authoritative — in writing, to everyone — and say why.
  5. Keep the others, marked superseded. Do not delete them; a published figure may have come from one.

Do this before the manuscript, not during revision. It is a bad week either way, and a far worse one under a deadline.

The awkward one: authorship and credit

Data-sharing disagreements are rarely about data. They are about credit, arriving late and in the wrong conversation.

Two sentences in that first email prevent most of it: what contribution is expected to warrant authorship, and who decides if it becomes unclear. It is uncomfortable to write when everyone is enthusiastic and nothing has been found. It is considerably more uncomfortable eighteen months later, with a result on the table.

Take it with you

The four questions, the seven-line note and the checklist are free to copy into your own practice. No attribution needed and no signup. If you take one thing, send the seven-line note with the next dataset that leaves your lab.