What kind of study are you running?
The capture method depends entirely on this answer, and many difficulties come from applying the method of one type to another.
| Type | What characterises it | The dominant constraint |
|---|---|---|
| Retrospective | The data already exists, written to treat | Documenting absence: "not noted" is not "absent" |
| Prospective | The protocol precedes observation | Sustaining complete follow-up over time |
| Registry | A population, with no planned end | The ascertainment rate of eligible cases |
| Case-control / cohort | A structured comparison | Comparability of groups, decided in advance |
The first two rows cover nearly all thesis and departmental work. Both methods are detailed in clinical data capture; the third case in the clinical registry.
The full path, and where each decision is made
- Write the question. A question that does not name a population, an exposure and a measurable endpoint is not yet a research question.
- Declare the primary endpoint. Before any data. This is the decision that protects the inference, and the only one no software can take for you — see the methodology.
- Build the variable dictionary. Type, scale, unit, bounds, categories, reasons for absence: the data dictionary.
- Capture, keeping the origin of every value. Capture and provenance.
- Control quality, before the freeze. Data quality.
- Freeze the database. The dataset analysed becomes exactly the dataset described: data management.
- Export and write up. Six pieces rather than a table: the dataset.
The four most expensive mistakes
- Capturing before defining. Cleaning after the fact means guessing what you meant; it never restores what was not written down.
- Choosing the endpoint after seeing the results. The error invalidates the inference without any figure being wrong, and it is undetectable from outside — hence the value of making it mechanically impossible.
- Treating all empty cells alike. "Not sought", "sought and not found" and "not applicable" do not analyse the same way.
- Postponing traceability. It is the only property on this list that can never be recovered: a value entered without its origin will not find it again.
What you will have to report
Writing the paper is easier when you recorded, during the study, what the publication will ask for. STROBE — and more broadly the guidelines gathered by EQUATOR — expect in particular:
- Eligibility criteria and the number of records eligible, included and excluded, with reasons.
- How each variable was measured, and from which source.
- The handling of missing data, and the rate per variable.
- Pre-specified analyses, distinguished from exploratory ones.
These four are by-products of well-run capture. Reconstructed afterwards they are approximate; recorded as you go they are exact and free.
Where to continue
This corpus can be read in any order. Three entry points, depending on your situation:
- You are starting a thesis or a study → clinical data capture, then the dictionary.
- You already have a file in progress → spreadsheet or platform, which says when a spreadsheet stops being adequate.
- You are building a departmental registry → the clinical registry, starting with the governance section.
Frequent questions
Does a retrospective record-based study need ethics committee approval?
That depends on the country, the institution and the nature of the data, and this page cannot decide in place of the competent committee. What is constant: the question arises before capture, not at submission, when it is too late to regularise.
How many records are needed?
That is a question of statistical power, computed from the expected effect and the endpoint — therefore after step 2, never before. No data management method compensates for insufficient numbers.
Can a single-centre study be published?
Yes; most published observational studies are exactly that. The limitation is to be described, not hidden: a single-centre cohort has restricted external validity, and saying so strengthens the paper rather than weakening it.