Getting More Value from Existing Research Data
Valuable evidence can remain in a dataset after its original study, evaluation or organisational analysis is complete. Finding it requires a defensible question, suitable variables and interpretation that respects what the data can—and cannot—support.
Data can outlive the question that originally produced it
Research studies, surveys, programme evaluations, baseline and endline assessments, routine monitoring, health research and organisational analyses all produce data for particular objectives. The original analysis is usually designed around those objectives; it does not necessarily exhaust every analytically defensible question represented by the variables.
That possibility is not permission to search indiscriminately for interesting results. Re-use must remain consistent with the data's nature, quality and provenance, as well as consent, governance and permissible use. A useful first step is to establish what was collected, why it was collected and what population and period it can reasonably describe.
Secondary analysis is more than rerunning the original statistics
A re-analysis can ask a different but defensible question, examine an unexplored relationship, construct a justified variable or reassess a subgroup pattern. It may compare plausible model specifications, revisit assumptions, improve visualisation or bring related variables into a more informative analytical framework.
Alternative methods are valuable only when they fit the question and data. A more complicated model is not intrinsically more rigorous, and changing techniques solely to obtain a preferred result undermines rather than strengthens the analysis.
Start with the question, not the statistical technique
Research question → Available variables → Data structure and quality → Appropriate analytical method → Interpretation
This sequence prevents a fashionable technique from becoming the rationale for the work. A question establishes the target of inference or prediction. The available variables then determine whether it is measurable, while the data structure and quality constrain which methods are credible. Only after those decisions should an analyst select a technique and define how its outputs will be interpreted.
For organisations considering this process, research and evaluation support can help sharpen the question, while data science and analytics support addresses the analytical design and implementation.
Reassess data quality before re-analysis
Data suitable for an original descriptive purpose may be unsuitable for a new modelling objective. Review missingness, inconsistent coding, duplicates, outliers, variable definitions, measurement limitations and completeness. Sample characteristics, potential selection processes and the quality of documentation also affect what can be learned.
These checks are analytical, not merely administrative. Missingness may be systematically related to an outcome; an apparent subgroup may reflect coding changes; and a variable name may conceal changes in definition. Sophisticated modelling cannot repair data that are fundamentally unsuited to the question.
New questions must remain within what the data can support
Association, prediction, explanation and causation are not interchangeable
An association asks whether variables vary together. Prediction focuses on estimating an outcome for new observations. Explanation or inference may quantify particular relationships and their uncertainty. Causal inference asks what would happen under a change or intervention and requires assumptions and design features beyond a fitted association.
An observational or cross-sectional dataset should not casually be converted into a causal claim. A new model does not change the underlying study design, restore an unmeasured confounder or establish temporal order that was never observed.
When advanced analytics may add value
Regression modelling can estimate adjusted relationships; subgroup analysis can explore meaningful variation; and sensitivity analysis can test how conclusions respond to defensible choices. Predictive modelling and machine learning may help where prediction, nonlinear structure or complex interactions are central. Visual analytics can reveal structure and communicate uncertainty, while automation can make repeated, well-defined steps more consistent.
The appropriate method depends on the research question, sample, outcome, predictors, study design and realistic validation opportunities. The site's research analytics contributions provide context for how analytical work connects with research reporting, without suggesting that any single method suits every dataset.
Re-analysis can also test robustness
Secondary analysis is not only a search for something “new.” It can assess whether conclusions are sensitive to modelling choices, remain consistent across reasonable specifications, or change when influential observations are examined. It can also test whether assumptions or alternative operational definitions materially alter interpretation.
Robustness checks show how results behave under specified alternatives; they do not guarantee validity. Consistent estimates can still reflect shared measurement problems, selection bias or unsupported assumptions.
The value of reproducibility
Documented analytical decisions make re-analysis easier to examine and extend. Reproducible code, where appropriate, can preserve a trace from raw inputs through cleaning, transformations and outputs. Clear variable definitions, version control and transparent reporting help distinguish planned decisions from later exploration and allow errors or changes to be located.
Reproducibility does not make a weak design strong, but it makes the analytical process more inspectable and its limitations easier to communicate.
When re-analysis is not appropriate
Further analysis may add little or may be indefensible when:
- critical variables were never collected or measurement quality is inadequate;
- the sample cannot address the new question or missingness is prohibitive;
- data provenance is unclear or documentation cannot resolve important definitions;
- data-use permissions do not allow the proposed work; or
- the intended inference exceeds what the study design can support.
Deciding not to fit another model can be the sound methodological conclusion. Where a dataset is suitable, examples in selected project contexts can also clarify how analytical questions relate to wider practical objectives without implying that all existing data should be re-used.
Practical takeaway
The value of an existing dataset depends less on how old it is than on whether the data, research question and analytical method are appropriately aligned.
A proportionate assessment of provenance, permissions, variables, quality and design should precede any re-analysis. Sometimes that assessment identifies a credible new enquiry; sometimes it establishes that new data are needed.