Data analysis is the process of organising, examining, interpreting, and communicating research evidence so that it can answer the study’s research questions.
The appropriate analysis depends on:
Data analysis should therefore be planned as part of the research design rather than treated as an activity that begins only after data collection.
A useful principle is:
Research question → Data → Analytical method → Findings → Interpretation → Conclusion
Data analysis involves transforming collected research information into meaningful findings.
Depending on the study, this may involve:
Analysis should produce evidence that contributes directly to answering the research problem.
The terms are related but are not always used identically.
Data analysis commonly refers to the process of examining a dataset or body of evidence to answer particular questions.
Data analytics can refer more broadly to the use of analytical methods, technologies, systems, and processes to generate insights from data.
In academic research, the emphasis is generally on the analytical procedures needed to answer the research questions.
Poor analysis can undermine an otherwise well-designed study.
A research project may have:
and still produce weak conclusions if the analysis is inappropriate.
Good analysis helps researchers:
The analysis should be considered during research planning.
Before collecting data, ask:
Planning early can prevent collecting large amounts of information that cannot meaningfully answer the research questions.
Different research questions require different forms of analysis.
For example:
Descriptive question
What proportion of customers use online support?
Potential analysis:
Comparative question
Does satisfaction differ between two customer groups?
Potential analysis:
Relationship question
Is employee training associated with productivity?
Potential analysis:
Experience question
How do employees experience workplace automation?
Potential analysis:
Research data can be analysed using quantitative, qualitative, or mixed approaches.
Usually focuses on numerical measurements and may include:
Usually focuses on:
Combines quantitative and qualitative evidence in a coordinated research design.
Before analysing data, researchers should understand what they have collected.
Preparation may include:
Data preparation is part of rigorous analysis.
Data cleaning involves identifying and addressing problems that could affect analysis.
Potential problems include:
Researchers should document important cleaning decisions.
A data dictionary describes the variables in a dataset.
It may contain:
A data dictionary can make analysis more consistent and reproducible.
Coding converts responses into a format that can be analysed.
For example:
Employment status
The numerical codes are labels for categories.
They do not necessarily indicate mathematical quantities.
Qualitative coding involves assigning labels to meaningful portions of text or other qualitative material.
For example, interview data about remote work might generate codes such as:
Codes can later be organised into broader categories and themes.
A typical quantitative workflow may involve:
The exact workflow varies by study.
A qualitative workflow may involve:
The process is often iterative rather than strictly linear.
Descriptive analysis summarises the observed data.
Common outputs include:
Descriptive analysis should usually precede more complex statistical modelling.
Frequency analysis counts how often each category or value occurs.
For example:
| Employment Type | Frequency |
|---|---|
| Full-time | 120 |
| Part-time | 45 |
| Self-employed | 35 |
The appropriate denominator should be clear when percentages are reported.
Percentages make group sizes easier to compare.
For example:
60% of respondents reported using online learning platforms.
The researcher should clarify whether the percentage is based on:
Common measures include:
The mean is useful for many numerical variables but can be influenced by extreme observations.
The median is often useful when distributions are skewed.
The mode identifies the most frequently occurring value or category.
Measures of dispersion describe how widely observations vary.
Examples include:
The appropriate measure depends on the data and analytical objective.
Researchers should understand the distribution of important variables.
Useful tools include:
Distributional characteristics can influence the choice and interpretation of statistical methods.
An outlier is an observation that is unusually distant from other observations according to a defined criterion.
Outliers may represent:
Do not automatically remove an outlier simply because it affects a result.
Investigate it first.
Missing data should be examined before analysis.
Questions include:
The appropriate strategy depends on the research design and missing-data mechanism.
Possible approaches include:
There is no universal method that should be applied to every dataset.
Visualisation can help identify:
Common visualisations include:
Choose visualisations based on the analytical purpose.
Tables are useful when readers need precise numerical values.
Graphs are useful when readers need to understand:
Avoid presenting the same information repeatedly in multiple formats without a clear reason.
Cross-tabulation examines the distribution of one categorical variable across another.
For example:
| Employment Type | Satisfied | Not Satisfied |
|---|---|---|
| Full-time | 80 | 40 |
| Part-time | 25 | 20 |
Cross-tabulations can help identify potential group differences.
Correlation examines the association between variables.
Depending on the data and assumptions, researchers may consider measures such as:
The choice should reflect the measurement and distribution of the data.
Regression analysis can estimate relationships between an outcome and explanatory variables.
It may be used for:
The model must match the type of outcome and research design.
Researchers may need to compare:
The appropriate method depends on:
Do not select a test solely because it is commonly taught or available in software.
Hypothesis testing provides a formal framework for evaluating evidence against a specified null hypothesis.
A typical process is:
A p-value can be used within a hypothesis-testing framework to assess compatibility between observed data and a specified null hypothesis under the assumptions of the test.
A p-value should not be interpreted as:
Confidence intervals provide information about the uncertainty surrounding an estimate within a specified statistical framework.
They can be reported for:
Confidence intervals often provide more useful information than a significance decision alone.
Effect sizes communicate the magnitude of a finding.
Depending on the analysis, examples may include:
Reporting effect size can help readers assess practical importance.
A statistically significant result is not automatically important in practice.
For example, a very large dataset can detect a small difference that has little real-world relevance.
Interpret findings in relation to:
Many statistical methods rely on assumptions.
Depending on the analysis, researchers may need to consider:
The relevant assumptions depend on the chosen method.
Qualitative analysis often begins by identifying meaningful segments of data.
A researcher may code an interview excerpt as:
“I received very little training before the system was introduced.”
Possible code:
Insufficient training
Other excerpts with similar meaning may receive the same code.
Related codes can be grouped into categories.
For example:
Codes
may form:
Category: Implementation support
Categories help organise the dataset before broader interpretation.
Themes are broader patterns of meaning that help answer the research question.
For example:
Category: Implementation support
could contribute to:
Theme: Organisational readiness for technological change
A theme should have analytical relevance rather than simply being a frequently mentioned subject.
Thematic analysis can involve:
Researchers should use a clearly defined analytical approach rather than treating thematic analysis as simply highlighting interesting quotations.
Content analysis can be used to systematically examine textual, visual, or other material.
Depending on the approach, researchers may analyse:
Content analysis can be qualitative, quantitative, or combine both elements.
Comparative analysis examines similarities and differences.
Researchers may compare:
Comparison should be guided by the research questions.
When data are collected over time, researchers can examine:
Longitudinal analysis requires attention to the structure of repeated observations.
Mixed-methods research combines quantitative and qualitative evidence.
For example:
Quantitative findings
A survey identifies low employee satisfaction.
Qualitative findings
Interviews reveal that employees associate low satisfaction with workload and limited management communication.
The two forms of evidence can be integrated to produce a more comprehensive interpretation.
Integration can occur through:
Simply collecting quantitative and qualitative data does not automatically make a study genuinely integrated.
Researchers may analyse existing datasets rather than collecting new data.
Potential sources include:
Researchers should understand:
Administrative records can provide valuable research evidence.
Examples include:
Researchers must consider privacy, access, data quality, and appropriate use.
Research datasets may contain:
Sensitive information should be protected according to applicable ethical, institutional, contractual, and legal requirements.
Only collect and retain information that is necessary for the research purpose where possible.
Researchers may use tools such as:
Software selection should reflect the type of analysis, researcher competence, institutional requirements, and reproducibility needs.
Statistical or qualitative software can execute analytical procedures.
It cannot independently determine:
The researcher remains responsible for methodological decisions.
Where practical, maintain:
Reproducibility makes it easier to understand how findings were generated.
An audit trail can document important analytical decisions.
For example:
This can improve transparency.
Sensitivity analysis examines whether findings change when reasonable analytical assumptions or decisions are altered.
For example, researchers may compare results:
Sensitivity analysis can help assess robustness.
A finding is more convincing when reasonable analytical alternatives do not substantially change the substantive conclusion.
Robustness does not mean that every possible analysis must produce identical results.
The researcher should explain meaningful differences.
Unexpected findings should not automatically be treated as errors.
Consider:
Unexpected results can lead to useful discussion and future research.
A finding that does not support the original hypothesis can still be valuable.
Researchers should report it honestly.
A study does not become unsuccessful merely because the expected result was not observed.
Analytical decisions can introduce bias.
Potential sources include:
Transparent procedures can reduce these risks.
Testing many hypotheses or relationships can increase the likelihood of obtaining apparently significant findings by chance.
Researchers should plan analyses carefully and distinguish:
Appropriate statistical methods may be required when multiple comparisons are substantial.
Confirmatory analysis tests hypotheses or analytical plans established before examining the relevant results.
Exploratory analysis searches for potentially meaningful patterns that may generate new hypotheses.
Both can be valuable.
The distinction should be reported honestly.
A research proposal should explain how collected data will be analysed.
For quantitative research, specify where appropriate:
For qualitative research, specify where appropriate:
The proposed analysis should clearly connect to the research questions and objectives.
A strong methodology section may explain:
Avoid simply listing software names.
The results section should focus on what the analysis found.
It may include:
Interpretation should be appropriate to the structure of the dissertation, thesis, report, or article.
The discussion generally explains what the findings mean.
It may address:
The discussion should not simply repeat the results section.
Every major analytical result should contribute to one or more research objectives.
A useful review table can look like:
| Objective | Data Required | Analysis | Finding |
|---|---|---|---|
| Objective 1 | Variable A | Descriptive analysis | Finding 1 |
| Objective 2 | Variables B and C | Regression | Finding 2 |
| Objective 3 | Interview data | Thematic analysis | Finding 3 |
This can help identify analytical gaps.
A large dataset can produce many interesting results that are irrelevant to the study.
Keep the research questions central.
Do not begin with:
Which statistical test should I use?
Begin with:
What question am I trying to answer?
Then determine the appropriate method.
Consider effect size, uncertainty, context, and practical importance.
A statistical method may produce an output even when its assumptions are poorly satisfied.
Investigate unusual observations before deciding how to handle them.
Zero is not necessarily equivalent to missing.
Selective reporting can distort the evidence.
Report relevant results honestly.
Tables should improve understanding rather than overwhelm the reader.
The discussion should interpret rather than simply reproduce the results.
A description reports what the data show.
Interpretation explains what those findings may mean in the context of the research question.
Suppose a researcher investigates:
What factors are associated with employee intention to remain with an organisation?
A possible analysis plan might include:
Employee intention to remain.
Summarise participant characteristics and variables.
Examine distributions, missing values, and relationships.
Use an appropriate regression model if justified by the outcome measurement and study design.
Report estimated relationships, uncertainty, effect sizes where appropriate, and limitations.
The exact analysis would depend on how each variable is measured and the research design.
Suppose the research question is:
How do employees experience the adoption of AI-based workplace tools?
A possible analysis plan could involve:
Semi-structured interviews.
Transcription and organisation of interview material.
Identify statements concerning:
Group related codes.
Develop broader patterns such as:
Compare themes with the research questions and relevant literature.
Before finalising your analysis, ask:
The first step is to understand what the research questions require and what data are available to answer them.
Data cleaning and preparation should then be conducted before substantive analysis.
Yes.
Planning the analysis early helps ensure that the study collects the variables and information required to answer the research questions.
Analysis involves systematically examining the data.
Interpretation considers what the findings mean in relation to the research questions, theory, previous evidence, and study limitations.
Yes.
Mixed-methods research can intentionally integrate both types of evidence.
There is no single universal test.
The appropriate method depends on:
Usually not.
Where appropriate, report the effect estimate, uncertainty, and relevant statistical information rather than relying on a p-value alone.
Thematic analysis is an approach for identifying and interpreting meaningful patterns or themes in qualitative data.
Yes.
Software can assist with organising, coding, retrieving, and managing qualitative data.
The software does not replace the researcher’s analytical judgement.
AI tools can assist with some analytical tasks, but their use should be carefully evaluated.
Researchers remain responsible for:
Confidential or sensitive research data should only be processed through AI systems when appropriate permissions and safeguards exist.
A strong analysis should:
Continue exploring the Lamtas academic research resources:
Explore research planning, data analysis, research reports, evidence-based projects, and consulting services.
Explore data analysis, analytics, artificial intelligence, automation, and data-driven technology services.
Explore tutoring, training, learning resources, and research development support.
Explore research reports, editing, technical writing, documentation, and professional content services.
Good research data analysis is not simply the process of running statistical tests or highlighting interesting quotations.
It is a structured process of turning collected evidence into defensible findings.
The essential chain is:
Research Questions → Data → Preparation → Analysis → Findings → Interpretation → Conclusions
The strongest analyses are aligned with the research design, transparent about analytical decisions, careful about uncertainty, and honest about what the evidence can and cannot establish.
For academic research, the goal is not to produce the most complicated analysis possible.
The goal is to use an appropriate and defensible analytical approach that answers the research questions clearly.