Last updated: 2026-10-06
Collecting Evidence, Evaluating It, and Validating the Result
A project's conclusions are only as strong as the evidence behind them, and the evidence is only as useful as the way it was gathered and checked. Students often collect more data than they can analyse, and then evaluate it with methods chosen after the fact. The result is a report that describes what happened without showing that it supports the claims made.the trap of choosing methods after data arrives
This page covers three linked steps: planning data collection so that it answers the question, choosing evaluation methods before the data arrives, and checking that the result is valid. The design of controlled comparisons is covered separately in Designing Experiments and Writing Critical Analysis.
Plan Collection From the Question Backwards
Begin with the success criteria from Scope, Feasibility, and Defining Success. For each criterion, write down the observation that would show whether it is met. Then write down what you will collect to make that observation. If a criterion has no observation attached, it cannot be checked. If an observation has no criterion, it is probably not needed.
A short table keeps this clear:
| Success criterion | Observation needed | How it will be collected | Sample or scope |
|---|---|---|---|
| Task completion is improved | Time and success rate per task | Logged sessions on the same tasks, before and after | Defined user group, fixed task list |
| Results are reproducible | Same output from repeated runs | Scripted runs with recorded seeds and versions | Five runs per configuration |
The columns force the decisions that projects most often leave until too late: what exactly is measured, from whom or what, and how much of it.this is the blueprint for your methods section
Choose Evaluation Methods Before the Data Arrives
Decide how the data will be analysed before you collect it. This protects you from choosing the method that gives the result you hoped for. For quantitative data, state the comparison, the measure, and the threshold for a meaningful difference. For qualitative data, state how you will code or categorise it and who will check the coding. An evaluation plan that could be applied to any dataset is usually too vague.see the experiment design page
The data-preparation steps that come before analysis are covered in Exploratory Data Analysis and Preprocessing. Where the analysis depends on sampling or uncertainty, Probability and Statistics for Computing sets out the underlying ideas.
Validation: Does the Result Mean What You Think?
Validation asks whether the evidence supports the claim you are making, not only whether the numbers are correct. Three questions are useful:
- Does the measure capture the thing? A reduction in clicks may reflect a better interface, or a confusing one that needs more clicks to correct. The measure needs a justification.
- Would the result hold in another case? A result from one dataset, one user group, or one machine is a finding about that case. Say so, and say what you would expect to differ elsewhere.
- Could something else explain it? Name the most plausible alternative explanation, and state what evidence you have against it. The alternatives you cannot rule out belong in the limitations section.
For projects that study a single system, organisation, or group in depth, the guidelines by Runeson and Höst give a structured account of what a case study needs in order to be credible and how it should be reported[1]. They describe case study research as suited to studying contemporary phenomena in their natural context. That suits many student projects, which study a real system or group rather than a laboratory setting. The same guidelines stress that the design, the evidence, and the limits of the case must be reported clearly enough for a reader to judge them.
Triangulating Evidence
A single source of evidence is a single bearing. Where possible, collect evidence of a different kind for the same claim: logs and interviews, measurement and observation, or a benchmark and a user study. Agreement across different sources is much stronger than agreement within one. Disagreement is also informative, and often points to the most interesting part of the project. The idea is developed in Project Navigation and the Art of Triangulation.
Recording the Evidence Trail
Keep the raw data, the scripts that process it, and the decisions made at each stage. This is not only for reproducibility. A marker reading the report will often want to check a claim, and a clear trail makes that possible. The trail can be simple: a folder with dated raw files, the analysis script, and a short note of any decision that changed the plan.
Related Topics
- Designing Experiments and Writing Critical Analysis — controlled comparisons and the discussion of what the evidence shows.
- Literature Reviews That Do Analytical Work — the sources that the evidence must be set against.
- Testing Fundamentals — checking that a system behaves as specified, which is a different question from whether it solves the problem.
- The Report as the Front-End — why the evaluation chapter should be planned from the start.
References
- Runeson, P., & Höst, M. (2009). Guidelines for conducting and reporting case study research in software engineering. Empirical Software Engineering, 14, 131–164.