Abstract:
Various biases, knowingly or unknowingly can creep into medical research, affecting the validity of findings and often lead to disappointing results when applied in real-world settings. This article provides a concise overview of more than 30 types of biases that can adversely influence research outcomes. The objective is to raise readers’ awareness of these biases so that pre-emptive actions can be taken to ensure study validity, and thereby enhance the credibility of the findings.
Key words: Biases in Research, Validity of Findings, Selection Bias, Data Bias, Analysis Bias, Minimising Bias
Introduction
Random errors occur due to sampling fluctuations and other uncontrollable factors, whereas biases are systematic errors that can, at least in part, be avoided, or controlled. Biases tend to invalidate the results and give false findings. They affect validity and should not be confused with reliability.
The concepts of reliability and validity were explained in the second article of this series. While reliability refers to the consistency of results when studies are repeated in the same or similar population, validity on the other hand, concerns whether the study truly measures or reflects the intended target. A study can be reliable — producing the same results each time — but still invalid if those results are consistently wrong. Both reliability and validity contribute to the credibility of research findings, but validity is more important. Without validity, findings may deviate from reality, rendering the entire effort futile.
Since biases affect validity, it is important to give them full consideration during the planning and execution of a study and to avoid them wherever possible. It may not be possible to eliminate all biases, but any bias that persists despite best efforts should be explicitly stated as a study limitation, allowing for cautious interpretation of the findings.
Biases in medical investigations can occur at multiple stages. For convenience, they may be divided into 4 categories:
- When conceptualising the problem and defining the study population
- During the selection of cases and controls
- At the data collection stage
- During analysis and interpretation.
This article briefly describes these biases — further details can be found in Indrayan and Holt1 — and concludes with practical steps to avoid or minimise bias and control its impact on results.
Biases in Concepts and Definitions
I regularly see proposals comparing two procedures to determine which is better for a particular differential diagnosis, but the gold standard is not specified. The assessment of ‘better’ cannot be done without a clear reference. If the gold standard is as straightforward as a regular clinical assessment, then the value of evaluating a new procedure is questionable. Although no medical procedure is 100% accurate, the gold standard should be as close as possible to this ideal. Moreover, the gold standard should be relatively more complex or expensive than the procedure under evaluation to make the research more meaningful. Imprecisely defining the research problem at the outset leads to difficulties later, often without being recognised until too late.
Take, for example, studies on diabetes. Will you include cases with fasting blood glucose (FBG) level ≥ 126 mg/dL or glycated haemoglobin (HbA1c) ≥ 6.5 or both? How will you classify patients with elevated FBG but normal HbA1c? Similarly, for hypertension, will you define cases as those with systolic blood pressure (BP) level ≥ 140 mmHg and diastolic BP ≥ 90 mmHg, or just one or the other. What about a case with systolic BP of 138 mmHg and diastolic BP of 92 mmHg? Some studies might even use thresholds of 135/85 or 160/95. The results will vary depending on the chosen definition. If one study uses 140/90 and another uses 135/85, their results are not directly comparable, although both are evaluating cases of hypertension. A target population must be sharply defined with strict inclusion and exclusion criteria that are applied consistently when recruiting participants.
Similar issues arise when defining cancer populations. If you decide to study colon cancer cases, are you including only those who present to a hospital — some at an early stage and others at an advanced stage? Findings could differ significantly between these groups. Likewise, in studies of elderly patients, even specifying “age 70 years and above”, may be inadequate. Patients aged 80 or 90 years often respond very differently than those of age 70 years because of age-related degeneration. Pooling these groups without stratification may lead to results that fail to hold in practice. All these aspects should be clearly specified to avoid biased findings.
Another common source of bias is an ill-defined outcome measure. Merely stating that the outcome is “response” or “improvement after a regimen” is insufficient unless the exact measurement criteria are specified. For example, a reduction in pain by at least three points on the visual analogue scale (VAS), ability to walk without assistance for at least two meters at a fixed time-point after treatment initiation, duration of hospitalisation, mortality rate, or any such event. Even mortality assessment must include the time frame, since all patients eventually die, but the research interest may lie specifically in hospital mortality. Additionally, old age is a major confounding factor in any mortality analysis and must be accounted for.
Many researchers, particularly at postgraduate level, fail to define outcomes with the necessary precision. Without explicit criteria, subjectivity creeps in, undermining the study validity.
Biases in Selection of Subjects
An obvious bias occurs when the results derived from hospitalised cases are generalised to all cases of a disease. Patients admitted to hospital tend to be more severe and also represent those who have the means — financial, social, or logistical — to access hospital care. Socioeconomic factors can further distort outcomes; for example, differences in nutrition status across social classes may significantly influence results. To reduce such bias, it is highly recommended to draw a random sample from the target population. The sample size must be large enough to represent the full cross-section of subjects from the target population. A small sample, even if random, risks coverage bias.
Similarly, results from studies conducted on specific groups, such as employees of a company, or teachers, may not be appliable to the general population, as these groups may have different health or nutritional profiles. Studies based on convenience samples (subjects who are readily available or who volunteers) can also lead to bias.
Survival bias is another issue in studies on older populations: the individuals available for study are those who have already survived to that age and may therefore be healthier than those who died earlier and are no longer represented.
Mixing newly diagnosed cases with those who are already undergoing treatment can distort findings, as treatment modifies response patterns. In field studies, men may be underrepresented if they are away at work during data collection, leading to sex-related bias.
In case-control studies, controls must be matched to cases for all relevant factors except the disease itself. In practice, this matching is rarely assured in clinical studies. Historical controls can also introduce bias, as diagnostic methods and treatment strategies may have improved in the intervening period. In clinical trials, baseline imbalances between test and control groups can skew results; that is why random allocation is advocated.
Before–after study designs are particularly vulnerable to bias from the placebo effect and confounding factors over time. Such biases cannot be corrected without repeating the study using parallel control groups. For this reason, before–after studies are of limited value and, ideally, should be avoided altogether.
Finally, ethical considerations require obtaining an informed consent from all participants. However, individuals who are apprehensive, severely ill, or exceptionally healthy may refuse participation. This can introduce unintentional bias, and the results from such studies should be interpreted with caution.
Biases in Data Collection
Most biases occur at the stage of data collection. The tools and instruments used for assessment must be checked for their validity. Laboratory investigations should ideally be conducted in an accredited laboratory, preferably using automated equipment. For biochemical measurements such as creatinine, the analytical method (e.g., Jeffe Kinetic) must be specified, as alternative methods may yield different results. If a tool such as a questionnaire is used, it must be pretested and validated, and designed to avoid leading questions that suggest a particular answer. Interviewers must be adequately trained to establish rapport and get accurate responses. Similarly, scoring systems if used, must be validated for the specific patient group under study.
In clinical assessments, differences in assessor skill can distort results. Some clinicians are more proficient than others and cognitive bias — where an assessor’s values or preferences influence judgement — can easily creep in. This risk is not limited to humans; similar bias can appear in large language models2 currently being developed. Blinding is therefore frequently employed in clinical trials to obtain unbiased responses. Another issue is that some assessors reach conclusions too quickly, which may subsequently be found to be inadequate or incorrect.
All cases must be assessed at the same stage of disease progression, such as at onset, or at a fixed follow-up point (e.g., 3 months post-discharge). The varying time gap between onset and detection or admission in different cases, called length bias, may influence the results. Differences in disease severity between patients must also be statistically adjusted to avoid bias.
Non-adherence and non-response (partial or full) commonly introduce bias in the findings. Subjects may feel disillusioned, cheated, fatigued or overwhelmed by sustained questioning or intensive investigations. In follow-up studies, dropouts are common. Unequal or differential care — for example, when some patients are treated by one clinician and others by another — can bias findings, as can the pooling of data from patients who receive different treatments.
Telephonic interviews, sometimes done for follow-up of patients, are particularly prone to error: patients may be unavailable despite repeated calls, or they may provide careless or inaccurate information. High non-response rates produce serious bias. Even when patients are physically recalled, those who feel better may fail to attend, while those who deteriorate may move to other facilities.
Recall bias is common in patient interviews: some patients remember every event in detail, whereas others forget minor episodes and recall only major ones. Confounding bias,3 —where the effects of two or more factors become inextricably mixed — can sometimes be identified but often remains hidden, especially when knowledge is limited. Many patients, particularly older ones, take multiple treatment or supplements, and these can affect responses if not properly documented. Co-existing illnesses — whether treated or not — can also skew findings. Fortunately, statistical methods can help account for these biases when the number of variables is limited, and the study design is robust.
Another subtle source of error is digit preference, particularly documented in BP measurement.4 Observers tend to favour terminal digits such as 0 and 5, recording, for example, both 139 and 141 as 140. This reduces the standard deviation and biases results. To address this, BP categories such as 125–134, 135–144 are advocated instead of the usual 120–129, 130–139.
Biases in Analysis and Interpretation
While the use of inappropriate analytical methods is a concern, a more insidious issue is data dredging — sometimes described as “torturing the the data till it confesses”. Investigators may try multiple forms of analysis and selectively report whichever method produces the desired result.
Even simple factors can be manipulated in this way. Take anaemia, its impact may be analysed using haemoglobin values as a continuous variable or by applying a threshold to classify patients as anaemic or not. The choice of threshold — 10 mg/dL versus 9 mg/dL or 11 mg/dL — can substantially alter results. Occasionally, thresholds are deliberately chosen to fit the investigator’s hypothesis, a practice that often goes unnoticed by readers.
P-values are among the most widely used statistical measures to draw conclusions. In very large samples, even trivial differences can produce p < 0.05, giving a false impression of importance. Some statisticians recommend using p < 0.001 for extremely large samples5 to minimise errors, but this is rarely followed. Moreover an excessive use of statistical significance has overshadowed clinical or medical significance of the results. Encouragingly, awareness about the difference between statistical and medical significance, and the need to consider effect size — whether it is large enough to change current practice — is gradually rising, though progress is slow.
A common mistake in today’s era of prediction models is using sensitivity, specificity, and the area under the receiver operating characteristic (ROC) curve for assessing the predictive performance of a model and accepting the model if these indices are good. This approach is flawed for several reasons:
- These indices evaluate classification of known cases (disease present vs absent), not the prediction of disease in patients whose status is unknown.
- An ROC area of 0.70 — often deemed “acceptable” — still implies an error rate of up to 30%, which is far too high for clinical decision-making.
- These metrics apply to groups of patients, not individuals. The appropriate method is one-to-one agreement within clinical tolerance, which is seldom used.
Unfortunately, this same error is now being carried into the evaluation of artificial intelligence (AI)-based models, perpetuating biases in evaluation. There are many other such biases that routinely occur without anyone raising a question.
Lastly, do we know enough about the interaction of human body with mind and soul on one hand and with the environment on the other? One paradigm says, “what we do not know is much larger than what we know.” While this ignorance propels research, it rarely brings humility in our reporting. Without such humility, conclusions are prone to bias.
Publication bias is also well known: studies with positive findings get published more easily than those reporting negative results. This skews the literature towards favourable outcomes and creates a distorted picture during evidence reviews. The imbalance may also tempt investigators to overwork their data to achieve positive findings that will be publishable. Perhaps negative findings deserve twice the weight in publication decisions to correct this distortion.
Fortunately, awareness of analytical bias is increasing, and many researchers are now more careful. Unfortunately, these lessons are often ignored by some researchers, particularly at the post graduate level, laying a weak foundation that carries to the next generation of researchers.
Steps to Minimise Bias
Indrayan and Malhotra6 have suggested several actions to minimise bias in medical research results. Some of these, along with the related type of bias, have been highlighted in this article.
The foremost responsibility of a researcher is to recognise that research is a pursuit of truth — a goal that can only be achieved with personal honesty and a willingness to invest the necessary effort. A non-judgemental frame of mind is needed to accept whatever findings the research reveals. Failing to confirm a hypothesis is not a failure; it simply indicates that the truth lies elsewhere, or else the design, data, or analysis may have been inadequate. This provides an important lesson for future research. Journal editors and reviewers must also cultivate a similar attitude.
It is always desirable to multiple checks to assess validity. Peer review is one such measure. Statistically, the results should be checked for internal validity by splitting the sample into training and validation sets, while external validity should be assessed by replicating the results on another sample from the same or a similar population. Variations should remain within the limits of random sampling fluctuations. If discrepancies arise, the process, data, and analysis must be revisited to identify the problem and take corrective steps. Valid, unbiased results require deliberate effort, which is sometimes lacking due to an eagerness for quick recognition.
While statisticians often recommend large sample sizes, and these can indeed help address several problems, their importance is sometimes overstated. Large samples may lead to sloppy data collection caused by carelessness, fatigue, or limited resources, which can reduce the validity of the results in some situations. In contrast, intensive and in-depth investigations of even a small number of cases can yield highly valuable insights. Many Nobel Prizewinning studies have relied on the depth of investigation rather than a large sample size.
Abhaya Indrayan. Research Methodology and Biostatistics Series VII – Guard Against Biases for Validity
of the Findings. MMJ. 2025, September. Vol 2 (3).
References
- Indrayan A, Holt M P. Concise Encyclopedia of Biostatistics for Medical Professionals. CRC Press, 2016.
- Mahajan A, Obermeyer Z, Daneshjou R, et al. Cognitive bias in clinical large language models. npj Digit Med. 2025;8(1):428.
- Spisak T. Statistical quantification of confounding bias in machine learning models. GigaScience. 2022;11:giac082.
- Chandel T, Miranda V, Lowe A, et al. Zero End-Digit Preference in Blood Pressure and Implications for Cardiovascular Disease Risk Prediction—A Study in New Zealand. J Clin Med. 2024;13(22):6846.
- Serdar CC, Cihan M, Yücel D, et al. Sample size, power and effect size revisited: simplified and practical approaches in pre-clinical, clinical and laboratory studies. Biochemia medica. 2021;31(1):27–53.
- Indrayan A, Malhotra R. Medical Biostatistics. 4th Edition. CRC Press; 2017.