Reliability and Validity in Research (2026): Types, Threats & How to Demonstrate Them
Your examiners will ask one question above all others when reading your methodology chapter: “How do you know your measurements actually measured what you say they measured, and would they produce the same results if someone else repeated them?” These twin demands — that your instruments are consistent and that they are accurate — are what researchers call reliability and validity in research. Getting them wrong does not merely weaken one section of your thesis; it undermines every finding that follows.
This guide gives you the conceptual framework, the standard threshold values, the classical threats you must address, and the reporting language that satisfies doctoral examiners at institutions following Creswell, Shadish, and Lincoln & Guba traditions — all in a single, chapter-ready reference.
If you have not yet decided whether your study will be quantitative, qualitative, or mixed, start with our primer on qualitative vs quantitative research to orient the concepts here within your broader design.
Reliability vs Validity: Core Definitions
Think of reliability as a bathroom scale that always shows the same reading for the same object regardless of the day or who stands on it. Think of validity as whether that scale is actually measuring your weight — not the weight of your clothes, not fluid retention, but the underlying construct you care about. A scale that consistently over-reads by 5 kg is reliable but not valid. A scale whose readings fluctuate randomly is neither reliable nor valid. You need both.
In the measurement literature this relationship is expressed with a sporting analogy: a player who consistently hits the left edge of a target is reliable but not valid (biased); a player who scatters shots around the bullseye is valid on average but not reliable (inconsistent). The distinction matters practically because the strategies for improving each are different — reliability problems are usually addressed through instrument design, rater training, and sample size; validity problems require reconceptualisation of the construct or revision of the measurement approach.
Types of Reliability

Test-Retest Reliability
Test-retest reliability assesses the stability of a measure over time. You administer the same instrument to the same participants on two occasions — separated by an interval long enough that memory carry-over is minimised, but short enough that the underlying trait should not have genuinely changed — and correlate the two sets of scores. An intraclass correlation coefficient (ICC) or Pearson correlation of ≥ 0.70 is generally considered acceptable; ≥ 0.80 indicates good stability for most social science instruments. A two-to-four-week interval is a commonly used default for attitude scales in education and psychology research.
Inter-Rater Reliability
When data are produced by human observers or coders — qualitative coding, rubric-based assessment, clinical rating scales — you need inter-rater reliability (IRR) to confirm that different raters produce consistent scores. The appropriate statistic depends on the measurement level: Cohen’s Kappa (κ) for categorical data, weighted Kappa for ordinal categories, and the ICC for continuous ratings. For a detailed walkthrough of calculation and interpretation, see this guide to inter-rater reliability and Cohen’s Kappa. Following Landis and Koch (1977), κ ≥ 0.61 is considered substantial agreement and κ ≥ 0.81 almost perfect. If initial agreement falls below thresholds, conduct a calibration session in which raters discuss borderline cases, revise the coding scheme if needed, and re-code a subsample.
Internal Consistency
Internal consistency captures the degree to which items on a multi-item scale measure the same underlying construct. The most widely reported statistic is Cronbach’s alpha (α). Conventional thresholds are:
| Alpha value | Interpretation |
|---|---|
| ≥ 0.90 | Excellent (may indicate redundant items) |
| 0.80 – 0.89 | Good |
| 0.70 – 0.79 | Acceptable |
| 0.60 – 0.69 | Questionable — acceptable only for exploratory work |
| < 0.60 | Unacceptable for published research |
However, alpha assumes tau-equivalence — that every item loads equally on the latent factor — an assumption rarely met in practice. McDonald’s omega (ω) derives estimates from actual factor loadings and is increasingly preferred in methods-rigorous journals. Reporting both α and ω is recommended in 2026 when item loadings are unequal.
Parallel-Forms Reliability
Parallel-forms reliability is assessed when two equivalent versions of the same instrument exist. Participants complete both, and the correlation between scores indicates equivalence. This is used primarily in standardised testing contexts where memory effects from a single administration would inflate test-retest estimates.
Types of Validity
Content Validity
Content validity asks whether your instrument’s items adequately sample the full domain of the construct. It is established through expert judgement — a panel of five to ten subject-matter experts who rate each item on relevance and clarity. Lawshe’s Content Validity Ratio (CVR) or the Content Validity Index (CVI) quantify panel agreement; a CVI of ≥ 0.80 is the commonly cited acceptability threshold.
Construct Validity
Construct validity encompasses all evidence that your instrument actually measures the theoretical construct it purports to measure. It is examined through two complementary routes:
- Convergent validity: Scores should correlate positively with other instruments measuring the same construct.
- Discriminant validity: Scores should correlate weakly with instruments measuring theoretically distinct constructs.
In structural equation modelling (SEM), convergent validity is evaluated using Average Variance Extracted (AVE ≥ 0.50), and discriminant validity via the Fornell-Larcker criterion. Confirmatory Factor Analysis (CFA) is the primary tool for evaluating construct validity formally.
Criterion Validity
Criterion validity assesses whether your measure predicts or correlates with an external criterion:
- Concurrent validity: The measure correlates with an established criterion administered at the same time.
- Predictive validity: Scores at Time 1 predict an outcome measured at Time 2.
Internal Validity
Internal validity is the degree to which a study can support a causal inference — that changes in the independent variable caused changes in the dependent variable. Campbell and Stanley’s foundational framework identifies the most common threats:
- History: External events between pre- and post-measurement that affect the outcome independently of the intervention.
- Maturation: Natural changes in participants over time that mimic a treatment effect.
- Testing effects: Familiarity with the instrument from prior exposure improves performance independently of the intervention.
- Instrumentation: Changes in the measurement tool or rater behaviour over data collection.
- Statistical regression: Extreme scorers at pre-test tend toward the mean at post-test due to measurement error.
- Selection bias: Pre-existing differences between groups that explain post-intervention differences.
- Attrition: Non-random dropout of participants that alters group composition.
External Validity
External validity addresses the extent to which findings can be generalised beyond the immediate study context — across different populations, settings, and time points. This is directly linked to your choice of sampling methods in research: probability sampling supports stronger external validity claims than convenience or purposive sampling.
Threats to Reliability and Validity: Consolidated Overview

| Threat | Type affected | Mitigation strategy |
|---|---|---|
| Ambiguous or double-barrelled items | Reliability, content validity | Pilot test; cognitive interviewing |
| Rater drift over time | Inter-rater reliability | Calibration sessions; blind re-coding checks |
| Selection bias | Internal and external validity | Random assignment; defined inclusion criteria |
| Common method bias | Construct validity | Anonymous surveys; temporal separation of measures; Harman’s single-factor test |
| Attrition | Internal validity | Compare dropouts vs completers; intent-to-treat analysis |
| Social desirability bias | Criterion and construct validity | Anonymous response format; social desirability scale as covariate |
| Restricted range in criterion variable | Criterion validity | Sample across the full range; report restriction as a limitation |
Demonstrating Reliability & Validity in Quantitative Research
Claiming reliability and validity without evidence is insufficient for a thesis or dissertation. Your methodology chapter must present a coherent demonstration strategy. Best practice in 2026 typically follows these steps:
- Instrument selection or construction: When using an established scale, cite psychometric evidence from the original validation study and any relevant cross-sample adaptations. When constructing a new instrument, begin with a systematic item-generation phase grounded in existing theory.
- Expert review for content validity: Convene a panel of experts who rate each item on relevance and clarity. Calculate the CVI and report it. Revise or remove items falling below threshold.
- Pilot testing: Administer to a small sample (n ≥ 30) drawn from the same population as your main study. Compute alpha and omega; identify and revise items with item-total correlations below 0.30.
- Confirmatory Factor Analysis: For instruments claiming to measure a specific latent construct, CFA provides the most rigorous evidence of construct validity. Report CFI ≥ 0.95, RMSEA ≤ 0.06, SRMR ≤ 0.08, standardised factor loadings, and AVE.
- Inter-rater reliability (where applicable): Calculate and report the appropriate IRR statistic before main data collection begins. Document calibration procedures in your methods.
Qualitative Trustworthiness: The Lincoln & Guba Framework
Quantitative concepts of reliability and validity do not transfer cleanly to qualitative inquiry, where the goal is rich, contextually grounded understanding rather than statistical replication. Lincoln and Guba (1985) proposed four parallel criteria for establishing trustworthiness, each mapping conceptually onto a quantitative counterpart but demonstrated through procedural rather than statistical means.
| Qualitative criterion | Quantitative parallel | Primary strategies |
|---|---|---|
| Credibility | Internal validity | Member checking; triangulation; prolonged engagement; peer debriefing |
| Transferability | External validity | Thick description of context and participants; purposive sampling rationale |
| Dependability | Reliability | Audit trail of methodological decisions; inquiry audit; inter-coder agreement |
| Confirmability | Objectivity | Reflexivity statement; audit trail; negative case analysis |
Credibility
Credibility concerns whether your findings accurately reflect participants’ perspectives. The most powerful strategy is member checking (respondent validation): returning interpreted data or preliminary themes to participants and asking whether the interpretation resonates with their experience. Combined with triangulation — using multiple data sources, methods, or investigators — member checking substantially strengthens credibility claims. Participant disagreement with your interpretations is not a problem; it is analytically informative.
Transferability
Unlike external validity in quantitative research, transferability does not claim that findings generalise statistically to a defined population. Instead, the researcher provides thick description — sufficiently detailed contextual information about the setting, participants, and conditions — to enable readers to judge whether findings might be applicable to their own contexts. The responsibility for transfer is shared: you create the basis for it; readers evaluate its relevance.
Dependability and Confirmability
Both criteria are supported through maintaining an audit trail — a transparent, documented record of research design decisions, data collection procedures, analytic choices, and how interpretations evolved throughout the inquiry. A reflexivity statement acknowledging your positionality and how prior assumptions may have shaped interpretation is standard practice for confirmability in qualitative studies that employ thematic analysis and other interpretive frameworks.
How to Report in Your Methodology Chapter
The research methodology chapter of a dissertation or thesis should present reliability and validity evidence systematically under a dedicated subsection — typically titled “Reliability and Validity of Instruments” or, for qualitative studies, “Rigour and Trustworthiness.”
For quantitative studies, include:
- The statistical evidence collected (alpha, omega, ICC, CFA fit indices, CVI scores)
- The benchmark thresholds applied, with citations for each (e.g., Nunnally, 1978 for alpha; Landis & Koch, 1977 for Kappa)
- Whether evidence derives from the original validation study, a prior study with your population, or your own pilot
- Any remediation steps taken when initial values were subthreshold
For qualitative studies, include:
- Explicit reference to Lincoln and Guba’s trustworthiness framework
- The specific strategies employed for each of the four criteria and when in the research process they were applied
- A reflexivity statement acknowledging your role as the primary instrument of data collection and analysis
Frequently Asked Questions
What is the difference between reliability and validity in research?
Reliability refers to the consistency of a measure — the degree to which it produces the same results under the same conditions over time or across different raters. Validity refers to the accuracy of a measure — whether it actually captures the theoretical construct it is intended to measure. A reliable measure is not automatically valid, but a valid measure must first be reliable. If your questionnaire produces wildly different scores on two administrations a week apart, it cannot be validly measuring a stable trait.
What is an acceptable Cronbach’s alpha value?
The widely cited convention, originating from Nunnally (1978), is that α ≥ 0.70 is acceptable for research purposes, with α ≥ 0.80 considered good and α ≥ 0.90 excellent but potentially indicating item redundancy. Values below 0.60 are generally unacceptable for published research. Alpha is also sensitive to the number of items — a long scale can appear highly reliable even when individual items are poorly correlated. Reporting McDonald’s omega alongside alpha is increasingly recommended in 2026 as it is a more accurate estimator when the tau-equivalence assumption is violated.
How do qualitative researchers demonstrate validity?
Qualitative researchers use Lincoln and Guba’s (1985) trustworthiness framework rather than quantitative validity criteria. Credibility (parallel to internal validity) is demonstrated through member checking, triangulation, and peer debriefing. Transferability (parallel to external validity) is established through thick description of the research context. Dependability (parallel to reliability) relies on an audit trail of methodological decisions. Confirmability (parallel to objectivity) is achieved through reflexivity statements and negative case analysis.
Can a measure be reliable but not valid?
Yes. A measure can consistently produce the same result (high reliability) while systematically measuring the wrong construct (low validity). The classic illustration is using wrist circumference to measure intelligence — the measurement would be perfectly reproducible but would have no relationship to the construct of intelligence. Validity requires theoretical grounding, expert review, and empirical evidence from correlation studies or factor analysis.
What are the main threats to internal validity?
The main threats identified by Campbell and Stanley include: history (external events that affect the outcome), maturation (natural changes in participants over time), testing effects (familiarity with the instrument), instrumentation (changes in measurement procedures or rater behaviour), statistical regression to the mean, selection bias (pre-existing differences between groups), and attrition (non-random dropout). Randomisation and the use of a control group are the most powerful strategies for controlling multiple threats simultaneously.
What is member checking in qualitative research?
Member checking (respondent validation) is a qualitative credibility strategy in which the researcher returns preliminary findings, themes, or interpretations to study participants and invites them to assess whether the account accurately reflects their experience. It does not require participants to “approve” the findings — disagreement is itself analytically informative. Member checking is one of the most widely used strategies for establishing credibility within Lincoln and Guba’s trustworthiness framework.
