Measurement, samples, and protocol
In this lesson
By the end, you’ll be able to
- Operationalize a construct into a concrete, defensible measure
- Distinguish validity from reliability and judge a measure on both
- Reason about sampling frames and why underpowered studies are not 'small but suggestive'
Sign in to save your progress
You can keep reading without an account, but completed lessons won’t be saved.
Between a good question and real data sits a step that quietly decides whether a study means anything: measurement. A construct like 'stress', 'engagement', or 'skill' is an idea, not a number. Turning it into something you can record, well or badly, determines what your results are actually about. Get this wrong and every later statistic is precise nonsense.
To operationalize a construct is to turn it into a specific, recordable procedure: not 'engagement' but 'minutes of active time on task per session', not 'skill' but 'score on this defined task'. A good operationalization has a definition you could apply the same way twice, produces something you can record, and is either validated or justified against the alternatives. The danger is choosing a measure because it is easy or because it flatters your hypothesis, rather than because it captures the thing.
Signature interactive
Operationalization Gauntlet
Here is a construct: stress. Before you see how anyone else does it, propose how you would measure it in a study, the actual instrument or procedure.
Place on the grid
Place each bathroom scale on the validity (accuracy) × reliability (consistency) grid.
Reliability = same answer each time. Validity = the right answer. A scale can be reliably wrong.
Accurate but inconsistent
Accurate & consistent
Inaccurate & inconsistent
Consistently wrong
Two properties, often confused. Validity is accuracy: does the measure capture the real thing? Reliability is consistency: does it give the same answer under the same conditions? They are independent: a measure can be reliably wrong (a scale always 3 kg high) or valid on average but noisy. Reliability is necessary but not sufficient; a consistent measure of the wrong construct is still the wrong measure. When you read a study, ask of every key variable: is this measure both accurate and consistent, and how do the authors know?
Select all that apply
A construct like 'engagement' is well operationalized when the measure... Select all that apply.
Concrete, recordable, and justified, not impressive-sounding or result-flattering.
Who you measure matters as much as what you measure. A sampling frame that omits whole groups (daytime workers, non-English speakers, people without a phone) biases every conclusion about the wider population, and a bigger biased sample is just as biased. Power is the flip side: an underpowered study often cannot detect a real effect, so a null result from it is uninformative, not 'small but suggestive'. Deciding the sample size, and committing to the analysis in advance through preregistration, are how honest studies avoid reading noise as signal. And any study touching human data carries ethical duties: consent, privacy, and IRB review where required.
Written response
Lab. Write a methods paragraph (about 150–200 words) for the research question you built in Lesson 4, complete enough that a stranger could run it. Cover: the population and how you would sample it; how you operationalize each variable (the actual measures); the design; and the analysis plan.
Population + sampling → measures (operationalized) → design → analysis. Write it so someone else could run it.
0/120 words
Checkpoint · item 1 of 5
A measure is reliable but not valid. What does that mean?
Reliable = same answer each time. Valid = the right answer.
Reflection