Measurement, samples, and protocol

Lesson 6 36 minDesigning

In this lesson

By the end, you’ll be able to

  • Operationalize a construct into a concrete, defensible measure
  • Distinguish validity from reliability and judge a measure on both
  • Reason about sampling frames and why underpowered studies are not 'small but suggestive'

Sign in to save your progress

You can keep reading without an account, but completed lessons won’t be saved.

Sign in

Between a good question and real data sits a step that quietly decides whether a study means anything: measurement. A construct like 'stress', 'engagement', or 'skill' is an idea, not a number. Turning it into something you can record, well or badly, determines what your results are actually about. Get this wrong and every later statistic is precise nonsense.

validity can leak at every linkConstruct“stress”Operationaldefinitionhow you'll pin it downInstrumentthe actual measureDatathe numbers
Every study runs a chain from an abstract construct to numbers: define it, choose an instrument, collect data. Validity can leak at every link; the numbers can be flawless and still measure the wrong thing.

To operationalize a construct is to turn it into a specific, recordable procedure: not 'engagement' but 'minutes of active time on task per session', not 'skill' but 'score on this defined task'. A good operationalization has a definition you could apply the same way twice, produces something you can record, and is either validated or justified against the alternatives. The danger is choosing a measure because it is easy or because it flatters your hypothesis, rather than because it captures the thing.

Signature interactive

Operationalization Gauntlet

Here is a construct: stress. Before you see how anyone else does it, propose how you would measure it in a study, the actual instrument or procedure.

Place on the grid

Place each bathroom scale on the validity (accuracy) × reliability (consistency) grid.

Reliability = same answer each time. Validity = the right answer. A scale can be reliably wrong.

Validity
InconsistentConsistent

Accurate but inconsistent

    Accurate & consistent

      Inaccurate & inconsistent

        Consistently wrong

          InaccurateAccurate
          Shows your exact weight every time
          Always reads exactly 3 kg too high
          Jumps around your true weight, right on average
          Shows a different random number, unrelated to weight
          How sure?

          Two properties, often confused. Validity is accuracy: does the measure capture the real thing? Reliability is consistency: does it give the same answer under the same conditions? They are independent: a measure can be reliably wrong (a scale always 3 kg high) or valid on average but noisy. Reliability is necessary but not sufficient; a consistent measure of the wrong construct is still the wrong measure. When you read a study, ask of every key variable: is this measure both accurate and consistent, and how do the authors know?

          Select all that apply

          A construct like 'engagement' is well operationalized when the measure... Select all that apply.

          Concrete, recordable, and justified, not impressive-sounding or result-flattering.

          How sure?

          Who you measure matters as much as what you measure. A sampling frame that omits whole groups (daytime workers, non-English speakers, people without a phone) biases every conclusion about the wider population, and a bigger biased sample is just as biased. Power is the flip side: an underpowered study often cannot detect a real effect, so a null result from it is uninformative, not 'small but suggestive'. Deciding the sample size, and committing to the analysis in advance through preregistration, are how honest studies avoid reading noise as signal. And any study touching human data carries ethical duties: consent, privacy, and IRB review where required.

          Written response

          Lab. Write a methods paragraph (about 150–200 words) for the research question you built in Lesson 4, complete enough that a stranger could run it. Cover: the population and how you would sample it; how you operationalize each variable (the actual measures); the design; and the analysis plan.

          Population + sampling → measures (operationalized) → design → analysis. Write it so someone else could run it.

          0/120 words

          How sure?

          Checkpoint · item 1 of 5

          A measure is reliable but not valid. What does that mean?

          Reliable = same answer each time. Valid = the right answer.

          How sure?

          Reflection