QCE Psychology - Unit 2 - Intelligence

Psychometric intelligence, test scores, validity and reliability

Learn psychometric intelligence, test scores, validity and reliability for QCE Psychology Unit 2 through a complete model, worked evidence and bounded evaluation.

Part of the free QCE Psychology notes library for Unit 2: Intelligence.

Updated 2026-08-13 - 7 min read

QCAA official coverage - Psychology 2025 v1.3

Exact syllabus points covered

  1. Describe the psychometric approach to intelligence (i.e. intelligence quotient, or IQ).
  2. Describe common methods by which intelligence is measured with reference to IQ tests and scales, including Stanford–Binet scale
  3. Describe common methods by which intelligence is measured with reference to IQ tests and scales, including Wechsler’s intelligence scales for adults (WAIS-IV) and children (WISC-V).
  4. Discuss the degree to which intelligence tests are valid and reliable.
  5. Consider the validity and reliability of IQ and EQ testing to determine if these tests can be misleading and/or inaccurate.

Explain how standardised intelligence scores are constructed and evaluate Stanford–Binet, WAIS-IV and WISC-V evidence using reliability, validity and uncertainty. This note develops the connected model and the evidence needed to use it, rather than reducing the syllabus to a list of terms.

Psychometric intelligence, test scores, validity and reliability diagram

Original Sylligence diagram for psychology u12 iq measurement.

Psychometric intelligence, test scores, validity and reliability diagram

Build the psychological model

The psychometric approach treats intelligence as measurable variation inferred from performance on standardised tasks. A test converts a pattern of responses into norm-referenced scores by comparing a person with an appropriate standardisation sample. An IQ score is therefore an estimate under specified conditions, not a direct quantity of intelligence or a permanent biological label. Reliability asks whether measurement is consistent; validity asks whether the interpretation and use of scores are justified.

Psychological science separates a construct from the way it is measured. The mechanism for this lesson is Observed performance reflects the construct plus systematic and random influences. The most useful evidence is Standard score, norm group, confidence interval and subtest pattern. Neither a construct label nor a brain image explains a result by itself; the response must show how an operational measure connects to a theory prediction.

Connect the ideas

1. The Stanford–Binet provides age-spanning cognitive assessment; Wechsler scales use age-appropriate batteries such as WAIS-IV for adults and WISC-V for children, with index and overall scores derived from multiple subtests

The Stanford–Binet provides age-spanning cognitive assessment; Wechsler scales use age-appropriate batteries such as WAIS-IV for adults and WISC-V for children, with index and overall scores derived from multiple subtests.

2. Standard scores often use a mean of 100 and standard deviation of 15, but interpretation requires current norms, administration rules, confidence intervals and attention to uneven profiles

Standard scores often use a mean of 100 and standard deviation of 15, but interpretation requires current norms, administration rules, confidence intervals and attention to uneven profiles.

3. Internal consistency, test–retest and inter-rater evidence address different reliability questions

Internal consistency, test–retest and inter-rater evidence address different reliability questions. Content, construct, criterion and predictive evidence address different aspects of validity; high reliability is necessary but not sufficient for validity.

These ideas should not be memorised as isolated theorist names or definitions. Compare the mechanism each account proposes, the observation it predicts, and the result that would count against it. When two theories can explain the same surface result, a stronger investigation changes a condition that makes their predictions diverge.

Trace the proposed mechanism

  1. Define the construct and intended decision, then sample tasks expected to represent relevant cognitive performance.
  2. Standardise administration and scoring so irrelevant variation is reduced, and compare performance with an appropriate normative sample.
  3. Estimate reliability and measurement error, expressing scores as a range rather than false point precision.
  4. Gather converging and discriminating validity evidence for the intended population and use, then monitor adverse or misleading consequences.

Read each arrow critically. Does it name an observed association, a proposed causal process or an interpretation? Those claims require different designs. The lesson's responsible conclusion is The estimate is above the norm mean but contains measurement uncertainty. Its boundary is equally important: A test score estimates performance relative to a norm group; it is not a complete person.

Worked evidence

The worked conclusion identifies what was measured before explaining it. It avoids mind-reading, biological determinism, diagnosis from classroom evidence and universal claims from a group average. If numerical evidence is supplied, use the value and comparison; if qualitative evidence is supplied, identify the coding or interpretive boundary.

Investigate it properly

Research question. Does a short reasoning measure show stable scores and predict a separate criterion in the intended student population?

Design. Administer parallel or repeated forms under standardised conditions, prespecify the criterion and interval, record relevant accessibility factors and separate reliability from criterion analysis.

Evidence. Report score distributions, test–retest or alternate-form association, measurement error, criterion relationship and subgroup patterns rather than one overall coefficient.

Limitation and improvement. Practice can inflate retest performance and a narrow criterion can make a test appear valid only because it resembles the test. Use suitable intervals, independent criteria and broader construct evidence.

A defensible investigation should use standardised administration and test reliability, validity and fairness. Reliability asks whether the evidence is consistent under comparable conditions. Validity asks whether the method supports the intended inference. A larger sample can improve precision, but it cannot repair a confound, an invalid measure or a conclusion that exceeds the design.

Ethical reasoning

Explain limits, protect results and avoid deterministic labelling. Ethical quality is not a paragraph added after the method. It shapes recruitment, consent, risk, privacy, withdrawal, data handling, reporting and the consequences of applying a finding to individuals or groups.

Repair the inference

Scores are norm-referenced estimates from sampled tasks. Consistency does not guarantee the intended construct is measured, and overlapping confidence intervals plus contextual effects make tiny rank differences unsafe to interpret.

The tempting overclaim is The person has exactly 115 units of fixed intelligence. Replace it with the supported inference and explicitly preserve a test score estimates performance relative to a norm group; it is not a complete person. Good psychological writing can be confident about a measured pattern while remaining cautious about mechanism, diagnosis and generalisation.

Transfer to an unfamiliar study

Audit an unfamiliar school or workplace test by asking who was normed, what decision is intended, which reliability and validity evidence exists, how uncertainty is reported and what harm may follow from misclassification.

Use this five-part routine:

  1. Define the construct and its operational measure.
  2. Identify the design, comparison and observed result.
  3. Trace or compare the proposed mechanism.
  4. Evaluate validity, reliability, sample, ethics and one alternative explanation.
  5. State a bounded conclusion that could be revised by new evidence.

Quick check

Syllabus coverage

This lesson develops the following current QCAA Psychology 2025 subject matter:

  • Describe the psychometric approach to intelligence (i.e. intelligence quotient, or IQ).
  • Describe common methods by which intelligence is measured with reference to IQ tests and scales, including Stanford–Binet scale
  • Describe common methods by which intelligence is measured with reference to IQ tests and scales, including Wechsler’s intelligence scales for adults (WAIS-IV) and children (WISC-V).
  • Discuss the degree to which intelligence tests are valid and reliable.
  • Consider the validity and reliability of IQ and EQ testing to determine if these tests can be misleading and/or inaccurate.

The official syllabus remains the authority for required subject matter. This note adds connected explanation, worked reasoning and evidence routines so that the statements can be learned and applied.

Sources

Finished reading? Practise this topic free

Open Psychology past questions with this Unit 2 topic carried into the question bank, then save your progress for the next review.

Practise this topic free. Free to start. No payment details are required. Exact question coverage depends on the available past-paper syllabus mapping.