The science of assessment

Methodologies

How influential intelligence tests present problems, measure performance, and turn responses into scores.

Intelligence tests sample reasoning through language, numbers, visual patterns, memory, and construction. Standardized directions define the task; scoring rules determine which responses receive credit. Timed group examinations, untimed reasoning measures, and individual performance tests each organize that process differently.

Group mental-ability testing

Otis Gamma

The Gamma level of Arthur S. Otis’s Quick-Scoring Mental Ability Tests was designed for high-school and college examinees. The 1954 revision used 80 items in one examination rather than separately timed subtests. Published administrations allowed 30 minutes of working time, with directions given before the timed period.

The revised series supported hand or machine scoring. Performance was interpreted through published score-conversion tables, including deviation IQ: a score expressing standing relative to a reference group.

Georgia Department of Education: testing-program guide (1962)Otis Gamma administration in the original research report

Self-administering group examination

Otis Higher Examination

The Higher Examination, Form A, presents 75 items in a single 30-minute session. Vocabulary, arithmetic, analogies, directions, and practical reasoning appear together. Printed instructions and worked examples establish how answers should be recorded, allowing a group to proceed through the same examination under common timing.

Each correct response earns one point. The revised manual relates this total to age norms and provides college percentile reference points. Its historical IQ calculation adds 100 to the raw score and subtracts the published norm for the examinee’s age.

Otis’s revised manual (1928), reproduced in Josserand’s appendix

Untimed reasoning and psychometric analysis

ICAR-16

The 16-item International Cognitive Ability Resource sample test contains four items from each of four families: letter and number series, matrix reasoning, verbal reasoning, and three-dimensional rotation. Examinees identify sequences, complete visual relationships, solve verbal problems, and recognize rotated objects. The initial online validation used untimed items.

Responses are keyed as correct or incorrect. Condon and Revelle evaluated the items through internal-consistency analysis, factor analysis, and item-response models. These methods examine how responses relate across items, how the item families share a general ability component, and how questions differ in difficulty and discrimination.

Condon and Revelle: development and initial validation (2014)Scored ICAR item data and documentation

Brief, speeded cognitive assessment

Wonderlic Personnel Test

E. F. Wonderlic developed a short group examination for personnel assessment. The classic format contains 50 questions with a 12-minute time limit. Word comparisons, number series, arithmetic, directions, geometric analysis, and logical problems sample several kinds of reasoning within a compact session.

Items progress in difficulty, and the working period emphasizes how many problems an examinee can answer correctly. The traditional total is the number of correct answers. Published reference data relate these totals to comparison groups and occupational uses.

Wonderlic: development of the Personnel TestResearch description of the Personnel Test’s format and scoring

Eight-section written battery

Army Alpha

Army Alpha was developed for group examination of U.S. Army recruits during the First World War. Its eight sections covered following directions, arithmetic, practical judgment, synonyms and antonyms, sentence ordering, number series, analogies, and general information. Examiners read prescribed instructions and controlled each section’s timing.

Most sections counted correct answers; the synonym–antonym and sentence-ordering sections subtracted incorrect answers. The eight section scores were added to a maximum of 212 and interpreted using the Army’s published letter-rating thresholds.

Yoakum and Yerkes: Army Mental Tests (1920), Alpha procedures

Seven-section nonverbal battery

Army Beta

Army Beta was designed for recruits who could not readily complete an English-language written examination. Its seven sections used mazes, cube counting, symbol sequences, digit–symbol coding, number checking, picture completion, and geometric construction. Examiners demonstrated the tasks through gestures and examples, then administered separately timed sections.

Scoring used task-specific rules, including partial credit for mazes, a weighting for coding, and an error deduction for number checking. Section scores were summed to a maximum of 118 and interpreted using published Army letter ratings.

Yoakum and Yerkes: Army Mental Tests (1920), Beta procedures

School-based group assessment

National Intelligence Tests, Scale A

Published under the National Research Council in 1920, the National Intelligence Tests applied group-testing methods to school assessment. Scale A combined arithmetic reasoning, sentence completion, logical selection, same–opposite judgments, and symbol–digit substitution. Each section began with practice exercises and had its own time limit.

Scoring used different weights across sections, partial credit in logical selection, and an error deduction in same–opposite judgments. The five scores were added into a total. The manual interpreted results through school-grade norms and age-based percentile comparisons, with adjustments for when during the school year testing occurred.

National Intelligence Tests: Scale A directions and scoring keys

Oral directions, visual responses, and memory

Dearborn Group Tests

Walter F. Dearborn’s 1920 Series I organized assessment for the first three school grades into examiner-led activities using pictures, markings, and oral directions. Color–form tasks required pupils to distinguish visual features and apply demonstrated response rules. The Ladders task presented number groups orally, which pupils recalled by marking numbered positions.

Examiners followed scripted demonstrations and task-specific pacing. The manual assigned different scoring rules to each activity, including deductions for incorrect color–form markings and credit for completely correct memory responses. Published grade-related standards provided a basis for interpreting examination totals.

Dearborn: Manual of Directions, Series I (1920)

Individual, age-graded examination

Stanford–Binet, 1916

Lewis Terman’s Stanford revision organizes individually administered tasks into age levels, from early childhood through adult groups. Vocabulary, comprehension, memory, numerical reasoning, and practical judgment are examined through spoken questions, pictures, objects, and drawing tasks. The examiner follows prescribed wording and scoring criteria, establishes a level at which all tasks are passed, and tests progressively higher levels.

Mental age combines this basal credit with additional months earned for successful tasks above it. For children, the intelligence quotient expresses mental age divided by chronological age, multiplied by 100.

Terman: The Measurement of Intelligence (1916), Chapter VIII

Individual construction and spatial analysis

Kohs Block Design

Samuel Kohs’s performance test uses sixteen colored cubes to reproduce a sequence of seventeen increasingly difficult designs. Each cube has solid-colored and diagonally divided faces. An examiner presents one design at a time, gives standardized directions or demonstrations, and records completion time and block movements.

Designs receive different point values, with deductions for excess time and moves; a design not completed within its time limit receives no credit. The total is interpreted through the manual’s mental-age conversion tables. The tasks require visual analysis, spatial organization, and the assembly of parts into a matching whole.

Kohs: Intelligence Measurement (1923), Chapter II