TestsimulationAufgabe 1 von 21
Diese Aufgabe00:00
Restzeit Psychologieverständnis englisch36:00

Text

Text 1 — Construct Validity

Construct validity asks whether test scores behave as they should if they truly reflect the intended latent attribute. In classical terms, an observed score contains both the "true" attribute and random error. High reliability means the score is consistent, but not necessarily right. Two complementary lines of evidence are central. First, convergent validity: scores on the focal measure should correlate strongly with other well-founded indicators of the same attribute. Second, discriminant validity: the same scores should relate only weakly to measures of distinct constructs. The multitrait--multimethod (MTMM) approach makes this concrete: imagine measuring anxiety and depression with self-report, peer-report, and behavioral tasks. A credible pattern shows higher associations for anxiety--anxiety across different methods than for anxiety--depression within the same method (e.g., self-report), which would otherwise signal method bias. Factor-analytic models provide structured tests. In exploratory analysis, items intended to reflect a construct should cluster together: vocabulary, reading comprehension, and analogy items forming a verbal factor, while matrices and spatial rotation form a separate factor. In confirmatory analysis, a model with a latent "verbal ability" factor should fit better than a one-factor model that forces all items together. Key diagnostics are interpretable parameters: target items load strongly on their factor, cross-loadings are small, and any correlated residuals have a clear rationale (e.g., two nearly identical paraphrase items). A practical summary is how much item variance the factor accounts for---if a "mindfulness" factor explains little variance in its items, the construct interpretation is weak. Overusing post-hoc correlated errors (e.g., linking many items just because they share wording) can hide method artifacts and blur what the factor actually represents.

Threats come in two forms. Construct underrepresentation occurs when the measure samples too narrow a slice of the domain. A "leadership effectiveness" test that only asks about goal setting ignores influence, communication, and conflict management. Construct-irrelevant variance arises when scores reflect nuisance features. A math test that relies on long word problems may tap reading comprehension; a personality inventory with many positively keyed items invites acquiescent responding. Mitigations include balanced keying (half positive, half negative), adding marker variables for response styles, using multiple methods (e.g., supervisor ratings plus work samples), or modeling explicit method factors. Partialing out a measured nuisance can help, but only if the nuisance is well measured and truly separate (e.g., statistically controlling reading speed when evaluating numerical reasoning items that contain text).

Generalization is another pillar. If the same construct interpretation is claimed across groups or conditions, measurement should be invariant. For example, a depression scale used with adolescents and older adults should show the same pattern and strength of item--construct relations; otherwise, group differences in means or correlations may be artifacts. In item response terms, differential item functioning is a warning sign: if an item like "I enjoy social media" is easier to endorse for teens than for adults with the same latent depression level, scores are not comparable.

Finally, criterion-related evidence must be interpreted through the construct lens. A cognitive ability test might predict job performance because it proxies access to training, not because it measures the targeted ability. Conversely, weak prediction can stem from a noisy criterion (e.g., supervisor ratings with halo bias) or restricted variability (e.g., only top applicants). Sound validity arguments integrate structural evidence from measurement models, convergent and discriminant relations, experimental sensitivity (e.g., a stress manipulation raising test anxiety scores but not conscientiousness), and invariance across samples and settings. Test construction should follow theory: if subtests reflect distinct facets of "executive function" (updating, shifting, inhibition), weights should mirror their theoretical roles rather than simply maximizing internal consistency. Overly redundant items can inflate precision while eroding coverage. A disciplined validation program iterates: define the construct and its network, design indicators that sample its breadth while managing method variance, test models and invariance, check predicted relations with neighboring and distant constructs, and revise when patterns diverge for reasons that theory can explain.

Psychologieverständnis englisch

Within the multitrait--multimethod example, which correlation pattern most clearly indicates method bias rather than construct coherence?