My research is about individual differences in cognitive abilities, knowledge, and creativity, how they develop, and how well we can measure them. A recurring theme is that observed performance reflects more than one kind of variation: people differ from one another, the same person can perform differently across occasions and situations, and both matter when we use assessment to draw conclusions about ability or potential. Much of my current work studies these questions in childhood, although they are not restricted to children.

In childhood, many of these questions come together in the study of talent and its development. Cognitive potential, interests, motivation, self-regulation, creative activity, and the opportunities available to a child can all matter for what develops over time. This makes childhood a particularly useful setting for studying how stable individual differences, development, and shorter-term variation relate to one another, and what this means for psychological and educational measurement.

What I work on

  • Variation within a child

    Why does the same child show more of an ability or activity on one day than on another? In Within-Person Variability in Creative Activities Among Primary School Children: An Intensive Longitudinal Study, funded by the DFG, we follow the daily creative activities of primary…

    Why does the same child show more of an ability or activity on one day than on another? In Within-Person Variability in Creative Activities Among Primary School Children: An Intensive Longitudinal Study, funded by the DFG, we follow the daily creative activities of primary school children through parent reports over consecutive days. We examine how much variation occurs within children and how much between them, and whether daily changes are related to mood, motivation, self-regulation, and what the home environment offers on a particular day.

    The broader question is whether variation within a person contains information of its own. Rather than treating every fluctuation around a person’s average as measurement error, I am interested in when those fluctuations are systematic, what explains them, and whether repeated measurement changes the conclusions we draw about a person.

  • How differences develop

    How stable are individual differences over development, and how much can an early assessment tell us about what follows? Much of my current work on this question is connected to PINGUIN, a four-site project developing a tablet-based assessment of cognitive potential and early…

    How stable are individual differences over development, and how much can an early assessment tell us about what follows? Much of my current work on this question is connected to PINGUIN, a four-site project developing a tablet-based assessment of cognitive potential and early skills in language, literacy, and mathematics at school entry.

    The instrument is one part of the project. I am equally interested in what happens afterwards: how strongly early differences predict later learning and school performance, how these relations differ across children and backgrounds, how fair such assessments are, and how information from them can be made useful for teachers. More generally, I am interested in how individual differences emerge, change, and become more or less stable over time.

  • Interests and the choices they lead to

    Children do not only differ in what they can do, but also in what they want to do. Interests develop early and can shape which activities, subjects, and opportunities children move toward.

    Children do not only differ in what they can do, but also in what they want to do. Interests develop early and can shape which activities, subjects, and opportunities children move toward.

    With colleagues, I study how interest patterns develop during primary school, how interests and abilities jointly predict later choices, and what happens at the point where a choice is actually made. One example comes from a statewide enrichment program, where a course title and a few lines of description may be most of the information a family has when deciding whether to participate. We are studying how this framing affects course choice, and why girls and boys choose STEM-related courses at different rates.

  • Identification, assessment, and fairness

    How should potential be identified when performance is still developing, varies across situations, and is measured imperfectly? This is one of the places where my interests in individual differences and educational measurement come together.

    How should potential be identified when performance is still developing, varies across situations, and is measured imperfectly? This is one of the places where my interests in individual differences and educational measurement come together.

    I work on the validity and fairness of talent identification, the consequences of false-positive and false-negative decisions, socio-demographic differences in selection, and the evaluation of enrichment programs. Much of this work is connected to the Hector Children’s Academies, which we monitor and evaluate.

  • Creativity and its assessment

    Creativity has taken up an increasing part of my work in recent years, both as a substantive topic and as a particularly interesting measurement problem. On the substantive side, I study how creative potential relates to creative activity in everyday life and which cognitive,…

    Creativity has taken up an increasing part of my work in recent years, both as a substantive topic and as a particularly interesting measurement problem. On the substantive side, I study how creative potential relates to creative activity in everyday life and which cognitive, motivational, and contextual factors contribute to it.

    On the measurement side, creativity makes many familiar problems especially visible. Open-ended responses have to be scored, human ratings are costly, and results depend on how the rating process is designed. I have worked on planned missing designs for human ratings, on measures of creative activity for children and adolescents, and on the assessment and interpretation of creative thinking in PISA 2022.

  • Language data and automated scoring

    A good deal of what people know, think, or can produce appears as language rather than as a selected response. That makes text data useful, but also difficult to measure well.

    A good deal of what people know, think, or can produce appears as language rather than as a selected response. That makes text data useful, but also difficult to measure well.

    Together with colleagues, I develop procedures for automatically scoring open-ended and creative responses using large language models, in German, English, and increasingly multilingual settings. I am particularly interested in when automated scores capture the same information as human judgments, where they introduce new problems, and how such procedures can be used without losing sight of validity and fairness.

    A related direction I would like to develop further is the use of semantic networks to represent differences in knowledge and conceptual structure at the level of groups and individuals.

  • Measuring new competencies

    Some measurement problems arise because the construct itself is still taking shape. AI literacy is a good example.

    Some measurement problems arise because the construct itself is still taking shape. AI literacy is a good example. Article 4 of the EU AI Act requires organizations to ensure that staff have sufficient AI competence, but leaves open what that competence consists of and how it should be assessed.

    With colleagues, I work on how such broad requirements can be translated into measurable competencies and what evidence would be needed before an assessment can be used for practical decisions. This includes questions about performance-based assessment, validity, reliability, fairness, and the extent to which a measure remains useful as the technology itself changes.

Selected publications

All 42 publications