Research

Research

Research by Jiayi Deng on AI fairness, psychometric validity, subgroup comparability, rapid guessing, and evidence synthesis.

Jiayi's research connects psychometric validity, subgroup comparability, behavioral evidence, and responsible AI evaluation. The throughline is practical: build evidence that high-impact assessment and selection systems are fair, valid, interpretable, and appropriately monitored.

Research themes

AI fairness and human-AI evaluation

Evaluation of model outputs, end-to-end workflows, human overrides, counterfactual behavior, monitoring, and responsible-use boundaries in high-impact decision contexts.

Psychometric validity and subgroup comparability

Item response theory, measurement invariance, differential item/distractor functioning, reliability, equating, calibration, and fairness-oriented validation.

Behavioral and process data

Response time, rapid guessing, aberrant response patterns, cognitive labs, think-aloud protocols, and structured behavioral coding.

Evidence synthesis

Systematic review and meta-analysis, including current work on how AI-enabled interventions affect student motivation and engagement.

Selected publications

Selected peer-reviewed publications

Google Scholar

2026 · Educational Assessment

Optimizing measurement precision and diagnosticity for a two-dimensional assessment of reading comprehension

Biancarosa, G., Kennedy, P. C., Lee, S. B., DeWeese, J. N., Wong, Y. L., Deng, J., Weiss, D. J., & Davison, M. L. (2026).

Educational Assessment, 1-20.

2025 · Large-scale Assessments in Education

Linking errors introduced by rapid guessing responses when employing multigroup concurrent IRT scaling

Deng, J. (2025).

Large-scale Assessments in Education, 13(1), Article 28.

2025 · Educational Measurement: Issues and Practice

Digital Module 37: Introduction to item response tree (IRTree) models

Kim, N., Deng, J., & Wong, Y. L. (2025).

Educational Measurement: Issues and Practice, 44(1), 109-110.

2025 · Educational and Psychological Measurement

Is effort-moderated scoring robust to multidimensional rapid guessing?

Rios, J. A., & Deng, J. (2025).

Educational and Psychological Measurement, 85(1), 134-155. Online first 2024.

2024 · Applied Psychological Measurement

aberrance: An R package for detecting aberrant behavior in test data

Gorney, K., & Deng, J. (2024).

Applied Psychological Measurement, 48(4-5), 230-231.

2022 · Applied Psychological Measurement

Investigating the effect of differential rapid guessing on population invariance in equating

Deng, J., & Rios, J. A. (2022).

Applied Psychological Measurement, 46(7), 589-604.

View all peer-reviewed publications and software

Methods and software

Reproducible measurement and evaluation tools

Jiayi's work includes open-source statistical software, simulation studies, meta-analysis, response-time methods, and reproducible project artifacts such as model cards, dataset cards, monitoring plans, and decision logs.

aberrance is available through CRAN with package DOI 10.32614/CRAN.package.aberrance.

Human-AI Fairness Audit Lab

A synthetic, non-hiring audit demonstration that evaluates model behavior, reviewer overrides, calibration, counterfactual sensitivity, subgroup outcomes, and monitoring regressions.

View the case study

Historical publication, presentation, and teaching records remain available for continuity.