PDF(1564 KB)
A Robust Parameter Estimation Method for Addressing Aberrant Responses in Computerized Adaptive Testing
Gao Xuliang, Zhao Ying, Wang Fang
Journal of Psychological Science ›› 2026, Vol. 49 ›› Issue (4) : 1011-1023.
PDF(1564 KB)
PDF(1564 KB)
A Robust Parameter Estimation Method for Addressing Aberrant Responses in Computerized Adaptive Testing
Aberrant responses, including random guessing, cheating, and mistakes caused by stress, fatigue, carelessness, or misreading, pose significant challenges in educational and psychological assessments. These responses are particularly problematic in Computerized Adaptive Testing (CAT) compared with traditional paper-and-pencil tests. In CAT, examinees cannot revise their answers, and their ability levels are dynamically estimated based on their responses to individual items. This adaptive nature means that early responses can substantially influence the trajectory of subsequent item selection, making CAT especially vulnerable to the effects of aberrant responses. Once these irregular responses occur, they can distort the ability estimation process, leading to a cascade of inappropriate item selections throughout the test.
When examinees produce aberrant responses, considerable deviations may arise in the estimation of their ability parameters. These distortions disrupt the adaptive testing process and may cause the system to select items that are either too difficult or too easy based on inaccurate estimates. Such misalignment compromises the precision of item selection and diminishes the overall validity of the assessment. As these errors accumulate, they can severely weaken the reliability and accuracy of the CAT process. Therefore, effectively managing aberrant responses is essential to maintaining the integrity and effectiveness of CAT systems.
Robust parameter estimation methods are vital for mitigating the adverse effects of aberrant responses. Such methods seek to minimize the influence of anomalous data, thereby improving the reliability of ability estimation and enhancing the overall accuracy of the test. However, research on robust estimation techniques specifically designed to address aberrant responses in CAT remains relatively limited. Therefore, developing more effective and resilient estimation strategies is crucial for advancing the performance and precision of CAT systems.
To address this limitation, a novel method called Weighted Maximum A Posteriori (WMAP) was proposed. Simulation results show that WMAP substantially enhances the accuracy of ability parameter estimation for examinees exhibiting aberrant responses. Compared with traditional estimation approaches, WMAP effectively mitigates the adverse effects of aberrant data, yielding more accurate ability estimates. Notably, it also improves estimation precision for examinees with normal response patterns, offering a dual advantage. This improvement stems from WMAP’s innovative weighting mechanism, which assigns greater weight to responses from items that better correspond to the examinee’s ability level. This feature is particularly beneficial during the early stages of CAT, when item selection may not yet be fully aligned with the examinee’s true ability.
Further validation with empirical data confirms the robust performance of WMAP in real-world testing contexts. Compared with traditional methods, WMAP is more effective in identifying and correcting biases caused by aberrant responses. By minimizing their adverse impact on ability estimation, WMAP produces more accurate and reliable test results. Its weighting mechanism not only diminishes the influence of aberrant responses but also promotes more appropriate item selection aligned with the examinee’s true ability. Consequently, WMAP enhances both the precision of ability estimation and the adaptability of the CAT system, resulting in more reliable assessment outcomes across diverse testing scenarios.
Beyond its immediate applications, WMAP represents a significant advancement in developing robust methodologies for CAT. Its value goes beyond mitigating aberrant responses. Future research could integrate WMAP into response-time models to gain deeper insights into examinee behavior and further reduce measurement error. Moreover, WMAP could be further adapted to handle more complex patterns of aberrant responding, such as those stemming from emotional fluctuations or cognitive fatigue, thereby extending its applicability and practical value. In parallel, complementary advances in item selection strategies could amplify the benefits of WMAP. Dynamically adjusting selection algorithms can account for potential estimation biases, such strategies can minimize errors and achieve more precise item-examinee matching, even under conditions of aberrant behavior. The integration of WMAP with adaptive item selection approaches would significantly enhance the overall performance of CAT systems, particularly in high-stakes testing environments where precision, reliability, and fairness are paramount.
computerized adaptive testing / aberrant responses / robust parameter estimation
| [1] |
陈平, 丁树良. (2008). 允许检查并修改答案的计算机化自适应测验. 心理学报, 40(6), 737-747.
|
| [2] |
陈平, 辛涛. (2011). 认知诊断计算机化自适应测验中的项目增补. 心理学报, 43(7), 836-850.
|
| [3] |
张龙飞, 王晓雯, 蔡艳, 涂冬波. (2020). 心理与教育测验中异常反应侦查新技术: 变点分析法. 心理科学进展, 28(9), 1462-1477.
变点分析法(change point analysis, CPA)近些年才引入心理与教育测量学, 相较于传统方法, CPA不仅可以侦查异常作答被试, 还能自动精确地定位变点位置, 高效清洗作答数据。其原理在于:判断作答序列中是否存在可将该序列划分为具有不同统计学属性两部分的点(即变点), 并且需使用被试拟合统计量(person-fit statistic, PFS)来量化两个子序列之间的差异。未来可将单变点分析拓展至多变点, 结合反应时等信息, 构建非参数化指标以及将现有指标拓展至多级计分或多维测验, 以提高CPA的适用广度及效力。
|
| [4] |
张心, 涂冬波. (2014). 计算机化自适应测验中几种常用能力估计方法的特性与评价. 中国考试, 5, 18-25.
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
Item compromise persists in undermining the integrity of testing, even secure administrations of computerized adaptive testing (CAT) with sophisticated item exposure controls. In ongoing efforts to tackle this perennial security issue in CAT, a couple of recent studies investigated sequential procedures for detecting compromised items, in which a significant increase in the proportion of correct responses for each item in the pool is monitored in real time using moving averages. In addition to actual responses, response times are valuable information with tremendous potential to reveal items that may have been leaked. Specifically, examinees that have preknowledge of an item would likely respond more quickly to it than those who do not. Therefore, the current study proposes several augmented methods for the detection of compromised items, all involving simultaneous monitoring of changes in both the proportion correct and average response time for every item using various moving average strategies. Simulation results with an operational item pool indicate that, compared to the analysis of responses alone, utilizing response times can afford marked improvements in detection power with fewer false positives.
|
| [11] |
The clinical assessment of mental disorders can be a time-consuming and error-prone procedure, consisting of a sequence of diagnostic hypothesis formulation and testing aimed at restricting the set of plausible diagnoses for the patient. In this article, we propose a novel computerized system for the adaptive testing of psychological disorders. The proposed system combines a mathematical representation of psychological disorders, known as the "formal psychological assessment," with an algorithm designed for the adaptive assessment of an individual's knowledge. The assessment algorithm is extended and adapted to the new application domain. Testing the system on a real sample of 4,324 healthy individuals, screened for obsessive-compulsive disorder, we demonstrate the system's ability to support clinical testing, both by identifying the correct critical areas for each individual and by reducing the number of posed questions with respect to a standard written questionnaire.
|
| [12] |
|
| [13] |
The focus of this paper is on the improvement of substance use disorder (SUD) screening and measurement. Using a multi-dimensional item response theory model, the bifactor model, we provide a psychometric harmonization between SUD, depression, anxiety, trauma, social isolation, functional impairment and risk-taking behavior symptom domains, providing a more balanced view of SUD. The aims are to (1) develop the item-bank, (2) calibrate the item-bank using a bifactor model that includes a primary dimension and symptom-specific subdomains, (3) administer using computerized adaptive testing (CAT) and (4) validate the CAT-SUD in Spanish and English in the United States and Spain.Item bank construction, item calibration phase, CAT-SUD validation phase.Primary care, community clinics, emergency departments and patient-to-patient referrals in Spain (Barcelona and Madrid) and the United States (Boston and Los Angeles).Calibration phase: the CAT-SUD was developed via simulation from complete item responses in 513 participants. Validation phase: 297 participants received the Composite International Diagnostic Interview (CIDI) and the CAT-SUD.A total of 252 items from five subdomains: (1) SUD, (2) psychological disorders, (3) risky behavior, (4) functional impairment and (5) social support. CAT-SUD scale scores and CIDI SUD diagnosis.Calibration: the bifactor model provided excellent fit to the multi-dimensional item bank; 168 items had high loadings (> 0.4 with the majority > 0.6) on the primary SUD dimension. Using an average of 11 items (four to 26), which represents a 94% reduction in respondent burden (average administration time of approximately 2 minutes), we found a correlation of 0.91 with the 168-item scale (precision of 5 points on a 100-point scale).strong agreement was found between the primary CAT-SUD dimension estimate and the results of a structured clinical interview. There was a 20-fold increase in the likelihood of a CIDI SUD diagnosis across the range of the CAT-SUD (AUC = 0.85).We have developed a new approach for the screening and measurement of SUD and related severity based on multi-dimensional item response theory. The bifactor model harmonized information from mental health, trauma, social support and traditional SUD items to provide a more complete characterization of SUD. The CAT-SUD is highly predictive of a current SUD diagnosis based on a structured clinical interview, and may be predictive of the development of SUD in the future.© 2020 Society for the Study of Addiction.
|
| [14] |
A plausible s-factor solution for many types of psychological and educational tests is one that exhibits a general factor and s − 1 group or method related factors. The bi-factor solution results from the constraint that each item has a nonzero loading on the primary dimension and at most one of the s − 1 group factors. This paper derives a bi-factor item-response model for binary response data. In marginal maximum likelihood estimation of item parameters, the bi-factor restriction leads to a major simplification of likelihood equations and (a) permits analysis of models with large numbers of group factors; (b) permits conditional dependence within identified subsets of items; and (c) provides more parsimonious factor solutions than an unrestricted full-information item factor analysis in some cases.
|
| [15] |
|
| [16] |
Participant attentiveness is a concern for many researchers using Amazon's Mechanical Turk (MTurk). Although studies comparing the attentiveness of participants on MTurk versus traditional subject pool samples have provided mixed support for this concern, attention check questions and other methods of ensuring participant attention have become prolific in MTurk studies. Because MTurk is a population that learns, we hypothesized that MTurkers would be more attentive to instructions than are traditional subject pool samples. In three online studies, participants from MTurk and collegiate populations participated in a task that included a measure of attentiveness to instructions (an instructional manipulation check: IMC). In all studies, MTurkers were more attentive to the instructions than were college students, even on novel IMCs (Studies 2 and 3), and MTurkers showed larger effects in response to a minute text manipulation. These results have implications for the sustainable use of MTurk samples for social science research and for the conclusions drawn from research with MTurk and college subject pool samples.
|
| [17] |
Self-report data are common in psychological and survey research. Unfortunately, many of these samples are plagued with careless responses, due to unmotivated participants. The purpose of this study was to propose and evaluate a robust estimation method to detect careless or unmotivated responders, while leveraging item response theory (IRT) person-fit statistics. First, we outlined a general framework for robust estimation specific for IRT models. Subsequently, we conducted a simulation study covering multiple conditions in order to evaluate the performance of the proposed method. Ultimately, we showed that robust maximum marginal likelihood (RMML) estimation significantly improves detection rates for careless responders and reduces bias in item parameters across conditions. Furthermore, we applied our method to a real data set, to illustrate the utility of the proposed method. Our findings suggest that robust estimation coupled with person-fit statistics offers a powerful procedure to identify careless respondents for further review and to provide more accurate item parameter estimates in the presence of careless responses.
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
Computerized adaptive testing (CAT) is an active current research field in psychometrics and educational measurement. However, there is very little software available to handle such adaptive tasks. The R package catR was developed to perform adaptive testing with as much flexibility as possible, in an attempt to provide a developmental and testing platform to the interested user. Several item-selection rules and ability estimators are implemented. The item bank can be provided by the user or randomly generated from parent distributions of item parameters. Three stopping rules are available. The output can be graphically displayed.
|
| [22] |
|
| [23] |
When data are collected via anonymous Internet surveys, particularly under conditions of obligatory participation (such as with student samples), data quality can be a concern. However, little guidance exists in the published literature regarding techniques for detecting careless responses. Previously several potential approaches have been suggested for identifying careless respondents via indices computed from the data, yet almost no prior work has examined the relationships among these indicators or the types of data patterns identified by each. In 2 studies, we examined several methods for identifying careless responses, including (a) special items designed to detect careless response, (b) response consistency indices formed from responses to typical survey items, (c) multivariate outlier analysis, (d) response time, and (e) self-reported diligence. Results indicated that there are two distinct patterns of careless response (random and nonrandom) and that different indices are needed to identify these different response patterns. We also found that approximately 10%-12% of undergraduates completing a lengthy survey for course credit were identified as careless responders. In Study 2, we simulated data with known random response patterns to determine the efficacy of several indicators of careless response. We found that the nature of the data strongly influenced the efficacy of the indices to identify careless responses. Recommendations include using identified rather than anonymous responses, incorporating instructed response items before data collection, as well as computing consistency indices and multivariate outlier analysis to ensure high-quality data.
|
| [24] |
|
| [25] |
Maximum likelihood estimates of subjects' abilities in item response curve models are overly sensitive to disturbances that are common in educational testing, such as carelessness and random guessing. Recently Waller (1974) and Wainer and Wright (1980) have suggested more robust methods of estimation. This paper presents an alternative estimator, based on the principle of Tukey's biweight—a robust estimator of location. Among the advantages of the biweight estimator of latent ability are the following: (1) the nature and extent of disturbances are not assumed to be the same for all subjects; (2) each response is utilized in proportion to its apparent value; (3) the biweight estimate agrees with the maximum likelihood estimate when no disturbances are present; and (4) computation requires only a minor modification of the Newton-Raphson method commonly used to obtain maximum likelihood estimates of ability. The biweight estimator is demonstrated with artificial data, and shown to produce smaller meansquared errors than the maximum likelihood estimator even in the presence of only slight disturbances.
|
| [26] |
|
| [27] |
\n Background:\n The Patient-Reported Outcomes Measurement Information System Upper Extremity (PROMIS UE) computer adaptive test was developed to improve precision and reduce question burden. We hypothesized that in patients with carpal tunnel syndrome (CTS): (1) PROMIS UE would correlate with established patient-reported outcome measures (PROs); (2) the time and number of questions required would be lower than current metrics; (3) there would be no floor or ceiling effects; and (4) PROMIS UE would not correlate with disease severity.\n Methods:\n Patients undergoing electrodiagnostic evaluation found to have a primary diagnosis of unilateral CTS prospectively completed PROMIS UE, Quick Disabilities of the Arm, Shoulder and Hand (qDASH), and Boston Carpal Tunnel Syndrome Questionnaire (BCTQ). Electrophysiologic and clinical severity was recorded. The relationships among PROs were described with Spearman coefficients. A floor or ceiling effect was confirmed if >15% of patients achieved the lowest or highest possible score, respectively.\n Results:\n Fifty-one patients (average, 53.9 years) were enrolled. An excellent correlation was identified between PROMIS UE and qDASH (\n R\n = −0.76,\n P\n <.001). There was a good correlation between PROMIS UE and BCTQ (\n R\n = −0.58,\n P\n < 0.001). The PROMIS UE required less time and fewer questions than qDASH and BCTQ (\n P\n =.02 and\n P\n <.001). There were no floor or ceiling effects. Neither neurophysiologic nor clinical severity correlated with PROMIS UE (\n R\n = 0.24,\n P\n >.05 and\n R\n = −0.18,\n P\n >.05).\n Conclusions:\n The PROMIS UE has an excellent correlation with qDASH and a good correlation with BCTQ in patients with CTS. Furthermore, PROMIS UE required less time and fewer questions than established PROs. Used as a single PRO, PROMIS UE represents a practical alternative to current metrics in patients with CTS.\n
|
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
|
/
| 〈 |
|
〉 |