- Book Chapter
20
- 10.1016/s0169-7161(06)26025-5
25 Statistical Aspects of Adaptive Testing
- Jan 01, 2006
- Handbook of Statistics
- Wim J Van Der Linden + 1 more +1
25 Statistical Aspects of Adaptive Testing
To increase the number of items available for adaptive testing and reduce the cost of item writing, the use of techniques of item cloning has been proposed. An important consequence of item cloning is possible variability between the item parameters. To deal with this variability, a multilevel item response (IRT) model is presented which allows for differences between the distributions of item parameters of families of item clones. A marginal maximum likelihood and a Bayesian procedure for estimating the hyperparameters are presented. In addition, an item-selection procedure for computerized adaptive testing with item cloning is presented which has the following two stages: First, a family of item clones is selected to be optimal at the estimate of the person parameter. Second, an item is randomly selected from the family for administration. Results from simulation studies based on an item pool from the Law School Admission Test (LSAT) illustrate the accuracy of these item pool calibration and adaptive testing procedures. Index terms: computerized adaptive testing, item cloning, multilevel item response theory, marginal maximum likelihood, Bayesian item selection.
25 Statistical Aspects of Adaptive Testing
25 Statistical Aspects of Adaptive Testing
Fitting a Polytomous Item Response Model to Likert-Type Data
This study examined the application of the MML-EM algorithm to the parameter estimation problems of the normal ogive and logistic polytomous response models for Likert-type items. A rating-scale model was devel oped based on Samejima's (1969) graded response model. The graded response model includes a separate slope parameter for each item and an item response parameter. In the rating-scale model, the item re sponse parameter is resolved into two parameters: the item location parameter, and the category threshold parameter characterizing the boundary between re sponse categories. For a Likert-type questionnaire, where a single scale is employed to elicit different re sponses to the items, this item response model is ex pected to be more useful for analysis because the item parameters can be estimated separately from the threshold parameters associated with the points on a single Likert scale. The advantages of this type of model are shown by analyzing simulated data and data from the General Social Surveys. Index terms: EM algorithm, General Social Surveys, graded response model, item response model, Likert scale, marginal maximum likelihood, polytomous item response model, rating-scale model.
Read moreConstructing Rotating Item Pools for Constrained Adaptive Testing
Preventing items in adaptive testing from being over‐ or underexposed is one of the main problems in computerized adaptive testing. Though the problem of overexposed items can be solved using a probabilistic item‐exposure control method, such methods are unable to deal with the problem of underexposed items. Using a system of rotating item pools, on the other hand, is a method that potentially solves both problems. In this method, a master pool is divided into (possibly overlapping) smaller item pools, which are required to have similar distributions of content and statistical attributes. These pools are rotated among the testing sites to realize desirable exposure rates for the items. A test assembly model, motivated by Gulliksen's matched random subtests method, was explored to help solve the problem of dividing a master pool into a set of smaller pools. Different methods to solve the model are proposed. An item pool from the Law School Admission Test was used to evaluate the performances of computerized adaptive tests from systems of rotating item pools constructed using these methods.
Read moreA Comparison of Item‐Selection Methods for Adaptive Tests with Content Constraints
In test assembly, a fundamental difference exists between algorithms that select a test sequentially or simultaneously. Sequential assembly allows us to optimize an objective function at the examinee's ability estimate, such as the test information function in computerized adaptive testing. But it leads to the non‐trivial problem of how to realize a set of content constraints on the test—a problem more naturally solved by a simultaneous item‐selection method. Three main item‐selection methods in adaptive testing offer solutions to this dilemma. The spiraling method moves item selection across categories of items in the pool proportionally to the numbers needed from them. Item selection by the weighted‐deviations method (WDM) and the shadow test approach (STA) is based on projections of the future consequences of selecting an item. These two methods differ in that the former calculates a projection of a weighted sum of the attributes of the eventual test and the latter a projection of the test itself. The pros and cons of these methods are analyzed. An empirical comparison between the WDM and STA was conducted for an adaptive version of the Law School Admission Test (LSAT), which showed equally good item‐exposure rates but violations of some of the constraints and larger bias and inaccuracy of the ability estimator for the WDM.
Read moreStratified item selection and exposure control in unidimensional adaptive testing in the presence of two-dimensional data.
It is not uncommon to use unidimensional item response theory (IRT) models to estimate ability in multidimensional data. Therefore it is important to understand the implications of summarizing multiple dimensions of ability into a single parameter estimate, especially if effects are confounded when applied to computerized adaptive testing (CAT). Previous studies have investigated the effects of different IRT models and ability estimators by manipulating the relationships between item and person parameters. However, in all cases, the maximum information criterion was used as the item selection method. Because maximum information is heavily influenced by the item discrimination parameter, investigating a-stratified item selection methods is tenable. The current Monte Carlo study compared maximum information, a-stratification, and a-stratification with b blocking item selection methods, alone, as well as in combination with the Sympson-Hetter exposure control strategy. The six testing conditions were conditioned on three levels of interdimensional item difficulty correlations and four levels of interdimensional examinee ability correlations. Measures of fidelity, estimation bias, error, and item usage were used to evaluate the effectiveness of the methods. Results showed either stratified item selection strategy is warranted if the goal is to obtain precise estimates of ability when using unidimensional CAT in the presence of two-dimensional data. If the goal also includes limiting bias of the estimate, Sympson-Hetter exposure control should be included. Results also confirmed that Sympson-Hetter is effective in optimizing item pool usage. Given these results, existing unidimensional CAT implementations might consider employing a stratified item selection routine plus Sympson-Hetter exposure control, rather than recalibrate the item pool under a multidimensional model.
Read moreAn Evaluation of a Markov Chain Monte Carlo Method for the Rasch Model
The accuracy of the Gibbs sampling Markov chain monte carlo procedure was examined for estimating item and person ( .) parameters in the one-parameter logistic model. Four datasets were analyzed using the Gibbs sampling method, conditional maximum likelihood, marginal maximum likelihood, and joint maximum likelihood. Maximum likelihood and expected a posteriori. estimation methods were used with marginal maximum likelihood estimation of item parameters. Item parameter estimates from the four methods were almost identical;. estimates from Gibbs sampling were similar to those obtained from the expected a posteriori method.
Read moreThe Data-augmentation Techniques in Item Response Modeling: Current Approaches and New Developments
在心理与教育测量中,项目反应理论(Item Response Theory,IRT)模型的参数估计方法是理论研究与实践应用的基本工具。最近,由于IRT模型的不断扩展与EM(expectation-maximization)算法自身的固有问题,参数估计方法的改进与发展显得尤为重要。这里介绍了IRT模型中边际极大似然估计的发展,提出了它的阶段性特征,即联合极大似然估计阶段、确定性潜在心理特质“填补”阶段、随机潜在心理特质“填补”阶段,重点阐述了它的潜在心理特质“填补”(dataaugmentation)思想。EM算法与Metropolis-Hastings Robbins—Monro(MH-RM)算法作为不同的潜在心理特质“填补”方法,都是边际极大似然估计的思想跨越。目前,潜在心理特质“填补”的参数估计方法仍在不断发展与完善。
Read moreA General Approach to Algorithmic Design of Fixed-Form Tests, Adaptive Tests, and Testlets
The selection of items from a calibrated item bank for fixed-form tests is an optimal test design problem; this problem has been handled in the literature by mathematical programming models. A similar problem, however, arises when items are selected for an adaptive test or for testlets. This paper focuses on the similarities of optimal design of fixed-form tests, adaptive tests, and testlets within the framework of the general theory of optimal designs. A sequential design procedure is proposed that uses these similarities. This procedure not only enables optimal design of fixed-form tests, adaptive tests, and testlets, but is also very flexible. The procedure is easy to apply, and consistent estimates for the trait level distribution are obtained. Index terms: adaptive tests, consistency, efficiency, optimal test design, sequential procedure, test design, testlets.
Read moreEstimating Diffusion-Based Item Response Theory Models: Exploring the Robustness of Three Old and Two New Estimators
Diffusion-based item response theory models for responses and response times in tests have attracted increased attention recently in psychometrics. Analyzing response time data, however, is delicate as response times are often contaminated by unusual observations. This can have serious effects on the validity of statistical inference. In this article, we compare three established and two new estimation approaches for diffusion-based item response theory models with respect to their robustness. The three established approaches are the marginal maximum likelihood (ML) estimator for continuous time, the marginal ML estimator for discrete time, and the weighted least squares (WLS) estimator. The new approaches are two modifications of the WLS estimator with better robustness properties. The performance of the estimators is compared in a simulation study. The simulation study illustrates that the new approaches are robust against some forms of random independent contamination. The marginal ML estimator for discrete time also performs well. The marginal ML estimator for continuous time is heavily affected by contamination.
Read moreItem Response Models in Computerized Adaptive Testing: A Simulation Study
In the digital world, any conceptual assessment framework faces two main challenges: (a) the complexity of knowledge, capacities and skills to be assessed; (b) the increasing usability of web-based assessments, which requires innovative approaches to the development, delivery and scoring of tests. Statistical methods play a central role in such framework. Item response models have been the most common statistical methods used to address such kind of measurement challenges, and they have been used in computer-based adaptive tests, which allow the item selection adaptively, from an item pool, according to the person ability during test administration. The test is tailored to each student. In this paper we conduct a simulation study based on the minimum error-variance criterion method varying the item exposure rate (0.1, 0.3, 0.5) and the test maximum length (18, 27, 36). The comparison is done by examining the absolute bias, the root mean square-error, and the correlation. Hypotheses tests are applied to compare the true and estimated distributions. The results suggest the considerable reduction of bias as the number of item administered increases, the occurrence of ceiling effect in very small size tests, the full agreement between true and empirical distributions for computerized tests of length smaller than the paper-and-pencil tests.KeywordsItem response modelcomputerized adaptive testingmeasurement error
Read moreAdaptive nonparametric tests for a single sample location problem
Adaptive nonparametric tests for a single sample location problem
Calibration of an item pool for assessing the burden of headaches: an application of item response theory to the headache impact test (HIT).
Measurement of headache impact is important in clinical trials, case detection, and the clinical monitoring of patients. Computerized adaptive testing (CAT) of headache impact has potential advantages over traditional fixed-length tests in terms of precision, relevance, real-time quality control and flexibility. To develop an item pool that can be used for a computerized adaptive test of headache impact. We analyzed responses to four well-known tests of headache impact from a population-based sample of recent headache sufferers (n = 1016). We used confirmatory factor analysis for categorical data and analyses based on item response theory (IRT). In factor analyses, we found very high correlations between the factors hypothesized by the original test constructers, both within and between the original questionnaires. These results suggest that a single score of headache impact is sufficient. We established a pool of 47 items which fitted the generalized partial credit IRT model. By simulating a computerized adaptive health test we showed that an adaptive test of only five items had a very high concordance with the score based on all items and that different worst-case item selection scenarios did not lead to bias. We have established a headache impact item pool that can be used in CAT of headache impact.
Read moreSensitivity of Marginal Maximum Likelihood Estimation of Item and Ability Parameters to the Characteristics of the Prior Ability Distributions
The sensitivity of marginal maximum likelihood es timation of item and ability (θ) parameters was ex amined when the prior θ distributions are not matched to the underlying θ distributions. Thirty sets of 45-item test data were generated by specifi cation of three types of underlying θ distributions. They were then analyzed with PC-BILOG. Appropri ate specification of the prior θ distribution increased the accuracy of estimation for item and θ param eters when the sample size was large. With a small dataset, the appropriate specification of the prior increased the accuracy of θ parameter estimation, but it did not have that effect on item parameter esti mation. Only with a large dataset and matched under lying and prior θ distributions did increasing the number of quadrature points improve the accuracy of estimation of the item parameters. However, the ac curacy of θ estimation was increased by increasing the number of quadrature points, regardless of sample size and appropriateness of the prior θ distribution. The number of examinees had an im portant effect on the accuracy of item parameter estimation.
Read moreAssembling a Computerized Adaptive Testing Item Pool as a Set of Linear Tests
Test-item writing efforts typically results in item pools with an undesirable correlational structure between the content attributes of the items and their statistical information. If such pools are used in computerized adaptive testing (CAT), the algorithm may be forced to select items with less than optimal information, that violate the content constraints, and/or have unfavorable exposure rates. Although at first sight somewhat counterintuitive, it is shown that if the CAT pool is assembled as a set of linear test forms, undesirable correlations can be broken down effectively. It is proposed to assemble such pools using a mixed integer programming model with constraints that guarantee that each test meets all content specifications and an objective function that requires them to have maximal information at a well-chosen set of ability values. An empirical example with a previous master pool from the Law School Admission Test (LSAT) yielded a CAT with nearly uniform bias and mean-squared error functions for the ability estimator and item-exposure rates that satisfied the target for all items in the pool.
Read moreA multilevel item response theory model was investigated for longitudinal vision-related quality-of-life data
A multilevel item response theory model was investigated for longitudinal vision-related quality-of-life data