Search papers, labs, and topics across Lattice.
This study explores the feasibility of using pre-trained sentence embeddings to recover item parameters in multidimensional item response theory, potentially eliminating the need for traditional calibration methods. By simulating a computerized adaptive testing (CAT) environment with the IPIP Big-Five dataset, the authors found that semantic embeddings can approximate latent trait profiles with a correlation of 0.825, comparable to the 0.857 achieved with fitted parameters. However, the embedding-based approach introduced significantly higher measurement uncertainty, attributed to collinearity among dimensions, indicating a need for careful evaluation of item banks before employing text-derived loadings.
Semantic embeddings can rival traditional calibration methods in accuracy but come with a hidden cost of increased measurement uncertainty.
Multidimensional item response theory relies on calibrated item parameters, such as discrimination and category threshold values, which are usually estimated from large samples of human test responses. This study investigates whether the directional loadings of these parameters can be recovered directly from item text using pre-trained sentence embeddings, avoiding the need for initial item calibration. Using the open-source IPIP Big-Five dataset ($n=19{,}719$; 50 items), we built a multidimensional computerized adaptive testing (CAT) simulation using D-optimal item selection. We compared three item loading sources: fitted graded response model parameters, semantic text embeddings, and a lexical baseline. In simulation, semantic embeddings recovered latent trait profiles almost as accurately as fitted parameters (correlation $0.825$ vs. $0.857$), performing noticeably better than simple word overlap ($0.752$). However, the embedding-based model produced inflated posterior variance, showing nearly four times higher measurement uncertainty despite accurate point estimates. We attribute this to collinearity across dimensions, as embedding-derived loadings pointed in similar directions across traits (condition number $137$ vs. $1.0$; mean trait cosine $0.90$). This outcome reflects the shared vocabulary common in personality items. We propose a simple diagnostic metric based on the loading matrix condition number to evaluate whether an item bank is suitable for text-derived loadings prior to testing.