<< Back

Mathematics assessment in the era of large language models (#1786)

Read Article

Date of Conference

July 15-17, 2026

Published In

"Engineering without Borders: Artificial Intelligence, Knowledge, Innovation, and Alliances for a Future from the Americas"

Location of Conference

Santiago (Chile)

Authors

Quiroz-Chavil, Helga

Capuñay-Uceda, Oscar

Capuñay-Uceda, Carlos

Abstract

The rapid adoption of large language models (LLMs), such as ChatGPT, has introduced significant tensions in the assessment of mathematics in higher education, especially in engineering and science programs. Although these models demonstrate a high capacity for solving mathematical problems and generating coherent explanations, their integration challenges the validity of traditional assessment practices focused on the final product. This study presents a critical review of the literature on the use of LLMs in university mathematics assessment, based on a corpus of 27 articles indexed in Scopus, published between 2023 and 2026. A critical analysis methodology was employed, based on categories of pedagogical alignment, validity of evidence, assessment transformation, academic risks, and future projections. The results show that most studies use LLMs primarily as solution generators, while assessment instruments remain largely unchanged. Exam- and task-based assessments predominate, with little attention paid to process assessment, argumentation, or transfer. Likewise, methodological limitations in the empirical evidence and a predominantly declarative treatment of risks such as academic integrity and cognitive dependency are identified. Overall, the review highlights the need for a redesign of mathematics assessment that prioritizes evidence of reasoning, explanation, and verification, integrating LLMs in a critical and pedagogically aligned manner to preserve rigor, equity, and validity in higher education.

Read Article