Automated Classification of Tax Service Requests Using Machine Learning and Natural Language Processing (#2732)
Read ArticleDate of Conference
July 15-17, 2026
Published In
"Engineering without Borders: Artificial Intelligence, Knowledge, Innovation, and Alliances for a Future from the Americas"
Location of Conference
Santiago (Chile)
Authors
Ruete, David
Fredes-Contreras, Iván
Salinas, Omar
Caroca, Alejandro
San Martin Medina, Lilian
Taramasco, Carla
Maidana, Jean
Abstract
Classifying citizen service requests in public tax offices is harder than it looks. Citizens describe their problems in their own words, select the wrong service category more often than not, and force staff to reclassify every submission before routing can begin. This paper examines whether that reclassification burden can be automated using NLP and supervised machine learning, drawing on 4,021 real requests submitted to Chile's Taxpayer Ombudsman Office (DEDECON) in Spanish free text. The feature engineering strategy combines TF-IDF text representation with two domain-grounded signals: a proxy for the user's knowledge level (scored 0 to 2) and binary indicators of class-specific keyword presence. Six classifiers were evaluated under identical conditions: Logistic Regression, Support Vector Machines, Naive Bayes, Decision Trees, Random Forest, and K-Nearest Neighbors. Random Forest reached 92.18% accuracy and an F1-macro score of 0.921, competitive with transformer-based approaches at a fraction of the computational cost. The main source of residual error is semantic overlap between adjacent service categories, a problem that text features alone are unlikely to resolve. Beyond DEDECON, the pipeline is directly applicable to other Spanish-language public service institutions facing the same intake classification problem.