[Research Contribution] Applying PhoBERT to Automatically Classify Student Support Requests at UEH

1 October, 2026

Keywords: PhoBERT; Artificial Intelligence; Natural Language Processing; University Digital Transformation; CRM/Ticket; Student Support; UEH.

Nearly 47,000 student support requests generated on the CRM/Ticket system of the University of Economics Ho Chi Minh City (UEH) have become a data source for a study applying artificial intelligence to university governance. Through testing and comparing multiple machine learning models, the research proposes using PhoBERT to automatically classify and route requests to the responsible units. The work was awarded the Best Paper Award at the International Conference on Next-Generation Data Engineering and Analytics (INDEA-2026), organized with the University of Salford, Manchester, United Kingdom, in August 2026.

Thumb Lớn Thương Hiệu Học Thuật Mới (2)

In the process of university digital transformation, data generated from operational activities can become an important resource for identifying issues and improving the quality of service to learners. At UEH, the CRM/Ticket system is deployed to receive student requests related to tuition fees, course registration, examinations, the library, facilities, and other services.

As the volume of requests continues to grow, the process of reading, classifying, and forwarding each ticket to the correct functional unit poses requirements for speed, consistency, and the ability to handle at scale. From this challenge, the research team comprising Dr. Nguyen Quoc Hung and Chau Quoc Long conducted the study “Applying the PhoBERT Language Model for Automated Classification of Student Support Tickets at the University of Economics Ho Chi Minh City (UEH)”.

From an operational challenge to a research question

Student support requests are typically expressed in natural language, with varying styles of expression, length, and levels of clarity. Some tickets have specific content, while others may address multiple issues simultaneously or fall into the intersection of functional responsibilities between units.

This makes the automatic identification of content and determination of the appropriate handling unit a complex natural language processing problem. The study therefore focuses on answering the question: Can a language model developed specifically for Vietnamese support the system in understanding request content and automatically classifying tickets with sufficient accuracy for practical application?

To answer this question, the research team developed a process comprising data cleaning and normalization; establishing text fields and classification labels; splitting training, validation, and test sets; deploying experimental models; comparing results; analyzing errors; and evaluating the ability to integrate into the existing system.

The study used data recorded on the UEH CRM/Ticket system during the period from June 2022 to October 31, 2025. After cleaning and normalization, the dataset comprised 46,874 tickets, labeled according to six functional units.

The direct exploitation of data generated during operations helps the research closely reflect the linguistic characteristics, support needs, and request routing mechanisms within the university environment. On this basis, the research team analyzed the distribution among ticket groups, identified content characteristics and potentially confusing cases, laying the groundwork for model selection and experimental design.

PhoBERT achieved the highest performance among the models surveyed

The study selected PhoBERT – a language model developed specifically for Vietnamese – as its focus, while deploying multiple comparison models to evaluate results on a comparative basis.

The traditional machine learning methods surveyed included Logistic Regression, Linear SVM, and Random Forest. In the Transformer group, the study tested mBERT, XLM-R, and PhoBERT.

The results showed that Transformer models achieved higher performance than traditional machine learning methods in the student support request classification task. Among the models surveyed, PhoBERT achieved an Accuracy of 0.944 and a Macro-F1 of 0.946, the highest in the study’s comparative results table. This provides the empirical basis for the research team to propose PhoBERT for the automatic request classification task at UEH.

Beyond overall metrics, the research team also analyzed the confusion matrix and mispredicted cases. The results showed that errors typically occur when ticket content overlaps between the responsibilities of multiple units or contains multiple intents simultaneously. Identifying these limitations helps determine which cases can be automated and which still require professional staff to review and adjust.

Toward integration into the UEH CRM/Ticket system

From the experimental results, the study selected the PhoBERT Weighted configuration as the proposed option, built a ticket classification demo, and outlined the direction for deploying the model as a RESTful API to connect with the CRM/Ticket system.

Under this approach, when a student submits a request, the model can analyze the content, predict the appropriate functional group, and support routing the ticket to the responsible unit. The solution is expected to contribute to shortening intake time, reducing manual classification workload, and enhancing consistency in the routing process.

During application, AI plays a supporting role in initial identification and classification. Cases with low confidence, containing multiple intents, or with unclear responsible units still need to be forwarded to professional staff for review. The mechanism of collaboration between the model and humans is a necessary condition for maintaining accuracy, controllability, and the quality of student support.

From the challenge on the CRM/Ticket system, the research forms a relatively complete value chain: practical data – scientific research – technological solution – application capability – academic contribution.

From a master’s project to an international research work

The research was developed from the master’s project “Proposing the PhoBERT Model in Automatic Classification of Student Support Requests at the University of Economics Ho Chi Minh City (UEH),” conducted by Chau Quoc Long under the academic supervision of Dr. Nguyen Quoc Hung. Based on the project’s results, the authors continued to refine the methodology, systematize the experimental results, and develop it into an international scientific paper.

In August 2026, the work was awarded the Best Paper Award at the International Conference on Next-Generation Data Engineering and Analytics – INDEA-2026, held on August 21–22, 2026, with the University of Salford, Manchester, United Kingdom. The award recognizes the scientific quality and applied significance of the research, while demonstrating that a specific issue in university operations can be developed into a work with methodology, measurable results, and value shared on an international academic forum.

For UEH, nearly 47,000 student support tickets, when standardized and researched using appropriate methods, have become the basis for developing governance support solutions, improving processes, and enhancing the learner experience.

The work thereby suggests a practical approach to university digital transformation: starting from needs in operational activities, exploiting data to identify issues, and developing evidence-based solutions. From UEH’s data, the research creates knowledge and opens the possibility of application back into practice, contributing to promoting governance innovation and improving the quality of services for learners.

Authors: Dr. Nguyen Quoc Hung, Chau Quoc Long – University of Economics Ho Chi Minh City.

This article is part of the knowledge dissemination series from UEH with the message “Research Contribution For All.” UEH respectfully invites readers to stay tuned for the next edition of the UEH Research Insights newsletter.

Chân Trang (1)

News, images: Authors, Department of Communication and Partnerships