Home

  Editors

  Ethics

  Submission

  Volumes

  Indexing

  Copyright

  Fees

  Subscription

  Publisher

  Support

  EPPM

 

Journal of Engineering, Project, and Production Management, 2026, 16(5), 2026-0092

 

Fine-Tuning Large Language Models for Automated Resume Screening in Human Resource Management: A Comparative Performance Analysis

 

Xin Liu1 and Chunfang Su2

1 Associate Senior Engineer, Information Data Resources Management Department, Qingdao Human Resources Development Research and Promotion Center, Qingdao, Shandong Province, 266001, China,
E-mail: tiaw2623@outlook.com (corresponding author).
2 Associate Professor, Department of Computer Science, Jiangyin Polytechnic College, Wuxi, Jiangsu Province, China

 

Project Management

 

Received June 6, 2026; revised June 30, 2026; accepted July 10, 2026

 

Available online July 28, 2026

 

Abstract:  In today's fast‑paced recruitment landscape, the growing number of online applications has pushed manual resume screening to its limits, creating significant challenges in time, cost, and consistency for human resource professionals. Although classical machine-learning pipelines and earlier transformer baselines partially automate this task, they struggle with the unstructured nature of curricula vitae and free-text job descriptions, as well as the long-context reasoning required to match candidates to roles. This study compares six large language models (LLMs): BERT‑base, DistilBERT, RoBERTa‑large, Phi‑3‑mini‑4k, Mistral‑7B, and Llama‑3‑8B. These models were fine‑tuned using parameter‑efficient methods, specifically Low‑Rank Adaptation (LoRA) and 4-bit-quantized Low‑Rank Adaptation (QLoRA). The evaluation covered two tasks: resume classification and resume–job description matching. We curate a 24-category corpus of 2,484 real resumes, augmented with 8,000 synthetic resumes generated by DeepSeek-V3 and ChatGPT-4, and define a unified evaluation protocol covering accuracy, macro-F1, Matthews Correlation Coefficient (MCC), ranking quality (normalized Discounted Cumulative Gain, nDCG@5), and a fairness audit using the demographic parity gap. LoRA‑tuned Llama‑3‑8B achieved the highest test macro‑F1 score of 0.913 and MCC of 0.875, outperforming the BERT‑base baseline by 8.2 F1 points while training only 3.7 million trainable parameters, just 0.046% of the backbone. An ablation study isolates the effects of LoRA rank, prompt formulation, and synthetic data ratio. A fairness audit shows that an in-context debiasing prompt reduces the gender demographic parity gap from 6.4% to 2.1%. These findings offer evidence-based guidance for Human Resource Management (HRM) on selecting, fine-tuning and auditing LLMs for production-grade resume-screening pipelines.

 

Keywords:  Large language models; resume screening; fine-tuning; LoRA; QLoRA.

Copyright © Journal of Engineering, Project, and Production Management (EPPM-Journal).

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License.

Requests for reprints and permissions at eppm.journal@gmail.com.

Citation: Liu, X. and Su, C. (2026). Fine-Tuning Large Language Models for Automated Resume Screening in Human Resource Management: A Comparative Performance Analysis. Journal of Engineering, Project, and Production Management, 16(5), 2026-0092.

DOI: 10.32738/JEPPM-2026-0092

Full Text


Copyright © EPPM-Journal. All rights reserved.