![]() |
|
|
|
Journal of Engineering, Project, and Production Management, 2026, 16(5), 2026-0092
Fine-Tuning Large Language Models for Automated Resume Screening in Human Resource Management: A Comparative Performance Analysis
1
Associate Senior Engineer, Information Data Resources Management
Department, Qingdao Human Resources Development Research and Promotion
Center, Qingdao, Shandong Province, 266001, China,
Project Management
Received June 6, 2026; revised June 30, 2026; accepted July 10, 2026
Available online July 28, 2026
Abstract: In today's fast‑paced recruitment landscape, the growing number of online applications has pushed manual resume screening to its limits, creating significant challenges in time, cost, and consistency for human resource professionals. Although classical machine-learning pipelines and earlier transformer baselines partially automate this task, they struggle with the unstructured nature of curricula vitae and free-text job descriptions, as well as the long-context reasoning required to match candidates to roles. This study compares six large language models (LLMs): BERT‑base, DistilBERT, RoBERTa‑large, Phi‑3‑mini‑4k, Mistral‑7B, and Llama‑3‑8B. These models were fine‑tuned using parameter‑efficient methods, specifically Low‑Rank Adaptation (LoRA) and 4-bit-quantized Low‑Rank Adaptation (QLoRA). The evaluation covered two tasks: resume classification and resume–job description matching. We curate a 24-category corpus of 2,484 real resumes, augmented with 8,000 synthetic resumes generated by DeepSeek-V3 and ChatGPT-4, and define a unified evaluation protocol covering accuracy, macro-F1, Matthews Correlation Coefficient (MCC), ranking quality (normalized Discounted Cumulative Gain, nDCG@5), and a fairness audit using the demographic parity gap. LoRA‑tuned Llama‑3‑8B achieved the highest test macro‑F1 score of 0.913 and MCC of 0.875, outperforming the BERT‑base baseline by 8.2 F1 points while training only 3.7 million trainable parameters, just 0.046% of the backbone. An ablation study isolates the effects of LoRA rank, prompt formulation, and synthetic data ratio. A fairness audit shows that an in-context debiasing prompt reduces the gender demographic parity gap from 6.4% to 2.1%. These findings offer evidence-based guidance for Human Resource Management (HRM) on selecting, fine-tuning and auditing LLMs for production-grade resume-screening pipelines.
Keywords: Large language models; resume screening; fine-tuning; LoRA; QLoRA. Copyright © Journal of Engineering, Project, and Production Management (EPPM-Journal). This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. Requests for reprints and permissions at eppm.journal@gmail.com. Citation: Liu, X. and Su, C. (2026). Fine-Tuning Large Language Models for Automated Resume Screening in Human Resource Management: A Comparative Performance Analysis. Journal of Engineering, Project, and Production Management, 16(5), 2026-0092.
|