TY - GEN
T1 - Compact vs Mid-Scale LLMs for Customer Support
T2 - 22nd IFIP WG 12.5 International Conference on Artificial Intelligence Applications and Innovations, AIAI 2026
AU - Khan, Talhat
AU - Ryan, Conor
N1 - Publisher Copyright:
© IFIP International Federation for Information Processing 2027.
PY - 2027
Y1 - 2027
N2 - Large Language Models (LLMs) are transforming customer support automation, yet deployment requires balancing response quality with the computational costs of fine-tuning and inference. This work investigates whether a compact model (TinyLlama-1.1B) can rival mid-scale models (Llama-2-7B, Mistral-7B) when adapted for unified customer support tasks. We evaluate a joint-task setting where a single model simultaneously generates agent responses and predicts intent and category labels for routing. To establish a rigorous baseline, we compare these against larger non-fine-tuned models, including Llama-3.1-8B, Falcon-40B, and Llama-2-70B. Experiments across two datasets, general customer service (CSD) and heterogeneous IT support (CITD), demonstrate that domain-specific fine-tuning is indispensable; even 70B-parameter base models fail to maintain structured routing and produce generic replies. After parameter-efficient fine-tuning (QLoRA), TinyLlama achieves generation quality (BLEU/ROUGE-L) competitive with 7B models and approaches their routing accuracy on the CSD dataset, while training up to 5× faster on a single GPU. Furthermore, an author-led error analysis identifies critical failure modes, such as procedural omissions and corrupted outputs, that automated metrics fail to capture. Our results suggest that for specialised support domains, compact fine-tuned LLMs offer a high-performance, resource-efficient alternative to larger architectures, making them ideal for practical, cost-sensitive deployments.
AB - Large Language Models (LLMs) are transforming customer support automation, yet deployment requires balancing response quality with the computational costs of fine-tuning and inference. This work investigates whether a compact model (TinyLlama-1.1B) can rival mid-scale models (Llama-2-7B, Mistral-7B) when adapted for unified customer support tasks. We evaluate a joint-task setting where a single model simultaneously generates agent responses and predicts intent and category labels for routing. To establish a rigorous baseline, we compare these against larger non-fine-tuned models, including Llama-3.1-8B, Falcon-40B, and Llama-2-70B. Experiments across two datasets, general customer service (CSD) and heterogeneous IT support (CITD), demonstrate that domain-specific fine-tuning is indispensable; even 70B-parameter base models fail to maintain structured routing and produce generic replies. After parameter-efficient fine-tuning (QLoRA), TinyLlama achieves generation quality (BLEU/ROUGE-L) competitive with 7B models and approaches their routing accuracy on the CSD dataset, while training up to 5× faster on a single GPU. Furthermore, an author-led error analysis identifies critical failure modes, such as procedural omissions and corrupted outputs, that automated metrics fail to capture. Our results suggest that for specialised support domains, compact fine-tuned LLMs offer a high-performance, resource-efficient alternative to larger architectures, making them ideal for practical, cost-sensitive deployments.
KW - Artificial Intelligence
KW - Customer Support LLM
KW - Large Language Models
KW - Natural Language Processing
KW - Small Language Models
UR - https://www.scopus.com/pages/publications/105045690623
U2 - 10.1007/978-3-032-30801-6_5
DO - 10.1007/978-3-032-30801-6_5
M3 - Conference contribution
AN - SCOPUS:105045690623
SN - 9783032308009
T3 - IFIP Advances in Information and Communication Technology
SP - 58
EP - 72
BT - Artificial Intelligence Applications and Innovations - 22nd IFIP WG 12.5 International Conference, AIAI 2026, Proceedings
A2 - Maglogiannis, Ilias
A2 - Iliadis, Lazaros
A2 - Papaleonidas, Antonios
A2 - Zervakis, Michalis
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 16 July 2026 through 19 July 2026
ER -