Skip to main navigation Skip to search Skip to main content

Compact vs Mid-Scale LLMs for Customer Support: A Deployment-Oriented Benchmark of Unified Classification and Response Generation

  • University of Limerick

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large Language Models (LLMs) are transforming customer support automation, yet deployment requires balancing response quality with the computational costs of fine-tuning and inference. This work investigates whether a compact model (TinyLlama-1.1B) can rival mid-scale models (Llama-2-7B, Mistral-7B) when adapted for unified customer support tasks. We evaluate a joint-task setting where a single model simultaneously generates agent responses and predicts intent and category labels for routing. To establish a rigorous baseline, we compare these against larger non-fine-tuned models, including Llama-3.1-8B, Falcon-40B, and Llama-2-70B. Experiments across two datasets, general customer service (CSD) and heterogeneous IT support (CITD), demonstrate that domain-specific fine-tuning is indispensable; even 70B-parameter base models fail to maintain structured routing and produce generic replies. After parameter-efficient fine-tuning (QLoRA), TinyLlama achieves generation quality (BLEU/ROUGE-L) competitive with 7B models and approaches their routing accuracy on the CSD dataset, while training up to 5× faster on a single GPU. Furthermore, an author-led error analysis identifies critical failure modes, such as procedural omissions and corrupted outputs, that automated metrics fail to capture. Our results suggest that for specialised support domains, compact fine-tuned LLMs offer a high-performance, resource-efficient alternative to larger architectures, making them ideal for practical, cost-sensitive deployments.

Original languageEnglish
Title of host publicationArtificial Intelligence Applications and Innovations - 22nd IFIP WG 12.5 International Conference, AIAI 2026, Proceedings
EditorsIlias Maglogiannis, Lazaros Iliadis, Antonios Papaleonidas, Michalis Zervakis
PublisherSpringer Science and Business Media Deutschland GmbH
Pages58-72
Number of pages15
ISBN (Print)9783032308009
DOIs
Publication statusPublished - 2027
Event22nd IFIP WG 12.5 International Conference on Artificial Intelligence Applications and Innovations, AIAI 2026 - Chania, Greece
Duration: 16 Jul 202619 Jul 2026

Publication series

NameIFIP Advances in Information and Communication Technology
Volume793 IFIPAICT
ISSN (Print)1868-4238
ISSN (Electronic)1868-422X

Conference

Conference22nd IFIP WG 12.5 International Conference on Artificial Intelligence Applications and Innovations, AIAI 2026
Country/TerritoryGreece
CityChania
Period16/07/2619/07/26

Keywords

  • Artificial Intelligence
  • Customer Support LLM
  • Large Language Models
  • Natural Language Processing
  • Small Language Models

Fingerprint

Dive into the research topics of 'Compact vs Mid-Scale LLMs for Customer Support: A Deployment-Oriented Benchmark of Unified Classification and Response Generation'. Together they form a unique fingerprint.

Cite this