TY - GEN
T1 - How Summary Length Affects ROUGE Scores
T2 - 2nd IEEE International Conference on Advances in Data-Driven Analytics and Intelligent Systems, ADACIS 2025
AU - Rhazzafe, Soukaina
AU - Colreavy-Donnelly, Simon
AU - Nikolov, Nikola S.
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Extractive summarization is a technique that selects key sentences directly from a document to create a shorter version while preserving its core meaning. Since this technique reuses original text, it reduces the risk of introducing factual errors, an important advantage in clinical contexts where accuracy is critical. In this work, we apply a BERT extractive summarizer to clinical case reports, aiming to study how the number of retained sentences affects summary quality. We evaluated summaries of three different lengths, 3, 7 and 10 sentences, using three variants of ROUGE: ROUGE-1, ROUGE-2 and ROUGE-L metrics. Our results show a consistent trade-off: recall increases with summary length, while precision decreases. ROUGE-N F-scores peaked at seven sentences, suggesting this length offers the best balance between informativeness and conciseness. However, ROUGE-L scores declined steadily as more sentences were added, indicating that longer summaries may disrupt the structural coherence and narrative flow of the original text. These findings offer practical insights for optimizing extractive summarization in clinical applications, where summaries must be both accurate and easy to process.
AB - Extractive summarization is a technique that selects key sentences directly from a document to create a shorter version while preserving its core meaning. Since this technique reuses original text, it reduces the risk of introducing factual errors, an important advantage in clinical contexts where accuracy is critical. In this work, we apply a BERT extractive summarizer to clinical case reports, aiming to study how the number of retained sentences affects summary quality. We evaluated summaries of three different lengths, 3, 7 and 10 sentences, using three variants of ROUGE: ROUGE-1, ROUGE-2 and ROUGE-L metrics. Our results show a consistent trade-off: recall increases with summary length, while precision decreases. ROUGE-N F-scores peaked at seven sentences, suggesting this length offers the best balance between informativeness and conciseness. However, ROUGE-L scores declined steadily as more sentences were added, indicating that longer summaries may disrupt the structural coherence and narrative flow of the original text. These findings offer practical insights for optimizing extractive summarization in clinical applications, where summaries must be both accurate and easy to process.
KW - BERT Summarizer
KW - Clinical Case Reports
KW - Extractive Text Summarization
KW - ROUGE
KW - Summary Length
UR - https://www.scopus.com/pages/publications/105035997545
U2 - 10.1109/ADACIS65663.2025.11436781
DO - 10.1109/ADACIS65663.2025.11436781
M3 - Conference contribution
AN - SCOPUS:105035997545
T3 - Proceedings - 2025 IEEE International Conference on Advances in Data-Driven Analytics and Intelligent Systems, ADACIS 2025
BT - Proceedings - 2025 IEEE International Conference on Advances in Data-Driven Analytics and Intelligent Systems, ADACIS 2025
A2 - Kanzari, Dalel
A2 - Essalih, Mohamed
A2 - Madani, Kurosh
A2 - Marques, Rui
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 20 November 2025 through 22 November 2025
ER -