A comparison of local LLMs in paragraph decomposition context
Abstract
This study presents a comparative evaluation of three locally deployable large language models—LLaMA 3.1 8B, Mistral 7B, and Phi-3.5 Mini—on paragraph decomposition tasks using real-world student feedback. Unlike prior work that primarily targets large cloud-hosted models, this study operates entirely on local infrastructure using 4-bit quantization in a Google Colab environment. A standardized instruction-based prompting strategy with a single in-context example was applied across all models using a single experimental configuration with greedy decoding and a repetition penalty of ρ = 1.05. Outputs were manually evaluated against seven error categories and two caution labels covering structural, semantic, and stylistic dimensions. Across 285 generated sub-points, the models achieved an overall success rate of 60.0%. Mistral 7B attained the highest headline success rate (65.9%), yet introduced frequent language-level reformalization. LLaMA 3.1 8B produced the highest proportion of fully clean outputs (56.4%) but exhibited a strong tendency toward over-segmentation. Phi-3.5 Mini offered the fastest inference but showed the highest rate of style drift. Over-segmentation emerged as the dominant failure mode across all models, and decomposition quality declined consistently as the number of extracted points increased. These findings highlight the trade-offs between structural precision, linguistic faithfulness, and computational efficiency in local LLM-based text decomposition.
Authors: Nannicha Phraemetta, Nattawadee Wuttivoradit, Chayada Muangboonsri, Bunthit Watanapa, Vithida Chongsuphajaisiddhi