PL EN
Analysis of the impact of natural language selection in Large Language Models (LLMs) on energy consumption
 
More details
Hide details
1
Faculty of Electrical Engineering and Computer Science, Lublin University of Technology
 
These authors had equal contribution to this work
 
 
Publication date: 2026-08-31
 
 
Corresponding author
Tomasz Szymczyk   

Faculty of Electrical Engineering and Computer Science, Lublin University of Technology
 
 
 
KEYWORDS
TOPICS
ABSTRACT
Large language model (LLM) inference is computationally intensive, yet the extent to which the interaction language affects energy consumption remains insufficiently understood. This study investigates whether inference energy consumption varies across six European languages: Polish, English, German, French, Spanish, and Russian. Three open-weight models, Llama3.1-8b, Gemma-3-12b, and Qwen2.5-14b, were evaluated using ten semantically equivalent prompts, each translated into the six languages and repeated three times. Experiments were conducted on a workstation-class server equipped with an AMD Radeon RX 7900 XT GPU and the Ollama inference framework. Whole-system power consumption was measured using a calibrated HIOKI PW3337 power meter, and energy efficiency was expressed as energy per generated token (EPT). Mean EPT varied substantially among models: Llama3.1-8b consumed 6.43 J/token, Qwen2.5-14b 10.23 J/token, and Gemma-3-12b 10.87 J/token. Kruskal–Wallis tests identified statistically significant between-language differences only for Llama3.1-8b; no significant differences were found for Qwen2.5-14b or Gemma-3-12b. However, regression analysis for Llama3.1-8b showed that response length was the dominant predictor of EPT, as fixed inference overhead was amortized over longer outputs. The findings indicate that apparent language-dependent differences in inference energy are driven primarily by generated-output length rather than by language itself. Energy efficiency also depends strongly on model architecture, tokenizer behavior, and inference implementation, highlighting the importance of response-length control and whole-system measurement in Green AI benchmarking.
Journals System - logo
Scroll to top