PL EN
Analysis of the impact of natural language selection in Large Language Models (LLMs) on energy consumption
 
Więcej
Ukryj
1
Faculty of Electrical Engineering and Computer Science, Lublin University of Technology
 
Zaznaczeni autorzy mieli równy wkład w przygotowanie tego artykułu
 
 
Data publikacji: 31-08-2026
 
 
Autor do korespondencji
Tomasz Szymczyk   

Faculty of Electrical Engineering and Computer Science, Lublin University of Technology
 
 
 
SŁOWA KLUCZOWE
DZIEDZINY
STRESZCZENIE
Large language model (LLM) inference is computationally intensive, yet the extent to which the interaction language affects energy consumption remains insufficiently understood. This study investigates whether inference energy consumption varies across six European languages: Polish, English, German, French, Spanish, and Russian. Three open-weight models, Llama3.1-8b, Gemma-3-12b, and Qwen2.5-14b, were evaluated using ten semantically equivalent prompts, each translated into the six languages and repeated three times. Experiments were conducted on a workstation-class server equipped with an AMD Radeon RX 7900 XT GPU and the Ollama inference framework. Whole-system power consumption was measured using a calibrated HIOKI PW3337 power meter, and energy efficiency was expressed as energy per generated token (EPT). Mean EPT varied substantially among models: Llama3.1-8b consumed 6.43 J/token, Qwen2.5-14b 10.23 J/token, and Gemma-3-12b 10.87 J/token. Kruskal–Wallis tests identified statistically significant between-language differences only for Llama3.1-8b; no significant differences were found for Qwen2.5-14b or Gemma-3-12b. However, regression analysis for Llama3.1-8b showed that response length was the dominant predictor of EPT, as fixed inference overhead was amortized over longer outputs. The findings indicate that apparent language-dependent differences in inference energy are driven primarily by generated-output length rather than by language itself. Energy efficiency also depends strongly on model architecture, tokenizer behavior, and inference implementation, highlighting the importance of response-length control and whole-system measurement in Green AI benchmarking.
Journals System - logo
Scroll to top