Can Machines Detect Ultra-Processed Foods? A Head-To-Head Evaluation of Large Language Models Using NOVA Classification
| dc.authorid | https://orcid.org/0000-0002-7073-2907 | |
| dc.authorid | https://orcid.org/0000-0002-3356-7332 | |
| dc.authorid | https://orcid.org/0000-0001-7982-6988 | |
| dc.contributor.author | Bayram, Hatice Merve | |
| dc.contributor.author | Arslan, Sedat | |
| dc.contributor.author | Öztürkcan, Arda | |
| dc.date.accessioned | 2026-08-21T13:57:57Z | |
| dc.date.issued | 2026 | |
| dc.department | Sağlık Bilimleri Fakültesi | |
| dc.description.abstract | This cross-sectional study compared three large language models (LLMs) (Grok 4.1, Gemini 3, and ChatGPT 5.2) in classifying ultra-processed foods (UPF) using best-selling products from leading supermarket chains covering 53.2% of the national market. Of 3,001 products, 2,920 with complete ingredient information were included; two trained dietitians assigned NOVA groups as the reference standard. In the reference classification, 74.3% of products were UPF. Under the baseline prompt, all models underestimated UPF prevalence compared with the reference standard (p < .001). ChatGPT 5.2 yielded the highest binary UPF detection performance (accuracy: 69.01%; sensitivity: 59.01%; specificity: 98.00%; F1: 73.88%). Prompt sensitivity analyses revealed that a minimal prompt substantially outperformed the detailed baseline for most models (Gemini 3 F1: 94.20%; ChatGPT 5.2 F1: 92.62%), though Grok 4.1’s gain reflected a specificity trade-off (sensitivity: 97.79%; specificity: 21.09%). Low inter-run agreement (κ: 0.01–0.27) indicated sensitivity to model updates, supporting prompt calibration and human oversight. These findings suggest that off-the-shelf LLMs require prompt calibration and human oversight before UPF surveillance workflows. | |
| dc.identifier.citation | Hatice Merve Bayram, Sedat Arslan, Arda Ozturkcan, Can machines detect ultra-processed foods? A head-to-head evaluation of large language models using NOVA classification, International Journal of Food Science and Technology, Volume 61, Issue 2, 2026, vvag142, https://doi.org/10.1093/ijfood/vvag142 | |
| dc.identifier.doi | 10.1093/ijfood/vvag142 | |
| dc.identifier.issn | 0950-5423 | |
| dc.identifier.issue | 2 | |
| dc.identifier.scopus | 2-s2.0-105045939407 | |
| dc.identifier.scopusquality | Q1 | |
| dc.identifier.uri | https://hdl.handle.net/11363/12359 | |
| dc.identifier.volume | 61 | |
| dc.indekslendigikaynak | Scopus | |
| dc.institutionauthor | Bayram, Hatice Merve | |
| dc.institutionauthor | Öztürkcan, Arda | |
| dc.institutionauthorid | https://orcid.org/0000-0002-7073-2907 | |
| dc.institutionauthorid | https://orcid.org/0000-0001-7982-6988 | |
| dc.language.iso | en | |
| dc.publisher | Oxford University Press | |
| dc.relation.ispartof | International Journal of Food Science and Technology | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.subject | NOVA classification | |
| dc.subject | ultra-processed foods | |
| dc.subject | large language models | |
| dc.subject | food labelling | |
| dc.subject | supermarket foods | |
| dc.title | Can Machines Detect Ultra-Processed Foods? A Head-To-Head Evaluation of Large Language Models Using NOVA Classification | |
| dc.type | Article |










