An end-to-end vision–language framework with LLM-based reasoning for decision-oriented tomato leaf disease management in greenhouse environments

Abstract

Tomato foliar diseases pose a persistent threat to global food security, causing substantial yield losses and relying heavily on subjective, labor-intensive manual inspection. Although deep learning models achieve high accuracy under controlled laboratory conditions, their performance degrades markedly in real agricultural environments due to visual variability, background clutter, and inconsistent illumination. Moreover, existing automated systems fail to reflect how agricultural experts diagnose diseases by jointly considering global plant condition and localized symptom patterns. This paper presents an end-to-end framework for tomato leaf disease management that integrates vision–language disease recognition, lesion-aware representation learning, spatial aggregation, and large language model (LLM)–based reasoning for decision-oriented greenhouse deployment. At the core of the framework is the Dual-Level Contrastive Alignment Framework (DLCAF), which mirrors expert diagnostic reasoning by combining Global Disease-Description Alignment to accommodate linguistic and regional variability with Local Symptom Consistency Learning to focus on disease-specific lesions using expert annotations. Leaf-level predictions are aggregated into row- and greenhouse-level indicators to enable spatial analysis and temporal monitoring, which are subsequently interpreted by an LLM to generate actionable management recommendations. Extensive evaluation across laboratory, field, and greenhouse environments demonstrates the effectiveness of the proposed approach. DLCAF achieves 58.1% accuracy on PlantVillage, 72.5% under natural field conditions on FieldPlant, and 36.2% on the Taiwan Tomato Leaf Disease Dataset, consistently outperforming state-of-the-art vision–language models including CLIP, LongCLIP, and AgriCLIP. Additional validation on a newly collected real-world greenhouse dataset from Abu Dhabi (502 images) further confirms robust performance under practical deployment conditions, achieving 74.3% accuracy. Ablation studies show that expert-guided lesion supervision and multi-description semantic alignment provide more informative learning signals than random sampling or single-caption training. These results demonstrate that integrating global semantic understanding, localized symptom analysis, and LLM-based reasoning enables robust, interpretable, and practically deployable disease detection and decision-support systems for precision agriculture. To support reproducibility and future research, the implementation code is publicly available at GitHub, along with our collected AD Dataset released at Figshare. Keywords: Vision–language models; Contrastive learning; Lesion-aware learning; Agricultural decision support; Image–text alignment

Keywords

Citation

Shafay, M., Velayudhan, D., Owais, M., Hassan, T., Seneviratne, L., Hussain, I., & Werghi, N. (2026). An end-to-end vision–language framework with LLM-based reasoning for decision-oriented tomato leaf disease management in greenhouse environments. Computers and Electronics in Agriculture, 247, 111692.

Endorsement

Review

Supplemented By

Referenced By