- Research Article
- 10.1016/j.tifs.2026.105698
Embodied artificial intelligence in the food supply chain: Innovations, challenges, and future perspectives
- Jun 01, 2026
- Trends in Food Science & Technology
- Chongyu Wang + 11 more +11
Publications from 2021 to 2026
Showing 10 of 1,728 papers
Embodied artificial intelligence in the food supply chain: Innovations, challenges, and future perspectives
MTDMS-CMOEA: A Multi-Task dynamic membrane computing framework for Tri-Objective optimization of Latency, Energy, and security in IoT edge offloading
Could global warming cause a range expansion or shift of Lyme disease in the U.S. state of Maryland? A mathematical modelling approach
Abstract Lyme disease, transmitted by ticks, is endemic in several regions of the United States (U.S.) (including the Northeast), and the lifecycle of ticks is significantly affected by changes in local climatic variables. In this study, we modelled the dynamics of Lyme disease across the U.S. state of Maryland. We used a mechanistic model, calibrated using case and temperature data, to assess the effect of temperature fluctuations on the geospatial distribution and burden of Lyme disease across Maryland. Our results demonstrate that tick activity and Lyme disease intensity peak when the temperature reaches 17–20.5°C. We estimate that moderate projected global warming will cause a range expansion of Lyme disease, increasing burden in Central Maryland and extending risk into Western counties, while reducing the disease burden in Southern and most Eastern counties. High projected warming will cause a westward shift, with new Lyme disease hotspots emerging in Western counties, and reduction of burden in Central, Southern and Eastern regions. Maryland will experience reductions in overall Lyme disease burden under both projected global warming scenarios (with more reductions under the high warming scenario). Disease elimination is feasible using a hybrid strategy, which combines rodent baiting, habitat clearance and personal protection against tick bites, with moderate coverages.
Read moreHigh Performance Singular Value Decomposition on GPU Architectures
With the advancement of GPU architecture, matrix computation engines such as NVIDIA Tensor Cores now support double-precision (FP64) General matrix multiplications (GEMMs) with the same efficiency as single-precision (FP32) GEMMs. However, the adoption of this enhanced FP64 capability remains limited, primarily restricted to applications that involve multiple FP64 BLAS3 operations. Singular Value Decomposition (SVD), a fundamental decomposition in numerical linear algebra with numerous applications, can greatly benefit from exploiting this hardware feature. In this article, for FP32 SVD, we propose a novel algorithm, FP64 precision eigenvalue decomposition (EVD) based SVD, specifically designed to leverage the latest GPU architectural features. We provide a theoretical analysis demonstrating the feasibility of our approach on emerging GPU architectures and evaluate it from both accuracy and performance perspectives. Moreover, for FP64 SVD, we introduce a double-blocking band reduction technique combined with a GPU-based bulge chasing algorithm to further accelerate the overall SVD process. Experimental results show that, for FP32 SVD, our EVD-based SVD implementation achieves higher numerical accuracy and delivers speedups of up to 6.1× on H100 and 5.0× on A100 over the state-of-the-art cuSOLVER SVD solver. In the case of FP64 SVD, our method also achieves 4.9× and 4.8× speedups on H100 and A100, respectively. These results highlight the potential of our approach as a highly efficient and accurate solution for SVD on modern GPU platforms.
Read moreRCMoE: A Communication-Efficient Random Compression Framework for Resource-Constrained Mixture-of-Experts Training
Mixture-of-Experts (MoE) architecture with experts parallelism scales LLMs efficiently by activating only a subset of experts per input, avoiding proportional training costs. However, the intensive and heterogeneous communication substantially hinders the efficiency and scalability of MoE training in the resource-constrained scenario. Existing communication compression techniques fall short in MoE training due to: (i) Intensive training amplifies compression overhead, compromising training efficiency; (ii) Accumulated compression errors propagate through the network, degrading training quality. In this paper, we propose RCMoE, a communication-efficient Random Compression framework for MoE training with two core modules: (1) Local-Stochastic Quantization compresses the all-to-all communication by stochastically quantizing each row of the expert's intermediate computing results in parallel, effectively improving the compression efficiency and reducing compression error; (2) Probabilistic Thresholding Sparsification compresses the all-reduce communication by probabilistically sampling large gradients at high probability, thereby reducing the computational complexity and maintaining the convergence efficiency. Experiments on four typical MoE training tasks prove that RCMoE achieves higher 5.9x-8.1x total communication compression ratios and 1.3x-10.1x training speedup compared with the state-of-the-art compression techniques while maintaining the MoE training accuracy.
Read moreQuantifying the Potential to Escape Filter Bubbles: A Behavior-Aware Measure via Contrastive Simulation
Nowadays, recommendation systems have become crucial to online platforms, shaping user exposure by accurate preference modeling. However, such an exposure strategy can also reinforce users’ existing preferences, leading to a notorious phenomenon named filter bubbles. Given its negative effects, such as group polarization, increasing attention has been paid to exploring reasonable measures to filter bubbles. However, most existing evaluation metrics simply measure the diversity of user exposure, failing to distinguish between algorithmic preference modeling and actual information confinement. In view of this, we introduce Bubble Escape Potential (BEP), a behavior-aware measure that quantifies how easily users can escape from filter bubbles. Specifically, BEP leverages a contrastive simulation framework that assigns different behavioral tendencies (e.g., positive vs. negative) to synthetic users and compares the induced exposure patterns. This design enables decoupling the effect of filter bubbles and preference modeling, allowing for more precise diagnosis of bubble severity. We conduct extensive experiments across multiple recommendation models to examine the relationship between predictive accuracy and bubble escape potential across different groups. To the best of our knowledge, our empirical results are the first to quantitatively validate the dilemma between preferences modeling and filter bubbles. What's more, we observe a counter-intuitive phenomenon that mild random recommendations are ineffective in alleviating filter bubbles, which can offer a principled foundation for further work in this direction.
Read moreBeyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
The rapid evolution of generative technologies necessitates reliable methods for detecting AI-generated images. A critical limitation of current detectors is their failure to generalize to images from unseen generative models, as they often overfit to source-specific semantic cues rather than learning universal generative artifacts. To overcome this, we introduce a simple yet remarkably effective pixel-level mapping pre-processing step to disrupt the pixel value distribution of images and break the fragile, non-essential semantic patterns that detectors commonly exploit as shortcuts. This forces the detector to focus on more fundamental and generalizable high-frequency traces inherent to the image generation process. Through comprehensive experiments on GAN and diffusion-based generators, we show that our approach significantly boosts the cross-generator performance of state-of-the-art detectors. Extensive analysis further verifies our hypothesis that the disruption of semantic cues is the key to generalization.
Read morePathMind: A Retrieve-Prioritize-Reason Framework for Knowledge Graph Reasoning with Large Language Models
Knowledge graph reasoning (KGR) is the task of inferring new knowledge by performing logical deductions on knowledge graphs. Recently, large language models (LLMs) have demonstrated remarkable performance in complex reasoning tasks. Despite promising success, current LLM-based KGR methods still face two critical limitations. First, existing methods often extract reasoning paths indiscriminately, without assessing their different importance, which may introduce irrelevant noise that misleads LLMs. Second, while some methods leverage LLMs to dynamically explore potential reasoning paths, they require high retrieval demands and frequent LLM calls. To address these limitations, we propose PathMind, a novel framework designed to enhance faithful and interpretable reasoning by selectively guiding LLMs with important reasoning paths. Specifically, PathMind follows a "Retrieve-Prioritize-Reason" paradigm. First, it retrieves a query subgraph from KG through the retrieval module. Next, it introduces a path prioritization mechanism that identifies important reasoning paths using a semantic-aware path priority function, which simultaneously considers the accumulative cost and the estimated future cost for reaching the target. Finally, PathMind generates accurate and logically consistent responses via a dual-phase training strategy, including task-specific instruction tuning and path-wise preference alignment. Extensive experiments on benchmark datasets demonstrate that PathMind consistently outperforms competitive baselines, particularly on complex reasoning tasks with fewer input tokens, by identifying essential reasoning paths.
Read moreStop Mixing Things Up! BISCUIT Teaches Vision-Language Models to Learn New Concepts from Images on the Spot
Vision-Language Models (VLMs) have achieved impressive performance across various tasks, but often struggle to apply newly introduced visual concepts during inference. A common failure pattern is what we call Mixing Things Up: VLMs frequently confuse concept names, resulting in vague descriptions and failure to ground the concept correctly. Existing approaches mainly address person-related concepts through text prompts or tokenizer modifications. However, VLMs still miss or misinterpret untrained visual concepts, underscoring the need to learn new concepts directly from visual input, without relying on prior textual injection. To overcome these limitations, we propose BISCUIT (Basis-aligned Inference through Structured Concept Unification and Identification-aware Tuning), a two-step training method. Step I proposes a dual-stream structure-aware vision encoder that fuses RGB and edge-based embeddings within a shared basis space to enhance concept recognition. Step II enhances generation quality through identification-aware tuning, which encourages alignment between the generated text and the newly introduced visual concepts. Existing methods mainly focus on person concepts and lack comprehensive evaluation across diverse visual categories. We further propose a benchmark BiscuitVQA to evaluate VLMs performance on recognizing and applying novel image-introduced concepts across diverse concept types and task types, including real people, cartoons, animals, and symbolic content. We apply BISCUIT to LLaVA-1.5 and Qwen2.5-VL, achieving competitive results among open-source models and narrowing the gap to Gemini-2.5 and GPT-4o. Interestingly, our BISCUIT maintains strong generalization, showing minimal degradation on other downstream tasks.
Read moreParameter-, Memory-, Time-Efficient Multi-Task Dense Vision Adaptation
While adapting pretrained vision models to downstream dense prediction tasks is widely used, current methods often overlook adaptation efficiency, especially in the context of multi-task learning (MTL). Although parameter-efficient fine-tuning (PEFT) methods can enhance parameter efficiency, broader aspects such as GPU memory and training time efficiency remain underexplored. In this paper, we propose a new paradigm that simultaneously achieves efficiency in Parameters, GPU Memory, and Training Time for Multi-Task Dense Vision Adaptation. Specifically, we propose a dual-branch framework, in which a frozen pretrained backbone serves as the generic main branch, and the proposed Bi-Directional Task Adaptation (BDTA) modules are integrated in parallel to form a task bypass branch that extracts adaptation features required by multiple specific tasks. This adaptation module is lightweight, efficient, and does not require backpropagation through the large pre-trained backbone, thus avoiding resource-intensive gradient computations. Moreover, a Mixture of Task Experts mechanism (MoTE) is further proposed to integrate adaptation features across tasks and scales, thereby obtaining more robust representations tailored for dense prediction tasks. On the PASCAL-Context benchmark, our method achieves over 2× relative performance improvement compared to the best prior multi-task PEFT method, while using only ~30% of the parameters, ~50% of the memory, and ~60% of the training time, demonstrating superior overall adaptation efficiency.
Read more