AI Signal 268 2 feeds carried it
LLM mushroom identification tested against expert-verified FungiTastic dataset of 2.8k species
An evaluation tests off-the-shelf LLMs on zero-shot mushroom species identification using the FungiTastic dataset, focusing on edible species legally sold in Poland and deadly species, with attention to where the two categories dangerously overlap.
Foraging safety depends on correct species identification, and LLMs are increasingly used as identification tools despite no domain-specific training. The overlap between edible and deadly lists, exemplified by Tricholoma equestre, highlights that even expert-verified datasets carry contradictions that no model can resolve without contextual judgment.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
The FungiTastic dataset contains 340k observations and 615k photos across 2.8k species, with labels verified by experts and partially by DNA sequencing, split by year for train/validation/test.
The evaluation uses a zero-shot approach on the 2023 test split, testing LLMs as-is without any fine-tuning.
Tricholoma equestre appears on both Poland's official sellable mushroom list and Wikipedia's deadly species list, illustrating a real-world contradiction that complicates automated classification.
THE READ
What the cluster adds up to.
The evaluation uses the FungiTastic dataset, built by Czech computer-vision researchers from citizen science data, containing 340k observations across 2.8k species with expert-verified labels. The dataset is split temporally, training data up to 2021, validation in 2022, and test in 2023, which means the LLM evaluation uses the 2023 test split with no prior exposure to those observations. This is a zero-shot setup: the LLMs receive no fine-tuning and are tested as-is, simulating how a forager might use a chatbot in the field.
The species selection is grounded in real foraging practice. The author uses Poland's official list of 47 sellable mushroom species (37 present in the dataset) and a Wikipedia-derived list of deadly species (20 present after intersection). This yields 55 species total, chosen to reflect what someone in a Polish forest would actually encounter. The choice of Poland is deliberate: mushroom hunting is a widespread cultural practice there, making the safety stakes concrete rather than hypothetical.
The most striking finding is the overlap between the edible and deadly lists at Tricholoma equestre (yellow knight), which is considered edible in some countries and deadly in others, with documented fatal incidents. This is not a model error but a labeling contradiction embedded in the source data itself. Any classifier, LLM or otherwise, trained on labels that disagree across geographic and cultural contexts will inherit that ambiguity, and a zero-shot model has no mechanism to surface it.
The dataset includes more than standard photos, masks, captions, microscopic images, and satellite imagery, but the evaluation restricts itself to the kind of photo a smartphone user would take. This scopes the test to the realistic deployment scenario: a person in a forest photographing a mushroom and asking a chatbot whether it is safe to eat. The gap between that scenario and the dataset's controlled, expert-labeled conditions is where the risk lies.
Only one feed carried this story, and the article extract is partial, so the full quantitative results of the evaluation are not available in the provided material. What is clear is the methodology and the framing: the author is not asking whether LLMs can identify mushrooms in principle, but whether they can be trusted with health-critical decisions given contradictory real-world data and no domain-specific training.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER