Adverse Outcome Pathways (AOPs) are a conceptual framework that organises toxicological knowledge into structured causal chains linking a chemical stressor to an organism-level adverse effect, but their construction remains an expensive, expert-driven process, with the relevant evidence scattered across thousands of unstructured publications. This thesis designs and evaluates an end-to-end pipeline that addresses this gap with large language models, under the constraint of scarce annotated training data. To overcome annotation scarcity, a small set of manually curated examples is combined with weak labels generated by frontier models (GPT-5 and Gemini 2.5), yielding a corpus of roughly 3,450 events across 100 papers. Each weak label is filtered through a scorer that combines multi-granularity fuzzy matching with chemical anchoring, grounding each extracted event in its source text and discarding unverifiable ones. The corpus is then used to fine-tune, via LoRA, three open-source instruction-tuned models below ten billion parameters (BioMistral-7B, Mistral-7B-Instruct-v0.3, and Llama-3.1-8B-Instruct), chosen to isolate domain fit, instruction-following, and model recency and scale. Beyond the modelling experiments, the work provides a complete system spanning PDF-to-Markdown preprocessing, dataset construction, event extraction, grounding-based post-processing, and a desktop viewer that traces each event to its source passage. In the zero-shot regime all models fail, either over-generating or producing malformed output, while fine-tuning resolves both problems. All three models converge to broadly similar performance, so biomedical pretraining offers no clear advantage. Among them, the more recent and slightly larger Llama-3.1-8B-Instruct performs best. Under five-fold cross-validation it reaches a mean paper-level F1 of 0.43 ± 0.05. Evaluated on an independent, manually curated gold set, this model maintains its cross-validated performance and competes with frontier systems, indicating extraction rather than mere imitation of the teacher labels. The resulting pipeline establishes the feasibility of automated extraction and lays a working basis for a decision-support tool that could shorten AOP reconstruction in regulatory toxicology.

Automated extraction of toxicological Adverse Outcome Pathway events using fine-tuned LLMs

LUGLI, DAVIDE
2025/2026

Abstract

Adverse Outcome Pathways (AOPs) are a conceptual framework that organises toxicological knowledge into structured causal chains linking a chemical stressor to an organism-level adverse effect, but their construction remains an expensive, expert-driven process, with the relevant evidence scattered across thousands of unstructured publications. This thesis designs and evaluates an end-to-end pipeline that addresses this gap with large language models, under the constraint of scarce annotated training data. To overcome annotation scarcity, a small set of manually curated examples is combined with weak labels generated by frontier models (GPT-5 and Gemini 2.5), yielding a corpus of roughly 3,450 events across 100 papers. Each weak label is filtered through a scorer that combines multi-granularity fuzzy matching with chemical anchoring, grounding each extracted event in its source text and discarding unverifiable ones. The corpus is then used to fine-tune, via LoRA, three open-source instruction-tuned models below ten billion parameters (BioMistral-7B, Mistral-7B-Instruct-v0.3, and Llama-3.1-8B-Instruct), chosen to isolate domain fit, instruction-following, and model recency and scale. Beyond the modelling experiments, the work provides a complete system spanning PDF-to-Markdown preprocessing, dataset construction, event extraction, grounding-based post-processing, and a desktop viewer that traces each event to its source passage. In the zero-shot regime all models fail, either over-generating or producing malformed output, while fine-tuning resolves both problems. All three models converge to broadly similar performance, so biomedical pretraining offers no clear advantage. Among them, the more recent and slightly larger Llama-3.1-8B-Instruct performs best. Under five-fold cross-validation it reaches a mean paper-level F1 of 0.43 ± 0.05. Evaluated on an independent, manually curated gold set, this model maintains its cross-validated performance and competes with frontier systems, indicating extraction rather than mere imitation of the teacher labels. The resulting pipeline establishes the feasibility of automated extraction and lays a working basis for a decision-support tool that could shorten AOP reconstruction in regulatory toxicology.
2025
LLM
Toxicology
AOP
AI
Extraction
File in questo prodotto:
File Dimensione Formato  
Lugli.Davide.pdf

Accesso riservato

Dimensione 553.9 kB
Formato Adobe PDF
553.9 kB Adobe PDF

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14251/7321