Adverse Outcome Pathways (AOPs) are a conceptual framework that organises toxicological knowledge into structured causal chains linking a chemical stressor to an organism-level adverse effect, but their construction remains an expensive, expert-driven process, with the relevant evidence scattered across thousands of unstructured publications. This thesis designs and evaluates an end-to-end pipeline that addresses this gap with large language models, under the constraint of scarce annotated training data. To overcome annotation scarcity, a small set of manually curated examples is combined with weak labels generated by frontier models (GPT-5 and Gemini 2.5), yielding a corpus of roughly 3,450 events across 100 papers. Each weak label is filtered through a scorer that combines multi-granularity fuzzy matching with chemical anchoring, grounding each extracted event in its source text and discarding unverifiable ones. The corpus is then used to fine-tune, via LoRA, three open-source instruction-tuned models below ten billion parameters (BioMistral-7B, Mistral-7B-Instruct-v0.3, and Llama-3.1-8B-Instruct), chosen to isolate domain fit, instruction-following, and model recency and scale. Beyond the modelling experiments, the work provides a complete system spanning PDF-to-Markdown preprocessing, dataset construction, event extraction, grounding-based post-processing, and a desktop viewer that traces each event to its source passage. In the zero-shot regime all models fail, either over-generating or producing malformed output, while fine-tuning resolves both problems. All three models converge to broadly similar performance, so biomedical pretraining offers no clear advantage. Among them, the more recent and slightly larger Llama-3.1-8B-Instruct performs best. Under five-fold cross-validation it reaches a mean paper-level F1 of 0.43 ± 0.05. Evaluated on an independent, manually curated gold set, this model maintains its cross-validated performance and competes with frontier systems, indicating extraction rather than mere imitation of the teacher labels. The resulting pipeline establishes the feasibility of automated extraction and lays a working basis for a decision-support tool that could shorten AOP reconstruction in regulatory toxicology.
Automated extraction of toxicological Adverse Outcome Pathway events using fine-tuned LLMs
LUGLI, DAVIDE
2025/2026
Abstract
Adverse Outcome Pathways (AOPs) are a conceptual framework that organises toxicological knowledge into structured causal chains linking a chemical stressor to an organism-level adverse effect, but their construction remains an expensive, expert-driven process, with the relevant evidence scattered across thousands of unstructured publications. This thesis designs and evaluates an end-to-end pipeline that addresses this gap with large language models, under the constraint of scarce annotated training data. To overcome annotation scarcity, a small set of manually curated examples is combined with weak labels generated by frontier models (GPT-5 and Gemini 2.5), yielding a corpus of roughly 3,450 events across 100 papers. Each weak label is filtered through a scorer that combines multi-granularity fuzzy matching with chemical anchoring, grounding each extracted event in its source text and discarding unverifiable ones. The corpus is then used to fine-tune, via LoRA, three open-source instruction-tuned models below ten billion parameters (BioMistral-7B, Mistral-7B-Instruct-v0.3, and Llama-3.1-8B-Instruct), chosen to isolate domain fit, instruction-following, and model recency and scale. Beyond the modelling experiments, the work provides a complete system spanning PDF-to-Markdown preprocessing, dataset construction, event extraction, grounding-based post-processing, and a desktop viewer that traces each event to its source passage. In the zero-shot regime all models fail, either over-generating or producing malformed output, while fine-tuning resolves both problems. All three models converge to broadly similar performance, so biomedical pretraining offers no clear advantage. Among them, the more recent and slightly larger Llama-3.1-8B-Instruct performs best. Under five-fold cross-validation it reaches a mean paper-level F1 of 0.43 ± 0.05. Evaluated on an independent, manually curated gold set, this model maintains its cross-validated performance and competes with frontier systems, indicating extraction rather than mere imitation of the teacher labels. The resulting pipeline establishes the feasibility of automated extraction and lays a working basis for a decision-support tool that could shorten AOP reconstruction in regulatory toxicology.| File | Dimensione | Formato | |
|---|---|---|---|
|
Lugli.Davide.pdf
Accesso riservato
Dimensione
553.9 kB
Formato
Adobe PDF
|
553.9 kB | Adobe PDF |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14251/7321