Il presente studio esplora l’affidabilità dei sistemi di traduzione automatica neurale (NMT) e dei modelli linguistici di grandi dimensioni (LLM) nell’ambito del discorso specialistico, con particolare attenzione al linguaggio giuridico. Gli strumenti oggetto di indagine sono rispettivamente DeepL e Google Gemini. Data la crescente diffusione di tali sistemi automatizzati in contesti professionali, questo lavoro indaga la loro efficacia nel gestire le insidie del linguaggio giuridico. Al fine di rispondere alla domanda di ricerca, viene condotta un’analisi comparativa sulle prestazioni dei due strumenti summenzionati, nonché della risposta di Google Gemini a diverse tecniche di prompting. Un contratto di locazione di riferimento costituisce il corpus primario della presente ricerca. Lo studio adotta un approccio ibrido, che combina un’analisi qualitativa degli errori e una valutazione quantitativa tramite la metrica automatica “METEOR”. Tale elaborato esordisce contestualizzando questi nuovi strumenti, offrendo una panoramica storica dello sviluppo dei sistemi di traduzione automatica, un approfondimento sulle complessità del linguaggio giuridico, e fornendo al lettore gli strumenti necessari per un utilizzo più efficace e consapevole di tali tecnologie. Successivamente, viene condotta un’analisi a coppie divisa in due fasi, che prevede un confronto tra gli output prodotti da DeepL e Google Gemini, seguito da una valutazione dell’impatto di due tecniche di prompting. Infine, vengono presentati i risultati, che consentono di trarre importanti conclusioni e di fare ulteriore luce su questi nuovi strumenti. I risultati emersi suggeriscono che, sebbene entrambi dimostrino un livello elevato di competenza generale, Google Gemini mostra prestazioni leggermente superiori rispetto a DeepL. In relazione alle tecniche di prompting avanzate, l’analisi quantitativa ha prodotto risultati inaspettati. Contrariamente alla letteratura esistente, la tecnica di prompting “Domain-specific Prompt” ha ottenuto il punteggio più alto su METEOR. Tuttavia, il Self-Guided CoT/ToT si è dimostrato qualitativamente più efficace. Questo aspetto viene esaminato nel dettaglio esplorando i limiti della metrica automatica adottata. Inoltre, l’analisi rivela una limitazione significativa presentata dal modello linguistico di grandi dimensioni di riferimento. Di fronte a prompt specialistici, Google Gemini tende a riassumere anziché fornire la traduzione integrale del testo sorgente, influenzando negativamente la qualità dell’output, soprattutto all’interno dell’ambito giuridico, dove la precisione è imprescindibile. In conclusione, vengono offerti ulteriori spunti sul ruolo cruciale svolto dal post-editing umano. Pur riconoscendo l’efficacia di tali sistemi, si riscontra che essi siano ancora privi dell’affidabilità necessaria per un utilizzo autonomo. In ultima analisi, viene sottolineato come l’intervento umano rimanga fondamentale per garantire la qualità della traduzione giuridica automatizzata.

This dissertation explores the reliability of Neural Machine Translation (NMT) systems and Large Language Models (LLMs) when applied to specialised discourse, with specific focus on legal language. The tools under investigation are respectively DeepL and Google Gemini. Given the increasing reliance on automated tools in professional settings, the present study investigates whether these systems can effectively handle the pitfalls of legal language. To answer the research question, a comparative analysis on the performance of the two aforementioned prominent tools, as well as on the response of Google Gemini to different prompting techniques, is conducted. A reference contract, specifically a lease agreement, serves as the primary dataset. The study employs a hybrid approach, combining human-led qualitative error analysis with an automated quantitative evaluation by means of automatic metric “METEOR”. The research begins by contextually framing these new tools, providing a historical overview of MT systems’ development alongside insights into the intricacies of legal language and by equipping the reader with the necessary instruments to rely on these machines more effectively. Subsequently, a two-stage pairwise analysis is carried out, involving a direct comparison between DeepL’s and Google Gemini’s outputs, followed by an evaluation of the impact of two prompting techniques. Finally, results are displayed, allowing for drawing conclusions and shedding further light on these new tools. Findings suggest that while both systems demonstrate a high level of general proficiency, Google Gemini slightly outperforms DeepL. With regard to the advanced prompting techniques, quantitative analysis has yielded unexpected results. Contrary to the existing literature, Domain-specific Prompt scores the highest on METEOR. However, Self-Guided CoT/ToT proves to be the qualitatively most effective. This aspect is discussed in detail by exploring the shortcomings of the implemented automatic metric. Furthermore, the analysis reveals a significant limitation presented by the representative LLM. When faced with domain-specific prompts, Google Gemini tends to summarise rather than providing a full translation of the source text, negatively affecting the output quality, especially within the legal domain, where precision is critical. The conclusion offers further insights into the pivotal role played by human-post-editing. Whereas the effectiveness of such systems is recognised, they are still found to lack the necessary reliability for autonomous use without human supervision. Ultimately, it is underscored that human intervention remains crucial to validate automated legal translation.

Reliability and performance of Machine Translation systems in specialised discourse: A case study on legal language

FALCONE, MIRIANA
2025/2026

Abstract

Il presente studio esplora l’affidabilità dei sistemi di traduzione automatica neurale (NMT) e dei modelli linguistici di grandi dimensioni (LLM) nell’ambito del discorso specialistico, con particolare attenzione al linguaggio giuridico. Gli strumenti oggetto di indagine sono rispettivamente DeepL e Google Gemini. Data la crescente diffusione di tali sistemi automatizzati in contesti professionali, questo lavoro indaga la loro efficacia nel gestire le insidie del linguaggio giuridico. Al fine di rispondere alla domanda di ricerca, viene condotta un’analisi comparativa sulle prestazioni dei due strumenti summenzionati, nonché della risposta di Google Gemini a diverse tecniche di prompting. Un contratto di locazione di riferimento costituisce il corpus primario della presente ricerca. Lo studio adotta un approccio ibrido, che combina un’analisi qualitativa degli errori e una valutazione quantitativa tramite la metrica automatica “METEOR”. Tale elaborato esordisce contestualizzando questi nuovi strumenti, offrendo una panoramica storica dello sviluppo dei sistemi di traduzione automatica, un approfondimento sulle complessità del linguaggio giuridico, e fornendo al lettore gli strumenti necessari per un utilizzo più efficace e consapevole di tali tecnologie. Successivamente, viene condotta un’analisi a coppie divisa in due fasi, che prevede un confronto tra gli output prodotti da DeepL e Google Gemini, seguito da una valutazione dell’impatto di due tecniche di prompting. Infine, vengono presentati i risultati, che consentono di trarre importanti conclusioni e di fare ulteriore luce su questi nuovi strumenti. I risultati emersi suggeriscono che, sebbene entrambi dimostrino un livello elevato di competenza generale, Google Gemini mostra prestazioni leggermente superiori rispetto a DeepL. In relazione alle tecniche di prompting avanzate, l’analisi quantitativa ha prodotto risultati inaspettati. Contrariamente alla letteratura esistente, la tecnica di prompting “Domain-specific Prompt” ha ottenuto il punteggio più alto su METEOR. Tuttavia, il Self-Guided CoT/ToT si è dimostrato qualitativamente più efficace. Questo aspetto viene esaminato nel dettaglio esplorando i limiti della metrica automatica adottata. Inoltre, l’analisi rivela una limitazione significativa presentata dal modello linguistico di grandi dimensioni di riferimento. Di fronte a prompt specialistici, Google Gemini tende a riassumere anziché fornire la traduzione integrale del testo sorgente, influenzando negativamente la qualità dell’output, soprattutto all’interno dell’ambito giuridico, dove la precisione è imprescindibile. In conclusione, vengono offerti ulteriori spunti sul ruolo cruciale svolto dal post-editing umano. Pur riconoscendo l’efficacia di tali sistemi, si riscontra che essi siano ancora privi dell’affidabilità necessaria per un utilizzo autonomo. In ultima analisi, viene sottolineato come l’intervento umano rimanga fondamentale per garantire la qualità della traduzione giuridica automatizzata.
2025
This dissertation explores the reliability of Neural Machine Translation (NMT) systems and Large Language Models (LLMs) when applied to specialised discourse, with specific focus on legal language. The tools under investigation are respectively DeepL and Google Gemini. Given the increasing reliance on automated tools in professional settings, the present study investigates whether these systems can effectively handle the pitfalls of legal language. To answer the research question, a comparative analysis on the performance of the two aforementioned prominent tools, as well as on the response of Google Gemini to different prompting techniques, is conducted. A reference contract, specifically a lease agreement, serves as the primary dataset. The study employs a hybrid approach, combining human-led qualitative error analysis with an automated quantitative evaluation by means of automatic metric “METEOR”. The research begins by contextually framing these new tools, providing a historical overview of MT systems’ development alongside insights into the intricacies of legal language and by equipping the reader with the necessary instruments to rely on these machines more effectively. Subsequently, a two-stage pairwise analysis is carried out, involving a direct comparison between DeepL’s and Google Gemini’s outputs, followed by an evaluation of the impact of two prompting techniques. Finally, results are displayed, allowing for drawing conclusions and shedding further light on these new tools. Findings suggest that while both systems demonstrate a high level of general proficiency, Google Gemini slightly outperforms DeepL. With regard to the advanced prompting techniques, quantitative analysis has yielded unexpected results. Contrary to the existing literature, Domain-specific Prompt scores the highest on METEOR. However, Self-Guided CoT/ToT proves to be the qualitatively most effective. This aspect is discussed in detail by exploring the shortcomings of the implemented automatic metric. Furthermore, the analysis reveals a significant limitation presented by the representative LLM. When faced with domain-specific prompts, Google Gemini tends to summarise rather than providing a full translation of the source text, negatively affecting the output quality, especially within the legal domain, where precision is critical. The conclusion offers further insights into the pivotal role played by human-post-editing. Whereas the effectiveness of such systems is recognised, they are still found to lack the necessary reliability for autonomous use without human supervision. Ultimately, it is underscored that human intervention remains crucial to validate automated legal translation.
Machine Translation
Legal discourse
DeepL
Google Gemini
Human Post-Editing
File in questo prodotto:
File Dimensione Formato  
Falcone.Miriana.pdf

Accesso riservato

Dimensione 1.92 MB
Formato Adobe PDF
1.92 MB Adobe PDF

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14251/7061